When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
Benign fine-tuning can break refusal behavior by re-sharpening an output-side safety-routing path.
The paper argues against gradient conflict as the main explanation for this fragility, proposing a Fisher-geometric account instead. It says safety behavior sits in a low-rank geometry that alignment flattens while leaving an output-routing mechanism intact. After 100 benign fine-tuning examples, that route is selectively sharpened in output-side MLP modules, letting attack success rise while general utility only slips mildly. The authors also report that a few safety examples can restore refusals, and that LoRA and ASAM delay early collapse but weaken at larger fine-tuning scales. ArXiv · AI/CL/LG's note
The paper argues against gradient conflict as the main explanation for this fragility, proposing a Fisher-geometric account instead. It says safety behavior sits in a low-rank geometry that alignment flattens while leaving an output-routing mechanism intact. After 100 benign fine-tuning examples, that route is selectively sharpened in output-side MLP modules, letting attack success rise while general utility only slips mildly. The authors also report that a few safety examples can restore refusals, and that LoRA and ASAM delay early collapse but weaken at larger fine-tuning scales. ArXiv · AI/CL/LG's note
score 5