Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty
The paper proposes a training penalty meant to stop reasoning fine-tuning from pushing models away from safety behavior.
The authors study reasoning-induced misalignment in representation space and identify separate activation directions for reasoning ability and safety behavior. They argue those directions are coupled: reasoning fine-tuning can shift safety representations, and larger shifts correspond to worse safety outcomes. Their Safety-Direction Penalty limits movement along the learned safety direction during reasoning fine-tuning, with layer diagnostics used to expand the penalty where needed. On Qwen2.5-3B and 7B, they report restored safety while keeping benchmark reasoning performance. ArXiv · AI/CL/LG's note
The authors study reasoning-induced misalignment in representation space and identify separate activation directions for reasoning ability and safety behavior. They argue those directions are coupled: reasoning fine-tuning can shift safety representations, and larger shifts correspond to worse safety outcomes. Their Safety-Direction Penalty limits movement along the learned safety direction during reasoning fine-tuning, with layer diagnostics used to expand the penalty where needed. On Qwen2.5-3B and 7B, they report restored safety while keeping benchmark reasoning performance. ArXiv · AI/CL/LG's note
score 5