On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
The paper argues that on-policy generalization can be carried into SFT by constraining updates to an on-policy-derived direction.
The authors find that SFT tends to move parameters in consistent directions, while on-policy training keeps adjusting those directions over time. They propose OPSFT, a supervised fine-tuning method that uses directions identified from on-policy training to constrain later updates. In their account, a few on-policy steps can identify useful directions, after which SFT can train more efficiently while preserving generalization benefits. Source: HF Daily Papers' note
The authors find that SFT tends to move parameters in consistent directions, while on-policy training keeps adjusting those directions over time. They propose OPSFT, a supervised fine-tuning method that uses directions identified from on-policy training to constrain later updates. In their account, a few on-policy steps can identify useful directions, after which SFT can train more efficiently while preserving generalization benefits. Source: HF Daily Papers' note
score 5