Megadose Built for builders and researchers.

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

· HF Daily Papers ·
The paper argues that on-policy generalization can be carried into SFT by constraining updates to an on-policy-derived direction.

The authors find that SFT tends to move parameters in consistent directions, while on-policy training keeps adjusting those directions over time. They propose OPSFT, a supervised fine-tuning method that uses directions identified from on-policy training to constrain later updates. In their account, a few on-policy steps can identify useful directions, after which SFT can train more efficiently while preserving generalization benefits. Source: HF Daily Papers' note

score 5

Categories: Research