Megadose Built for builders and researchers.

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

· HF Daily Papers ·
The paper finds rollout policy is not the main driver of distillation behavior.

In controlled strong-to-weak distillation tests across Llama3 and Qwen2.5 models, token-level KL direction had clearer effects on task performance and output coverage. Learning rate was more tied to forgetting and update sparsity. Forward KL stayed strong across rollout policies, while reverse KL was more sensitive and favored student-generated rollouts. On-policy data helped on harder Countdown variants, but that gain did not reliably survive later RLVR. HF Daily Papers' note

score 4

Categories: Research