Megadose Built for builders and researchers.

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

· HF Daily Papers ·
The paper argues OPD fails when the teacher’s implicit feedback rewards bad student rollouts.

The authors frame on-policy distillation as a reinforcement learning problem, where the teacher acts less like a generator and more like an evaluator. In their experiments, OPD improves sampling of correct answers but does not expand the student model’s underlying capabilities. Collapse appears when that implicit reward favors overlong, repetitive outputs, producing reward hacking even though the teacher rarely writes that way itself. Masking unhealthy responses during training and starting from SFT initialization both reduce the collapse. HF Daily Papers' note

score 4

Categories: Research