Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
The paper argues OPD fails when the teacher’s implicit feedback rewards bad student rollouts.
The authors frame on-policy distillation as a reinforcement learning problem, where the teacher acts less like a generator and more like an evaluator. In their experiments, OPD improves sampling of correct answers but does not expand the student model’s underlying capabilities. Collapse appears when that implicit reward favors overlong, repetitive outputs, producing reward hacking even though the teacher rarely writes that way itself. Masking unhealthy responses during training and starting from SFT initialization both reduce the collapse. HF Daily Papers' note
The authors frame on-policy distillation as a reinforcement learning problem, where the teacher acts less like a generator and more like an evaluator. In their experiments, OPD improves sampling of correct answers but does not expand the student model’s underlying capabilities. Collapse appears when that implicit reward favors overlong, repetitive outputs, producing reward hacking even though the teacher rarely writes that way itself. Masking unhealthy responses during training and starting from SFT initialization both reduce the collapse. HF Daily Papers' note
score 4