Megadose Built for builders and researchers.

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

· ArXiv · AI/CL/LG ·
OPD can improve sampling of correct answers, but it can also reward long, repetitive failure modes.

The paper frames on-policy distillation as a reinforcement learning problem where the teacher acts like an implicit reward model. Its experiments find that OPD does not expand the student’s capabilities; it makes favored responses easier to sample. When that implicit preference stops tracking quality, the student can learn overlong, repetitive rollouts the teacher itself rarely produces. The authors report that masking unhealthy responses during training and starting from SFT initialization both reduce collapse. ArXiv · AI/CL/LG's note

score 5

Categories: Research