Megadose AI progress, ranked and analyzed.

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

· HF Daily Papers ·
The review argues OPSD’s central failure is collapse, driven by token weighting, privileged context, and teacher dynamics.

Robert and Qader frame on-policy self-distillation as a cheaper variant of on-policy distillation, with the model teaching itself using information it will not have at test time. The paper says that asymmetry can create useful training signal, but also biases the model toward a narrower set of reasoning paths. Its scope is mathematical reasoning, and it reports no new experiments. The stated contribution is a shared vocabulary and a separation between settled findings and disputed ones. HF Daily Papers' note

score 4

Categories: Research