Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
The paper argues OPD’s gains may come less from teacher guidance than from suppressing unlikely student tokens.
Ding and Zhang find teacher supervision in on-policy distillation is noisy, with more noise appearing as teacher scale increases. The student still reaches similar performance whether that noisy supervision is kept or removed. Their analysis says a fixed negative advantage can match teacher-provided advantages, pointing to a teacher-free mechanism. They propose OPSA, which uses entropy-adaptive negative advantages and reports large gains over Qwen3-1.7B and OPD on AIME24. HF Daily Papers' note
Ding and Zhang find teacher supervision in on-policy distillation is noisy, with more noise appearing as teacher scale increases. The student still reaches similar performance whether that noisy supervision is kept or removed. Their analysis says a fixed negative advantage can match teacher-provided advantages, pointing to a teacher-free mechanism. They propose OPSA, which uses entropy-adaptive negative advantages and reports large gains over Qwen3-1.7B and OPD on AIME24. HF Daily Papers' note
score 5