Megadose AI progress, ranked and analyzed.

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

· ArXiv · AI/CL/LG ·
The paper argues OPD’s gains may come less from the teacher and more from penalizing unlikely student tokens.

The authors report that teacher supervision in on-policy distillation is noisy, with more noise as teacher scale rises. Student performance stayed similar whether that noisy supervision was kept or removed. Their analysis finds a fixed negative advantage can match teacher-provided advantages, motivating a teacher-free method, OPSA. OPSA improves Qwen3-1.7B on AIME24 and beats OPD there in Avg@32. ArXiv · AI/CL/LG's note

score 5

Categories: Research