Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Verified on-policy scaffolds, not privileged written solutions, are the lever the authors say keeps self-distillation useful at larger model sizes.
The paper finds scaffold correctness matters more for downstream reasoning accuracy than whether the teacher sees the right context. Standard OPSD keeps supervising many unverified student rollouts, creating an imitation gap when the teacher has information the student lacks. The authors introduce OASIS, which uses mostly label-verified on-policy trajectories and only needs final-answer labels. Across Qwen3 1.7B, 4B, and 8B on AIME and HMMT benchmarks, OASIS gains 3.2–3.8 points on average, while OPSD’s gains nearly vanish at 8B. HF Daily Papers' note
The paper finds scaffold correctness matters more for downstream reasoning accuracy than whether the teacher sees the right context. Standard OPSD keeps supervising many unverified student rollouts, creating an imitation gap when the teacher has information the student lacks. The authors introduce OASIS, which uses mostly label-verified on-policy trajectories and only needs final-answer labels. Across Qwen3 1.7B, 4B, and 8B on AIME and HMMT benchmarks, OASIS gains 3.2–3.8 points on average, while OPSD’s gains nearly vanish at 8B. HF Daily Papers' note
score 4