DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation
DualOPSD updates the privileged teacher during training instead of leaving it fixed.
The paper proposes an asymmetric alternating setup: the student learns from the privileged teacher, then the teacher shifts toward the updated student on the same trajectory. That makes later supervision track the learner without requiring another rollout. On Qwen3-8B in non-thinking mode, the authors report avg@12 gains over OPSD of 23.61 points on AIME 2024, 13.89 on AIME 2025, and 10.00 on HMMT 2025. They also report reduced truncation across 1.7B, 4B, and 8B models, with scale affecting the accuracy gain. ArXiv · AI/CL/LG's note
The paper proposes an asymmetric alternating setup: the student learns from the privileged teacher, then the teacher shifts toward the updated student on the same trajectory. That makes later supervision track the learner without requiring another rollout. On Qwen3-8B in non-thinking mode, the authors report avg@12 gains over OPSD of 23.61 points on AIME 2024, 13.89 on AIME 2025, and 10.00 on HMMT 2025. They also report reduced truncation across 1.7B, 4B, and 8B models, with scale affecting the accuracy gain. ArXiv · AI/CL/LG's note
score 6