Megadose Built for builders and researchers.

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

· ArXiv · AI/CL/LG ·
The paper says standard multi-teacher on-policy distillation leaves most oracle-ensemble gains unused because token-level training budget is badly allocated.

On a controlled SmolLM3-3B-Base benchmark with oracle routing, the authors report standard M-OPD recovers only 35.6% of the available headroom. They attribute the gap to sequence-length disparities, uneven convergence rates, and stale rewards from asynchronous policy updates, not gradient conflict. Their Open-MOPD recipe adds token-share balancing, dynamic budget allocation based on remaining gaps, and reward refresh, raising reported headroom recovery to 83.4%. The authors say they are open-sourcing the post-training recipe, trajectories, and evaluation suites. ArXiv · AI/CL/LG's note

score 5

Categories: Research