D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
The paper’s claim is that shifting distillation data toward slower-improving domains can cut wasted rollout compute.
D^3-MOPD watches the per-domain reverse-KL signals already produced during multi-teacher on-policy distillation and adjusts sampling ratios while training continues. The scheduler runs outside the main training loop, so the authors describe it as zero-overhead for the core process. In their Qwen3.6-35B-A3B experiment with four domain teachers, it closed 97% of the average student-to-teacher gap versus 63% for vanilla MOPD. The paper also reports roughly 3x fewer rollout steps to reach the same peak performance, and student wins over specialist teachers on three of seven benchmarks. HF Daily Papers' note
D^3-MOPD watches the per-domain reverse-KL signals already produced during multi-teacher on-policy distillation and adjusts sampling ratios while training continues. The scheduler runs outside the main training loop, so the authors describe it as zero-overhead for the core process. In their Qwen3.6-35B-A3B experiment with four domain teachers, it closed 97% of the average student-to-teacher gap versus 63% for vanilla MOPD. The paper also reports roughly 3x fewer rollout steps to reach the same peak performance, and student wins over specialist teachers on three of seven benchmarks. HF Daily Papers' note
score 4