Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
The paper argues that distilling teacher-relative logit shifts can preserve what each teacher learned better than copying endpoint policies.
The proposed Δ-MOPD subtracts each teacher’s base behavior, then re-anchors the shift at the student’s frozen initialization. In composed-teacher settings, it outperformed endpoint composition with three teachers by 4.11 Math points and 1.95 points across five benchmarks. With two composed teachers it matched endpoint accuracy, and under interleaved routing the two approaches were comparable. In phased routing, it improved mean performance and narrowed the reported order gap from 10.50 to 6.42 points. HF Daily Papers' note
The proposed Δ-MOPD subtracts each teacher’s base behavior, then re-anchors the shift at the student’s frozen initialization. In composed-teacher settings, it outperformed endpoint composition with three teachers by 4.11 Math points and 1.95 points across five benchmarks. With two composed teachers it matched endpoint accuracy, and under interleaved routing the two approaches were comparable. In phased routing, it improved mean performance and narrowed the reported order gap from 10.50 to 6.42 points. HF Daily Papers' note
score 4