MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
The paper’s core move is routing distillation supervision teacher-by-teacher at each token, without prompt domain labels.
MOPD-Router replaces hard prompt-level teacher assignment with token-level selection across the full teacher pool. Its ExpertAlign metric scores whether a teacher’s current correction reflects that teacher’s post-training specialization. In the reported experiments, ExpertAlign was strongest across four distillation settings, including unlabeled and domain-labeled mixtures. On unlabeled data it beat Mean aggregation by 5.88 points, and on labeled data it beat standard MOPD by 3.95 points without using the labels. HF Daily Papers' note
MOPD-Router replaces hard prompt-level teacher assignment with token-level selection across the full teacher pool. Its ExpertAlign metric scores whether a teacher’s current correction reflects that teacher’s post-training specialization. In the reported experiments, ExpertAlign was strongest across four distillation settings, including unlabeled and domain-labeled mixtures. On unlabeled data it beat Mean aggregation by 5.88 points, and on labeled data it beat standard MOPD by 3.95 points without using the labels. HF Daily Papers' note
score 4