Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
The paper claims a student model can absorb multiple specialist LLMs by learning from both their logits and hidden representations.
Latent-MOPD routes each domain to the same specialist for representation and token supervision, then gradually shifts the training signal from hidden states toward predictions. The authors say it beats token-only, representation-only, and uniform-averaging baselines across nine math, code, and logic benchmarks in the same-family setting. They also report gains over single-channel baselines with larger cross-family teachers. HF Daily Papers' note
Latent-MOPD routes each domain to the same specialist for representation and token supervision, then gradually shifts the training signal from hidden states toward predictions. The authors say it beats token-only, representation-only, and uniform-averaging baselines across nine math, code, and logic benchmarks in the same-family setting. They also report gains over single-channel baselines with larger cross-family teachers. HF Daily Papers' note
score 4