Megadose Built for builders and researchers.

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

· HF Daily Papers ·
The paper claims a student model can absorb multiple specialist LLMs by learning from both their logits and hidden representations.

Latent-MOPD routes each domain to the same specialist for representation and token supervision, then gradually shifts the training signal from hidden states toward predictions. The authors say it beats token-only, representation-only, and uniform-averaging baselines across nine math, code, and logic benchmarks in the same-family setting. They also report gains over single-channel baselines with larger cross-family teachers. HF Daily Papers' note

score 4

Categories: Research