Megadose Built for builders and researchers.

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

· HF Daily Papers ·
The paper finds that implementation details in multi-teacher distillation can materially change what the student learns.

The authors study Qwen3-1.7B distilled from four RL-trained domain teachers that share its initialization. They report that loss averaging can quietly reweight teacher signals, with token averaging favoring longer responses. Adam’s first moment makes teacher update directions look more similar than their raw gradients, while BF16 rounding hides many small FP32 weight changes. A top-64 intersection KL gradient closely matches Qwen’s full-vocabulary gradient, but its task effect depends on the averaging rule. HF Daily Papers' note

score 4

Categories: Research