Megadose Built for builders and researchers.

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

· ArXiv · AI/CL/LG ·
The paper traces how multi-teacher distillation signals turn into actual student-model updates, and finds several places where the training machinery changes their weight.

The authors study Qwen3-1.7B with four RL-trained domain teachers from the same initialization, plus SmolLM3-3B diagnostics. They report that response length, averaging rules, Adam’s first moment, and BF16 rounding all shape how teacher gradients appear in parameter changes. One result: raw gradient differences shrink after Adam updates, with teacher cosine similarity rising to 0.83 and averaging-rule similarity to 0.96. Their top-64 intersection KL gradient closely matches Qwen’s full-vocabulary gradient, but its task effect depends on the averaging setup. ArXiv · AI/CL/LG's note

score 4

Categories: Research