Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
On-policy distillation appears to pass along reasoning behavior, not just answers.
The paper studies OPD by changing one generalization factor at a time, including distribution shifts, cross-domain transfer, and multi-teacher setups. It finds that same-origin teacher-student pairs generalize broadly, while cross-origin pairs tend to stay closer to the trained distribution. The authors also warn that multi-teacher OPD can create a capability seesaw, because each teacher’s influence is not confined to its routed domain. HF Daily Papers' note
The paper studies OPD by changing one generalization factor at a time, including distribution shifts, cross-domain transfer, and multi-teacher setups. It finds that same-origin teacher-student pairs generalize broadly, while cross-origin pairs tend to stay closer to the trained distribution. The authors also warn that multi-teacher OPD can create a capability seesaw, because each teacher’s influence is not confined to its routed domain. HF Daily Papers' note
score 4