On-Policy Distillation Teaches New Skills but Not New Knowledge
Reverse-KL OPD improved reasoning composition, but barely moved factual memory.
The paper tests on-policy distillation in a controlled setup separating facts from multi-step skill. Across four models, reverse-KL OPD transferred compositional reasoning to unseen structures while adding little new factual knowledge. Switching to forward KL restored more factual transfer, while student rollouts mainly helped execution of multi-step reasoning. The authors report the same asymmetry in factual QA and competition math experiments. HF Daily Papers' note
The paper tests on-policy distillation in a controlled setup separating facts from multi-step skill. Across four models, reverse-KL OPD transferred compositional reasoning to unseen structures while adding little new factual knowledge. Switching to forward KL restored more factual transfer, while student rollouts mainly helped execution of multi-step reasoning. The authors report the same asymmetry in factual QA and competition math experiments. HF Daily Papers' note
score 4