Scaling Properties of Same-Family On-Policy Distillation
A smaller RL-trained teacher can push a larger same-family student past the teacher’s own held-out score.
The paper studies on-policy distillation across weak-to-strong, same-base, and strong-to-weak teacher-student setups. It reports an early “useful-transfer” regime where held-out accuracy rises roughly linearly with the square root of reverse KL divergence from the student initialization. The authors fit power laws for peak score and transfer slope, finding teacher scale helps only up to about the student’s scale. At the same gold score, smaller teachers transferred better, so score alone did not determine supervision value. HF Daily Papers' note
The paper studies on-policy distillation across weak-to-strong, same-base, and strong-to-weak teacher-student setups. It reports an early “useful-transfer” regime where held-out accuracy rises roughly linearly with the square root of reverse KL divergence from the student initialization. The authors fit power laws for peak score and transfer slope, finding teacher scale helps only up to about the student’s scale. At the same gold score, smaller teachers transferred better, so score alone did not determine supervision value. HF Daily Papers' note
score 5