What Does Privileged Information Add to On-Policy Self-Distillation?
The paper finds that the extra “privileged” solution often adds less than the distillation setup itself.
The authors build AMPLE-Math, a 5,319-problem testbed with six reasoning views tied to the same answers. In Qwen3-1.7B, reference-free distillation explains much of the gain, with only modest added benefit from privileged references. SmolLM3-3B shows a clearer boost from complete traces at one checkpoint, but the effect depends on the student model. Longer thinking-style rollouts can reverse the gains even when the tasks, references, and evaluation stay fixed. HF Daily Papers' note
The authors build AMPLE-Math, a 5,319-problem testbed with six reasoning views tied to the same answers. In Qwen3-1.7B, reference-free distillation explains much of the gain, with only modest added benefit from privileged references. SmolLM3-3B shows a clearer boost from complete traces at one checkpoint, but the effect depends on the student model. Longer thinking-style rollouts can reverse the gains even when the tasks, references, and evaluation stay fixed. HF Daily Papers' note
score 4