Megadose AI progress, ranked and analyzed.

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

· HF Daily Papers ·
A single query recovered most of full-data on-policy distillation’s gains in the authors’ tests.

The paper says one-shot OPD kept improving for hundreds of steps across task domains and model families. The authors measure “state coverage,” finding one query reached 71.5% of the states visited by full-data OPD, while 16 semantically diverse queries reached 98.9% and matched full-data training. They argue OPD is “data-overfed but algorithm-starved”: rollouts expose broad supervision quickly, but the student absorbs it slowly. The result also held in multi-teacher OPD, and even content-light or off-domain queries came close to the real-query baseline. HF Daily Papers' note

score 5

Categories: Research