Rethinking On-Policy Distillation of Large Language Models II: One Training Example
A single-query OPD run recovered most of the gains from full-data training.
The paper tests on-policy distillation at the extreme low-data limit and finds that one training query can keep improving for hundreds of steps. The authors say one query reaches 71.5% of the state coverage seen in full-data OPD, while 16 semantically distinct queries reach 98.9% and match full-data training. Their read is that OPD is “data-overfed but algorithm-starved”: rollouts expose broad supervision quickly, but the student absorbs it slowly. ArXiv · AI/CL/LG's note
The paper tests on-policy distillation at the extreme low-data limit and finds that one training query can keep improving for hundreds of steps. The authors say one query reaches 71.5% of the state coverage seen in full-data OPD, while 16 semantically distinct queries reach 98.9% and match full-data training. Their read is that OPD is “data-overfed but algorithm-starved”: rollouts expose broad supervision quickly, but the student absorbs it slowly. ArXiv · AI/CL/LG's note
score 5