Megadose AI progress, ranked and analyzed.

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

· HF Daily Papers ·
The paper says sparse teacher supervision can match or beat full on-policy distillation when token choice accounts for gradient noise.

The authors study why helpful teacher feedback can still produce noisy updates when supervision is sampled at only some student-generated tokens. They introduce an information-efficiency ratio to estimate which tokens give more reliable gradient signal. On math and medical reasoning tasks, adding that criterion improved existing token selectors. In some 0.1% to 1% token-budget settings, the sparse method matched or exceeded full OPD. HF Daily Papers' note

score 4

Categories: Research