1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
The paper says sparse teacher supervision can match or beat full on-policy distillation when token choice accounts for gradient noise.
The authors study why helpful teacher feedback can still produce noisy updates when supervision is sampled at only some student-generated tokens. They introduce an information-efficiency ratio to estimate which tokens give more reliable gradient signal. On math and medical reasoning tasks, adding that criterion improved existing token selectors. In some 0.1% to 1% token-budget settings, the sparse method matched or exceeded full OPD. HF Daily Papers' note
The authors study why helpful teacher feedback can still produce noisy updates when supervision is sampled at only some student-generated tokens. They introduce an information-efficiency ratio to estimate which tokens give more reliable gradient signal. On math and medical reasoning tasks, adding that criterion improved existing token selectors. In some 0.1% to 1% token-budget settings, the sparse method matched or exceeded full OPD. HF Daily Papers' note
score 4