Megadose AI progress, ranked and analyzed.

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

· ArXiv · AI/CL/LG ·
Sparse teacher supervision can match full on-policy distillation when token selection accounts for gradient noise.

The paper studies why useful teacher feedback can still produce noisy updates when it is estimated from sampled next tokens. It proposes an information-efficiency ratio to choose tokens by both usefulness and gradient-estimation reliability. In mathematical and medical reasoning tasks, adding that measure improved existing selectors across multiple settings. Sparse runs using only 0.1% to 1% of tokens matched or beat full OPD without token selection. ArXiv · AI/CL/LG's note

score 5

Categories: Research