Megadose AI progress, ranked and analyzed.

Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

· HF Daily Papers ·
Direct-OPD can waste supervision on states where the teacher barely changed.

The paper argues that token-level log-ratio rewards can look unchanged even when the teacher and reference assign almost no mass to the sampled tokens. Its proposed fix, S²D-OPD, keeps Direct-OPD supervision only for the top 10% of response states ranked by teacher-reference Jensen-Shannon divergence. In tests across two teacher pairs and four student models from 1.7B to 8B parameters, it beats dense Direct-OPD on AIME and HMMT in seven of eight settings and matches it in the other. The method adds no extra forward passes. HF Daily Papers' note

score 4

Categories: Research