Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD
Direct-OPD can waste supervision on states where the teacher barely changed.
The paper argues that token-level log-ratio rewards can look unchanged even when the teacher and reference assign almost no mass to the sampled tokens. Its proposed fix, S²D-OPD, keeps Direct-OPD supervision only for the top 10% of response states ranked by teacher-reference Jensen-Shannon divergence. In tests across two teacher pairs and four student models from 1.7B to 8B parameters, it beats dense Direct-OPD on AIME and HMMT in seven of eight settings and matches it in the other. The method adds no extra forward passes. HF Daily Papers' note
The paper argues that token-level log-ratio rewards can look unchanged even when the teacher and reference assign almost no mass to the sampled tokens. Its proposed fix, S²D-OPD, keeps Direct-OPD supervision only for the top 10% of response states ranked by teacher-reference Jensen-Shannon divergence. In tests across two teacher pairs and four student models from 1.7B to 8B parameters, it beats dense Direct-OPD on AIME and HMMT in seven of eight settings and matches it in the other. The method adds no extra forward passes. HF Daily Papers' note
score 4