Megadose AI progress, ranked and analyzed.

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

· ArXiv · AI/CL/LG ·
Token agreement can look perfect even when the student model is looping through bad answers.

The paper says on-policy distillation can reward “degenerate agreement,” where a student matches a teacher token-by-token while producing globally flawed responses. It splits the problem into excess tokens the teacher would not choose and deficit tokens the student fails to sample. The proposed TIDE method bounds the excess corrections and injects teacher top-K probability mass for deficits. On math reasoning benchmarks with Qwen3 teacher-student pairs, TIDE beat standard OPD and other baselines, with larger gains under strong mismatch.

ArXiv · AI/CL/LG's note

score 5

Categories: Research