MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation
The paper argues that smaller models need Long CoT data filtered for learnability, not just copied from stronger reasoners.
The authors analyze why direct Long CoT supervision can underperform shorter rationales in distillation. They report that Long CoT produces larger, more concentrated gradients, especially as student capacity grows. MI-Distillation builds a spectrum of instruct-reasoning data through model interpolation, then uses SeqLSS to select trajectories that are informative but still aligned with the student. Experiments on reasoning benchmarks show consistent gains over strong Long CoT distillation baselines. ArXiv · AI/CL/LG's note
The authors analyze why direct Long CoT supervision can underperform shorter rationales in distillation. They report that Long CoT produces larger, more concentrated gradients, especially as student capacity grows. MI-Distillation builds a spectrum of instruct-reasoning data through model interpolation, then uses SeqLSS to select trajectories that are informative but still aligned with the student. Experiments on reasoning benchmarks show consistent gains over strong Long CoT distillation baselines. ArXiv · AI/CL/LG's note
score 4