Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
The paper argues OPD mostly changes how a student model uses features it already has, rather than adding new ones from the teacher.
Using sparse crosscoders, the authors compare the student before and after on-policy distillation with the stronger teacher. They introduce a “swap readout” to measure how each student checkpoint changes its use of shared features. Across three OPD settings, more than 98% of frequently used student features stayed within 20% of their prior firing rates. The SFT warm-up appears to do much of the same feature reweighting in advance, including shifts tied to conversation format, reasoning style, and mathematical notation. HF Daily Papers' note
Using sparse crosscoders, the authors compare the student before and after on-policy distillation with the stronger teacher. They introduce a “swap readout” to measure how each student checkpoint changes its use of shared features. Across three OPD settings, more than 98% of frequently used student features stayed within 20% of their prior firing rates. The SFT warm-up appears to do much of the same feature reweighting in advance, including shifts tied to conversation format, reasoning style, and mathematical notation. HF Daily Papers' note
score 4