Megadose AI progress, ranked and analyzed.

Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders

· HF Daily Papers ·
The paper argues OPD mostly changes how a student model uses features it already has, rather than adding new ones from the teacher.

Using sparse crosscoders, the authors compare the student before and after on-policy distillation with the stronger teacher. They introduce a “swap readout” to measure how each student checkpoint changes its use of shared features. Across three OPD settings, more than 98% of frequently used student features stayed within 20% of their prior firing rates. The SFT warm-up appears to do much of the same feature reweighting in advance, including shifts tied to conversation format, reasoning style, and mathematical notation. HF Daily Papers' note

score 4

Categories: Research