When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
The paper targets distillation signals that look useful to training but are weakly tied to the input.
The authors argue that token-level teacher feedback in on-policy distillation can be driven by language priors, formatting habits, or stock reasoning patterns. Their SA-OPD method estimates whether a token’s supervision is input-grounded, then filters tokens with both low groundedness and extreme distillation divergence. The paper reports gains over vanilla OPD and other selective OPD methods across LLM and VLM experiments. HF Daily Papers' note
The authors argue that token-level teacher feedback in on-policy distillation can be driven by language priors, formatting habits, or stock reasoning patterns. Their SA-OPD method estimates whether a token’s supervision is input-grounded, then filters tokens with both low groundedness and extreme distillation divergence. The paper reports gains over vanilla OPD and other selective OPD methods across LLM and VLM experiments. HF Daily Papers' note
score 5