Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
DCSD splits a teacher’s token-level credit into direction and magnitude, then uses both to correct self-distillation signals.
The paper argues that existing teacher-supervised step credit couples whether a token helped with how much it helped, leaving both exposed to teacher errors and preference variance. Its method uses belief-margin probing for credit direction and marginal information gain for credit magnitude. Across 11 benchmarks, DCSD beats GRPO, OPSD, RLSD, and RLCSD overall, with reported gains of 8.45 points on mathematical reasoning and 7.01 on multimodal reasoning over base models. HF Daily Papers' note
The paper argues that existing teacher-supervised step credit couples whether a token helped with how much it helped, leaving both exposed to teacher errors and preference variance. Its method uses belief-margin probing for credit direction and marginal information gain for credit magnitude. Across 11 benchmarks, DCSD beats GRPO, OPSD, RLSD, and RLCSD overall, with reported gains of 8.45 points on mathematical reasoning and 7.01 on multimodal reasoning over base models. HF Daily Papers' note
score 4