Megadose Built for builders and researchers.

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

· HF Daily Papers ·
R2-OPD suppresses teacher rewards when they disagree with an independent estimate of reasoning progress.

The paper argues that on-policy distillation can punish useful reasoning steps when they diverge from the teacher’s wording or path. Its method ranks reasoning spans two ways: by teacher-derived reward and by estimated progress reward. When those rankings conflict, the distillation reward is selectively reduced instead of applied uniformly. The authors report consistent gains over standard OPD, especially on reasoning performance. HF Daily Papers' note

score 5

Categories: Research