Megadose Built for builders and researchers.

On-Policy Distillation with Negative-Policy Rollouts

· HF Daily Papers ·
A weaker “negative policy” is used during rollouts to show the student what to move away from, while the teacher still supplies the supervision.

The paper introduces NP-OPD, a variant of on-policy distillation for cases where the teacher and student have limited distributional overlap. Instead of changing the reward, it adds rollouts from a lower-capability policy so teacher supervision repeatedly covers tokens the negative policy prefers. The authors report gains over OPD across model scales, generation modes, reasoning domains, and OPD variants. Their analysis says the student suppresses tokens favored by the negative policy and shifts away from it. HF Daily Papers' note

score 4

Categories: Research