Megadose AI progress, ranked daily.

On-policy Distillation with Verifiable Reward

· HF Daily Papers ·
OPDVR ties token-level distillation to verifiable task success without adding new tuning knobs.

The paper frames RLVR as too sparse and OPD as too dependent on the teacher’s limits. Its method gates the distillation reward so correct trajectories get non-negative rewards and incorrect ones get non-positive rewards. That makes sampled-token OPD usable as an RLVR-style objective with policy-gradient methods such as GRPO. The authors report consistent gains over standard OPD on six reasoning benchmarks. HF Daily Papers' note

score 5

Categories: Research