Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
A two-stage OPD-then-RL recipe beats mixing distillation and RLVR in one training step.
The paper says on-policy distillation first broadens the student model’s set of teacher-backed reasoning paths, then RLVR sharpens performance inside that support. Its experiments report consistent gains over pure OPD, pure RLVR, and joint signal-combination baselines on logic and math reasoning benchmarks. The authors also identify OPD validation score as the switch signal for moving into RL, and argue OPD is a stronger cold start than SFT. ArXiv · AI/CL/LG's note
The paper says on-policy distillation first broadens the student model’s set of teacher-backed reasoning paths, then RLVR sharpens performance inside that support. Its experiments report consistent gains over pure OPD, pure RLVR, and joint signal-combination baselines on logic and math reasoning benchmarks. The authors also identify OPD validation score as the switch signal for moving into RL, and argue OPD is a stronger cold start than SFT. ArXiv · AI/CL/LG's note
score 5