Megadose AI progress, ranked and analyzed.

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

· ArXiv · AI/CL/LG ·
A two-stage OPD-then-RL recipe beats mixing distillation and RLVR in one training step.

The paper says on-policy distillation first broadens the student model’s set of teacher-backed reasoning paths, then RLVR sharpens performance inside that support. Its experiments report consistent gains over pure OPD, pure RLVR, and joint signal-combination baselines on logic and math reasoning benchmarks. The authors also identify OPD validation score as the switch signal for moving into RL, and argue OPD is a stronger cold start than SFT. ArXiv · AI/CL/LG's note

score 5

Categories: Research