Pass the Baton: Trajectory-Relayed On-Policy Distillation
Relay-OPD lets the teacher step in only when a student’s reasoning path starts to fail.
The paper targets “prefix failure” in on-policy distillation, where a student’s early wrong turn makes later supervision less useful. Its method detects a teacher-student continuation mismatch and uses that as a label-free trigger for a short teacher handoff. The student then resumes from the relayed trajectory, keeping training close to its own policy while concentrating help near critical early positions. In math reasoning tests, the authors report average gains over standard OPD and FastOPD for the 1.7B student, with trajectory length cut by more than half. ArXiv · AI/CL/LG's note
The paper targets “prefix failure” in on-policy distillation, where a student’s early wrong turn makes later supervision less useful. Its method detects a teacher-student continuation mismatch and uses that as a label-free trigger for a short teacher handoff. The student then resumes from the relayed trajectory, keeping training close to its own policy while concentrating help near critical early positions. In math reasoning tests, the authors report average gains over standard OPD and FastOPD for the 1.7B student, with trajectory length cut by more than half. ArXiv · AI/CL/LG's note
score 5