Megadose AI progress, ranked and analyzed.

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

· HF Daily Papers ·
The paper’s key move is letting the student decide when the teacher has outlived its usefulness.

RetireOPD trains a skill-conditioned teacher with environment rewards, then uses it to give token-level supervision while a skill-free student also trains with RL. The method drops distillation once student-teacher discrepancy stops improving and the student reaches a target share of the teacher’s success rate. In the reported Qwen2.5 runs, it beats the RL baseline on ALFWorld and WebShop by double-digit margins and surpasses the teacher in every setting. HF Daily Papers' note

score 5

Categories: Research