PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
PivotOPD trains agents for the early wrong turns that usually break multi-step tasks, and for the recovery path after them.
The paper says failed multi-turn rollouts often contain an early “pivotal mistake” that pushes the agent farther from completion. PivotOPD has a teacher supply the correct action at that turn, then recovery actions for the next few turns. The method beats 13 baselines across ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students, and also improves a Nemotron-3.5 student on SWE-Bench Verified. ArXiv · AI/CL/LG's note
The paper says failed multi-turn rollouts often contain an early “pivotal mistake” that pushes the agent farther from completion. PivotOPD has a teacher supply the correct action at that turn, then recovery actions for the next few turns. The method beats 13 baselines across ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students, and also improves a Nemotron-3.5 student on SWE-Bench Verified. ArXiv · AI/CL/LG's note
score 5