PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
PivotOPD trains agents both to avoid early pivotal errors and to recover after making them.
The paper says failed multi-turn rollouts often contain an early action that pushes the agent farther from completion. Those errors are frequently still recoverable if a teacher guides the next few turns. PivotOPD adds teacher-provided gold actions at the mistake point and recovery actions afterward. It reports the best average results among 13 baselines on ALFWorld, WebShop, and Search-based QA, plus a +3.2% resolve-rate gain on SWE-Bench Verified for a Nemotron-3.5 student. HF Daily Papers' note
The paper says failed multi-turn rollouts often contain an early action that pushes the agent farther from completion. Those errors are frequently still recoverable if a teacher guides the next few turns. PivotOPD adds teacher-provided gold actions at the mistake point and recovery actions afterward. It reports the best average results among 13 baselines on ALFWorld, WebShop, and Search-based QA, plus a +3.2% resolve-rate gain on SWE-Bench Verified for a Nemotron-3.5 student. HF Daily Papers' note
score 4