Megadose AI progress, ranked and analyzed.

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

· HF Daily Papers ·
PivotOPD trains agents both to avoid early pivotal errors and to recover after making them.

The paper says failed multi-turn rollouts often contain an early action that pushes the agent farther from completion. Those errors are frequently still recoverable if a teacher guides the next few turns. PivotOPD adds teacher-provided gold actions at the mistake point and recovery actions afterward. It reports the best average results among 13 baselines on ALFWorld, WebShop, and Search-based QA, plus a +3.2% resolve-rate gain on SWE-Bench Verified for a Nemotron-3.5 student. HF Daily Papers' note

score 4

Categories: Research