Megadose Built for builders and researchers.

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

· HF Daily Papers ·
UndoBench finds agents can finish tasks while failing badly at recovering from lost acknowledgments.

The benchmark separates normal task competence from recovery behavior across paired trials with the same seeds. In the frozen lost-acknowledgment study, nominal competence was 83.54%, but conditional recovery success fell to 46.72%. Naive retry created duplicate external effects in 53.33% of trials. The paper says recovery depends on where the fault occurs, with verification and server-side idempotency helping most after commit but before acknowledgment. HF Daily Papers' note

score 5

Categories: Research