DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
DocOps tests whether agents can make complex document edits without losing consistency or damaging structure.
The paper introduces a deterministically verifiable benchmark built around real-world document operations, from atomic edits to more coupled workflows. Its evaluations of closed- and open-source models find that frontier agent setups still struggle on long-range, highly connected tasks. The authors identify three recurring failures: losing long-term state, checking semantics too shallowly, and destructively editing structural metadata. HF Daily Papers' note
The paper introduces a deterministically verifiable benchmark built around real-world document operations, from atomic edits to more coupled workflows. Its evaluations of closed- and open-source models find that frontier agent setups still struggle on long-range, highly connected tasks. The authors identify three recurring failures: losing long-term state, checking semantics too shallowly, and destructively editing structural metadata. HF Daily Papers' note
score 5