One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Claude Opus 5 led the benchmark at 66.50% pass@1, but dropped to 47.53% on repeated reliability.
Thinkingbox tests agents on 507 policy-conditioned business workflows where the outcome is checked against persistent backend state. The paper says agents must handle missing information, multi-turn interaction, domain rules, and dependent tools without causing extra effects. Many failures still ended cleanly and included valid state-changing actions, which the authors argue makes tool-call success a weak proxy for finishing the job. HF Daily Papers' note
Thinkingbox tests agents on 507 policy-conditioned business workflows where the outcome is checked against persistent backend state. The paper says agents must handle missing information, multi-turn interaction, domain rules, and dependent tools without causing extra effects. Many failures still ended cleanly and included valid state-changing actions, which the authors argue makes tool-call success a weak proxy for finishing the job. HF Daily Papers' note
score 5