τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
The benchmark asks coding agents to build the agent itself, then tests the shipped system against simulated users.
τ^τ-Bench gives a developer agent business records, client requirements, an inherited codebase, a production API, and serving-cost limits. Across 53 tasks in four domains, the strongest tested setup passed 23.9% of evaluation simulations, while an expert-authored reference reached 82.2%. The paper says failures clustered around shallow record use, poor client communication, and too little experimentation before shipping. HF Daily Papers' note
τ^τ-Bench gives a developer agent business records, client requirements, an inherited codebase, a production API, and serving-cost limits. Across 53 tasks in four domains, the strongest tested setup passed 23.9% of evaluation simulations, while an expert-authored reference reached 82.2%. The paper says failures clustered around shallow record use, poor client communication, and too little experimentation before shipping. HF Daily Papers' note
score 6