Megadose AI progress, ranked and analyzed.

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training

· ArXiv · AI/CL/LG ·
The paper argues that agent benchmarks can look successful even when the simulated user is breaking its script.

UserProxyBench adds a fidelity check for LLM user simulators on tau-bench-style enterprise tasks. With the agent held fixed at GPT-5.5, changing only the user proxy moved mean task reward by 15.2 points across 375 tasks. The authors report that 24.4% of successful episodes still had a user-specification violation, most often premature disclosure. That failure reduced tool use while leaving reward intact, meaning the benchmark interaction changed without being penalized. ArXiv · AI/CL/LG's note

score 5

Categories: Research