UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
The paper argues that agent benchmarks can look successful even when the simulated user is breaking its script.
UserProxyBench adds a fidelity check for LLM user simulators on tau-bench-style enterprise tasks. With the agent held fixed at GPT-5.5, changing only the user proxy moved mean task reward by 15.2 points across 375 tasks. The authors report that 24.4% of successful episodes still had a user-specification violation, most often premature disclosure. That failure reduced tool use while leaving reward intact, meaning the benchmark interaction changed without being penalized. ArXiv · AI/CL/LG's note
UserProxyBench adds a fidelity check for LLM user simulators on tau-bench-style enterprise tasks. With the agent held fixed at GPT-5.5, changing only the user proxy moved mean task reward by 15.2 points across 375 tasks. The authors report that 24.4% of successful episodes still had a user-specification violation, most often premature disclosure. That failure reduced tool use while leaving reward intact, meaning the benchmark interaction changed without being penalized. ArXiv · AI/CL/LG's note
score 5