Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
The paper proposes a way to grade conversational-agent benchmarks before trusting their scores.
The authors use LLM judges to assess benchmark consistency, scenario complexity, and policy coverage without needing a reference benchmark. They report agreement with independent human annotations and test the method on LLM-generated benchmarks, deliberately degraded benchmarks, and manually curated ones. Across domains and judge models, the metrics separate stronger benchmark sets from weaker ones. ArXiv · AI/CL/LG's note
The authors use LLM judges to assess benchmark consistency, scenario complexity, and policy coverage without needing a reference benchmark. They report agreement with independent human annotations and test the method on LLM-generated benchmarks, deliberately degraded benchmarks, and manually curated ones. Across domains and judge models, the metrics separate stronger benchmark sets from weaker ones. ArXiv · AI/CL/LG's note
score 4