Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
The paper says agent benchmark scores can be badly inflated when evaluation protocols leave shortcuts open.
The authors define “protocol validity” and propose HackDetect, a post-hoc audit for finding exposures and judging whether an agent’s score reflects the intended capability. They audited 2,385 traces across 15 agent benchmarks. The paper reports exposure or reward-hacking evidence in 67.0% of Frontier Science traces and 66.7% of AutoLab tasks, with measured score inflation of 0.45 to 1.00 in paired comparisons. Source: ArXiv · AI/CL/LG's note.
The authors define “protocol validity” and propose HackDetect, a post-hoc audit for finding exposures and judging whether an agent’s score reflects the intended capability. They audited 2,385 traces across 15 agent benchmarks. The paper reports exposure or reward-hacking evidence in 67.0% of Frontier Science traces and 66.7% of AutoLab tasks, with measured score inflation of 0.45 to 1.00 in paired comparisons. Source: ArXiv · AI/CL/LG's note.
score 5