WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
The benchmark tests whether AI agents can tell which experiment changes actually moved the score.
WhatWorkedBench has agents inspect code, choose measurements, and predict outcomes across component configurations. Exhaustive CPU runs provide the reference effects across 36 tasks, 30 data sources, and eight workflow types. The paper reports that Gaussian-process fitting improved effect recovery over the same agent observations, including gains in the original Flash cohort and another cohort. Encoding code-equivalent configurations also improved recovery on six-option workflows. HF Daily Papers' note
WhatWorkedBench has agents inspect code, choose measurements, and predict outcomes across component configurations. Exhaustive CPU runs provide the reference effects across 36 tasks, 30 data sources, and eight workflow types. The paper reports that Gaussian-process fitting improved effect recovery over the same agent observations, including gains in the original Flash cohort and another cohort. Encoding code-equivalent configurations also improved recovery on six-option workflows. HF Daily Papers' note
score 5