AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents
AutoSciBench uses agent failures to rewrite scientific benchmarks until shortcuts close.
The framework defines tasks by domain, data type, reasoning demand, and a recipe for building and verifying answers. Solver trajectories and judge feedback are fed back into revisions, pushing tasks toward raw-data checks, intermediate interpretation, and evidence integration. In tests across computational biology, materials science, and clinical imaging, its generated benchmarks lowered solver accuracy by 22.4 and 25.5 points in the first two domains versus human-curated sets. The authors say the generated tasks also received higher average quality ratings across all three domains. HF Daily Papers' note
The framework defines tasks by domain, data type, reasoning demand, and a recipe for building and verifying answers. Solver trajectories and judge feedback are fed back into revisions, pushing tasks toward raw-data checks, intermediate interpretation, and evidence integration. In tests across computational biology, materials science, and clinical imaging, its generated benchmarks lowered solver accuracy by 22.4 and 25.5 points in the first two domains versus human-curated sets. The authors say the generated tasks also received higher average quality ratings across all three domains. HF Daily Papers' note
score 5