AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
AutoDataBench tests whether agents can produce individual training tasks that a data pipeline would actually accept.
The paper says existing evaluations look at downstream training results, while real data work accepts or rejects samples one at a time. In AutoDataBench, an agent gets an original benchmark task and a target model’s attempt, then must write a new task meeting standards for validity, novelty, difficulty, and behavioral coverage. Across three executable agent-task benchmarks, evaluated agents scored below 20/100 at a 45-minute budget. More time helped the strongest agent, but did not materially lower the cost of each usable task. HF Daily Papers' note
The paper says existing evaluations look at downstream training results, while real data work accepts or rejects samples one at a time. In AutoDataBench, an agent gets an original benchmark task and a target model’s attempt, then must write a new task meeting standards for validity, novelty, difficulty, and behavioral coverage. Across three executable agent-task benchmarks, evaluated agents scored below 20/100 at a 45-minute budget. More time helped the strongest agent, but did not materially lower the cost of each usable task. HF Daily Papers' note
score 5