Megadose AI progress, ranked and analyzed.

AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

· HF Daily Papers ·
AutoDataBench tests whether agents can produce individual training tasks that a data pipeline would actually accept.

The paper says existing evaluations look at downstream training results, while real data work accepts or rejects samples one at a time. In AutoDataBench, an agent gets an original benchmark task and a target model’s attempt, then must write a new task meeting standards for validity, novelty, difficulty, and behavioral coverage. Across three executable agent-task benchmarks, evaluated agents scored below 20/100 at a 45-minute budget. More time helped the strongest agent, but did not materially lower the cost of each usable task. HF Daily Papers' note

score 5

Categories: Research