Megadose Built for builders and researchers.

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

· HF Daily Papers ·
The benchmark tests whether agents can act correctly inside a simulated enterprise warehouse, not just write SQL.

Argo-Bench builds a New York food-delivery business with 235 ERP-style tables and 7.5 billion rows, while withholding the simulator’s ground truth from the agent. Its 210 tasks require agents to reconstruct facts across the warehouse and then take actions such as banning fraud, allocating incentives, or issuing back pay. The authors say each task has an executable reference solution using only the warehouse. The best of 14 tested frontier and open-weight models reached a score of 95 or higher on 34.8% of tasks and averaged 59.5 points. HF Daily Papers' note

score 5

Categories: Research