Megadose Built for builders and researchers.

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

· ArXiv · AI/CL/LG ·
Argo-Bench tests whether data agents can act correctly inside a simulated enterprise warehouse, not just write SQL.

The benchmark contains 210 analytics and data-science tasks over a simulated NYC food-delivery business with 235 ERP-style tables and 7.5 billion rows. Agents must reconstruct hidden ground truth from the warehouse, then take actions such as banning fraud accounts, allocating courier incentives, or issuing back pay. The grader scores those actions by their simulated consequences. The strongest model tested scored 95 or higher on 34.8% of tasks and averaged 59.5 points. ArXiv · AI/CL/LG's note

score 6

Categories: Research