DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
The benchmark finds current agents still struggle to run full data-science jobs inside real computer environments.
DSAgentBench tests 275 tasks across wrangling, exploration, modeling, visualization, and validation. Its evaluator checks analytical correctness, visual outputs, and model performance, not just whether code ran. In the paper’s experiments, the strongest agent, Claude-4.6-Sonnet, reached 56.70% task success, while open-source agents stayed below 1%. HF Daily Papers' note
DSAgentBench tests 275 tasks across wrangling, exploration, modeling, visualization, and validation. Its evaluator checks analytical correctness, visual outputs, and model performance, not just whether code ran. In the paper’s experiments, the strongest agent, Claude-4.6-Sonnet, reached 56.70% task success, while open-source agents stayed below 1%. HF Daily Papers' note
score 6