Megadose AI progress, ranked and analyzed.

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

· HF Daily Papers ·
The benchmark finds current agents still struggle to run full data-science jobs inside real computer environments.

DSAgentBench tests 275 tasks across wrangling, exploration, modeling, visualization, and validation. Its evaluator checks analytical correctness, visual outputs, and model performance, not just whether code ran. In the paper’s experiments, the strongest agent, Claude-4.6-Sonnet, reached 56.70% task success, while open-source agents stayed below 1%. HF Daily Papers' note

score 6

Categories: Research