How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
The paper’s core finding is that research agents repeatedly fail because they do not reliably check their own work against evidence.
The authors introduce AutoResearchEval, a benchmark of 100 real-world frontier science tasks spanning the research lifecycle. They evaluate 8 harness-model combinations across 800 agent trajectories and annotate failures at the process and artifact level. The resulting ARFT taxonomy names 45 recurring failure patterns. The same patterns appear even in the strongest models tested, which the paper frames as a model-level metacognitive deficit rather than a scaffold-specific issue. HF Daily Papers' note
The authors introduce AutoResearchEval, a benchmark of 100 real-world frontier science tasks spanning the research lifecycle. They evaluate 8 harness-model combinations across 800 agent trajectories and annotate failures at the process and artifact level. The resulting ARFT taxonomy names 45 recurring failure patterns. The same patterns appear even in the strongest models tested, which the paper frames as a model-level metacognitive deficit rather than a scaffold-specific issue. HF Daily Papers' note
score 5