Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Seven frontier models were tested on 36 long-horizon AI R&D tasks, with the study grading how they worked, not just what they scored.
The paper says current agents look more like engineering optimizers than autonomous researchers. They can frame and implement practical solutions, but results vary sharply across runs. Their best work mostly adapts or combines known techniques, with genuine methodological novelty described as rare. The authors also find that reused experience can help or mislead later decisions, and that evaluation harness design affects stability. HF Daily Papers' note
The paper says current agents look more like engineering optimizers than autonomous researchers. They can frame and implement practical solutions, but results vary sharply across runs. Their best work mostly adapts or combines known techniques, with genuine methodological novelty described as rare. The authors also find that reused experience can help or mislead later decisions, and that evaluation harness design affects stability. HF Daily Papers' note
score 6