Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
The paper says current frontier agents still look more like uneven engineering optimizers than autonomous researchers.
The authors evaluate seven frontier models on 36 long-horizon AI R&D tasks, using metrics that track framing, execution, feedback control, and reuse of experience. Agents can build practical solutions, but results vary sharply across runs. Their best work mostly adapts or combines known methods, while real methodological novelty is rare. Experience from earlier attempts can improve later choices, but it can also steer agents wrong. ArXiv · AI/CL/LG's note
The authors evaluate seven frontier models on 36 long-horizon AI R&D tasks, using metrics that track framing, execution, feedback control, and reuse of experience. Agents can build practical solutions, but results vary sharply across runs. Their best work mostly adapts or combines known methods, while real methodological novelty is rare. Experience from earlier attempts can improve later choices, but it can also steer agents wrong. ArXiv · AI/CL/LG's note
score 6