Rethinking the Evaluation of Harness Evolution for Agents
The paper argues that reported gains from harness evolution may be search effects, not better harness design.
The authors test automatic harness evolution against simpler task-level search baselines under matched feedback and inference budgets. On Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, evolved harnesses do not consistently beat test-time scaling methods. They also show limited transfer to held-out tasks, raising overfitting concerns when search and evaluation use the same benchmark. HF Daily Papers' note
The authors test automatic harness evolution against simpler task-level search baselines under matched feedback and inference budgets. On Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, evolved harnesses do not consistently beat test-time scaling methods. They also show limited transfer to held-out tasks, raising overfitting concerns when search and evaluation use the same benchmark. HF Daily Papers' note
score 4