Megadose AI progress, ranked and analyzed.

Rethinking the Evaluation of Harness Evolution for Agents

· HF Daily Papers ·
The paper argues that reported gains from harness evolution may be search effects, not better harness design.

The authors test automatic harness evolution against simpler task-level search baselines under matched feedback and inference budgets. On Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, evolved harnesses do not consistently beat test-time scaling methods. They also show limited transfer to held-out tasks, raising overfitting concerns when search and evaluation use the same benchmark. HF Daily Papers' note

score 4

Categories: Research