Evo-Bench: Can Language Models Improve Agent Harness?
Evo-Bench tests whether models can improve the agent harness around them, not just solve fixed tasks.
The benchmark covers Search, Office, and General agent domains, with task construction meant to separate harness evolution from raw model strength. Across nine frontier and open-weight models, the paper reports gains up to 16.6 points, near human-engineered harness baselines. Autonomous evolution did best on Search and General tasks, but struggled on Office workflows requiring specific processing steps. HF Daily Papers' note
The benchmark covers Search, Office, and General agent domains, with task construction meant to separate harness evolution from raw model strength. Across nine frontier and open-weight models, the paper reports gains up to 16.6 points, near human-engineered harness baselines. Autonomous evolution did best on Search and General tasks, but struggled on Office workflows requiring specific processing steps. HF Daily Papers' note
score 5