HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
The paper tests whether models can build the execution scaffolding that makes agents work, not just solve tasks inside one.
HarnessDev asks LLMs to create a runnable agent harness from a minimal seed, then revise it using downstream feedback. The benchmark judges the resulting harnesses on held-out task success and execution-token cost across six creator models, four domains, and five downstream benchmarks. The generated harnesses still lag mature human-built references on code, search, and research, but match or beat the selected references on writing and machine-learning experimentation. Evolution helps sometimes, but the gains are unstable and only partly carry over to held-out tasks or other runtime models. ArXiv · AI/CL/LG's note
HarnessDev asks LLMs to create a runnable agent harness from a minimal seed, then revise it using downstream feedback. The benchmark judges the resulting harnesses on held-out task success and execution-token cost across six creator models, four domains, and five downstream benchmarks. The generated harnesses still lag mature human-built references on code, search, and research, but match or beat the selected references on writing and machine-learning experimentation. Evolution helps sometimes, but the gains are unstable and only partly carry over to held-out tasks or other runtime models. ArXiv · AI/CL/LG's note
score 5