Megadose AI progress, ranked and analyzed.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

· ArXiv · AI/CL/LG ·
The paper tests whether models can build the execution scaffolding that makes agents work, not just solve tasks inside one.

HarnessDev asks LLMs to create a runnable agent harness from a minimal seed, then revise it using downstream feedback. The benchmark judges the resulting harnesses on held-out task success and execution-token cost across six creator models, four domains, and five downstream benchmarks. The generated harnesses still lag mature human-built references on code, search, and research, but match or beat the selected references on writing and machine-learning experimentation. Evolution helps sometimes, but the gains are unstable and only partly carry over to held-out tasks or other runtime models. ArXiv · AI/CL/LG's note

score 5

Categories: Research