DarwinX: Evolving Agent Harnesses Through Natural Selection
DarwinX improves frozen LLM agents by evolving their harnesses, not their weights.
The paper treats prompts, tools, skills, and control flow as a population under selection, admitting only variants that add coverage without regressions. Its fitness signal comes from each benchmark’s verifier rather than gold solutions or manual picks. In the reported runs, one loop lifts average performance by about 17 points across four benchmarks, including Terminal-Bench, TerminalWorld, WebArena-Infinity, and transfer to SWE-bench Verified. The authors argue the gains reflect broader agent competence because the evolved harness survives changes in task, verifier, and base model. HF Daily Papers' note
The paper treats prompts, tools, skills, and control flow as a population under selection, admitting only variants that add coverage without regressions. Its fitness signal comes from each benchmark’s verifier rather than gold solutions or manual picks. In the reported runs, one loop lifts average performance by about 17 points across four benchmarks, including Terminal-Bench, TerminalWorld, WebArena-Infinity, and transfer to SWE-bench Verified. The authors argue the gains reflect broader agent competence because the evolved harness survives changes in task, verifier, and base model. HF Daily Papers' note
score 6