Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
HoH reports a 52.25% average gain over standalone coding-agent harnesses after three iterations.
The paper frames Harness-of-Harness as a layer that runs existing coding agents through repeated planning, coding, and testing loops. It says the system improves by keeping increments small, separating development tests from evaluation, reusing prior work, and preserving versioned histories. The authors test it with Codex/GPT-5.5, OpenCode/DeepSeek-V4-Pro, and Pi/MiniMax-M3 across three benchmarks. They also report a 70-plus-iteration run that built a playable first-person-shooter game with story, mechanics, visuals, and audio. ArXiv · AI/CL/LG's note
The paper frames Harness-of-Harness as a layer that runs existing coding agents through repeated planning, coding, and testing loops. It says the system improves by keeping increments small, separating development tests from evaluation, reusing prior work, and preserving versioned histories. The authors test it with Codex/GPT-5.5, OpenCode/DeepSeek-V4-Pro, and Pi/MiniMax-M3 across three benchmarks. They also report a 70-plus-iteration run that built a playable first-person-shooter game with story, mechanics, visuals, and audio. ArXiv · AI/CL/LG's note
score 5