Automated Discovery Has No Universally Superior Harness
The paper argues harness choice should be treated as a tunable variable, not a default recipe.
The authors break OpenEvolve-style search and TTT-Discover into components and test 30 budget-matched harnesses across 12 model-problem pairs. Using more than 3.1 million LLM rollouts, they find no fixed harness is reliably best, while OpenEvolve variants often trail simpler alternatives. Early run progress predicts final performance, so they test adaptive allocation that starts several harnesses, prunes weak runs, and shifts compute to stronger ones. Source: ArXiv · AI/CL/LG's note.
The authors break OpenEvolve-style search and TTT-Discover into components and test 30 budget-matched harnesses across 12 model-problem pairs. Using more than 3.1 million LLM rollouts, they find no fixed harness is reliably best, while OpenEvolve variants often trail simpler alternatives. Early run progress predicts final performance, so they test adaptive allocation that starts several harnesses, prunes weak runs, and shifts compute to stronger ones. Source: ArXiv · AI/CL/LG's note.
score 6