HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
The paper turns “improving the agent wrapper” into a scored benchmark task for frontier models.
HarnessOpt-Bench gives an optimizer LLM a seed harness, graded feedback, and a fixed evaluation budget, then scores its final harness on held-out tests. The setup treats prompts, tools, memory, control flow, and orchestration code as the object being optimized. Across 111 runs with five frontier models and four downstream tasks, model choice separated results more than the coding harness around it. Native harnesses were not consistently better, and gains varied by task and seed regime. ArXiv · AI/CL/LG's note
HarnessOpt-Bench gives an optimizer LLM a seed harness, graded feedback, and a fixed evaluation budget, then scores its final harness on held-out tests. The setup treats prompts, tools, memory, control flow, and orchestration code as the object being optimized. Across 111 runs with five frontier models and four downstream tasks, model choice separated results more than the coding harness around it. Native harnesses were not consistently better, and gains varied by task and seed regime. ArXiv · AI/CL/LG's note
score 5