Megadose AI progress, ranked and analyzed.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

· ArXiv · AI/CL/LG ·
The paper turns “improving the agent wrapper” into a scored benchmark task for frontier models.

HarnessOpt-Bench gives an optimizer LLM a seed harness, graded feedback, and a fixed evaluation budget, then scores its final harness on held-out tests. The setup treats prompts, tools, memory, control flow, and orchestration code as the object being optimized. Across 111 runs with five frontier models and four downstream tasks, model choice separated results more than the coding harness around it. Native harnesses were not consistently better, and gains varied by task and seed regime. ArXiv · AI/CL/LG's note

score 5

Categories: Research