HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
The benchmark tests whether LLMs can improve the agent harness around a model, not the model itself.
HarnessOpt-Bench gives an optimizer a seed harness, evaluation feedback, and a fixed budget, then scores the final harness on a hidden test partition. The paper evaluates five frontier LLMs across four downstream tasks and 111 scored runs. Its results say optimizer models separated more clearly than the coding harnesses they used, while native harnesses were not reliably better. HF Daily Papers' note
HarnessOpt-Bench gives an optimizer a seed harness, evaluation feedback, and a fixed budget, then scores the final harness on a hidden test partition. The paper evaluates five frontier LLMs across four downstream tasks and 111 scored runs. Its results say optimizer models separated more clearly than the coding harnesses they used, while native harnesses were not reliably better. HF Daily Papers' note
score 5