How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
A minimal coding-agent setup matched open-source MLE harnesses when the model backbone and time budget were held constant.
The paper argues that elaborate layers such as multi-agent orchestration and retrieval subagents did not improve results in this setting. Its ablations point to the frontier LLM itself as the main driver of benchmark performance. The authors conclude that hand-crafted harness complexity is giving poor returns on current autonomous ML engineering benchmarks. ArXiv · AI/CL/LG's note
The paper argues that elaborate layers such as multi-agent orchestration and retrieval subagents did not improve results in this setting. Its ablations point to the frontier LLM itself as the main driver of benchmark performance. The authors conclude that hand-crafted harness complexity is giving poor returns on current autonomous ML engineering benchmarks. ArXiv · AI/CL/LG's note
score 6