StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
StarHarness reports 20-35 point benchmark gains by evolving the agent harness while leaving model weights unchanged.
The framework changes prompts, task framing, tools, skills, MCP-backed providers, subagent structure, and loop settings for a given enterprise environment. It builds its search pool by stratifying baseline failures, then separates visible search tasks from hidden selection tasks and held-out generalization tasks. The paper tests it on ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, with gains after 4-12 accepted harness changes per environment. The authors say the improvements persist on excluded tasks and transfer across GPT and Qwen model families. ArXiv · AI/CL/LG's note
The framework changes prompts, task framing, tools, skills, MCP-backed providers, subagent structure, and loop settings for a given enterprise environment. It builds its search pool by stratifying baseline failures, then separates visible search tasks from hidden selection tasks and held-out generalization tasks. The paper tests it on ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, with gains after 4-12 accepted harness changes per environment. The authors say the improvements persist on excluded tasks and transfer across GPT and Qwen model families. ArXiv · AI/CL/LG's note
score 5