StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
StateM’s claim is that better agent runtime control, not new weights, pushed Terminal-Bench 2.1 accuracy to 95.3%.
The paper describes StateM as a harness that keeps durable state, phase-local context, checked transitions, recoverable runbooks, and versioned procedures around an agent. On Terminal-Bench 2.1, it reports GPT-5.6 Sol xhigh reaching 95.3% raw accuracy across 445 trials, with every one of 89 tasks solved at least once. The same runtime is also reported to cut final-score API usage to about $15, compared with $574.68 for the GPT reference. HF Daily Papers' note
The paper describes StateM as a harness that keeps durable state, phase-local context, checked transitions, recoverable runbooks, and versioned procedures around an agent. On Terminal-Bench 2.1, it reports GPT-5.6 Sol xhigh reaching 95.3% raw accuracy across 445 trials, with every one of 89 tasks solved at least once. The same runtime is also reported to cut final-score API usage to about $15, compared with $574.68 for the GPT reference. HF Daily Papers' note
score 6