Megadose AI progress, ranked and analyzed.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

· HF Daily Papers ·
The paper’s key result is a Pass@1 jump from 50.00% to 68.03% by sampling eight candidate terminal actions and using a stronger verifier before execution.

Mid-Harness sits between the model and terminal harness, checking candidate commands before one is allowed to run. The authors report that extra samples do little under weak verification, but a GPT-5.6 Sol verifier can pick better actions from the same TMAX-9B generator. Pairwise verification works best when TMAX-9B is also used as the verifier, and distilling the stronger verifier into TMAX-9B improves results without changing the generator. The paper says action scaling beats simply generating more trajectories on TerminalBench-Lite at lower estimated token cost. HF Daily Papers' note

score 5

Categories: Research