Megadose Built for builders and researchers.

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

· ArXiv · AI/CL/LG ·
The paper argues terminal-agent training fails when rollout data is detached from the harness that produced it.

CoTrace routes trajectories by harness provenance, using verified rollouts for SFT and fresh online interactions for RL. In the reported Tmax split, Qwen3.5-9B rises from 78 to 88 solved tasks with SFT, and to 90 with the online RL variant. The authors say a smaller harness-matched corpus beats much larger pooled data at lower compute. Transfer results on Terminal-Bench 2.1 and SWE-bench Lite hinge on keeping the training and evaluation runtimes compatible. ArXiv · AI/CL/LG's note

score 5

Categories: Research