Megadose AI progress, ranked and analyzed.

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

· HF Daily Papers ·
A 122B MoE terminal agent trained in a real shell reports a 64.0% resolve rate on Terminal-Bench 2.1.

The paper introduces T1, a reinforcement-trained agent that can run 300-plus tool-call turns in a cloud sandbox and is rewarded by each task’s verifier. Its training recipe centers on stabilizing long rollouts, including TITO token alignment and replaying MoE routing choices during training. The authors say the pipeline lifts the base model from 43.8% to 64.0% on Terminal-Bench 2.1. On Long-Horizon Terminal Bench, they report 27.9%, ahead of GPT-5.4 and GLM-5.1. HF Daily Papers' note

score 6

Categories: Model Releases, Research