Environment Evolution for Terminal Agents
The paper reports off-policy “environment evolution” that made terminal-agent tasks harder and improved Qwen terminal-benchmark scores.
The authors argue that scratch-built environments stop giving strong models enough useful training signal. Their method raises difficulty generation by generation, using a loop-engineered multi-agent harness rather than relying on on-policy rollout failures alone. Rollout tests with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol found the evolved environments were consistently harder. In RL training, Qwen3.6-27B and Qwen3.6-35B-A3B improved by 14.4 and 18.0 percentage points on Terminal-Bench 2.1. ArXiv · AI/CL/LG's note
The authors argue that scratch-built environments stop giving strong models enough useful training signal. Their method raises difficulty generation by generation, using a loop-engineered multi-agent harness rather than relying on on-policy rollout failures alone. Rollout tests with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol found the evolved environments were consistently harder. In RL training, Qwen3.6-27B and Qwen3.6-35B-A3B improved by 14.4 and 18.0 percentage points on Terminal-Bench 2.1. ArXiv · AI/CL/LG's note
score 5