Megadose AI progress, ranked and analyzed.

Environment Evolution for Terminal Agents

· ArXiv · AI/CL/LG ·
The paper reports off-policy “environment evolution” that made terminal-agent tasks harder and improved Qwen terminal-benchmark scores.

The authors argue that scratch-built environments stop giving strong models enough useful training signal. Their method raises difficulty generation by generation, using a loop-engineered multi-agent harness rather than relying on on-policy rollout failures alone. Rollout tests with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol found the evolved environments were consistently harder. In RL training, Qwen3.6-27B and Qwen3.6-35B-A3B improved by 14.4 and 18.0 percentage points on Terminal-Bench 2.1. ArXiv · AI/CL/LG's note

score 5

Categories: Research