Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
Microsoft’s v1.0 framework trains agents through the same harness they use in deployment, with a 3,500-line control plane.
Agent Lightning puts an OpenAI-compatible proxy between the harness and the model, so existing agent code can keep its own loop while RL training records model calls. The release adds native Kubernetes rollout support and a collocated async training mode meant to reduce idle GPU time. In Microsoft’s coding-agent example, Qwen3.5-9B improved from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. Source: Microsoft Research AI's note.
Agent Lightning puts an OpenAI-compatible proxy between the harness and the model, so existing agent code can keep its own loop while RL training records model calls. The release adds native Kubernetes rollout support and a collocated async training mode meant to reduce idle GPU time. In Microsoft’s coding-agent example, Qwen3.5-9B improved from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. Source: Microsoft Research AI's note.
score 6