Agent Lightning v1.0: Towards Harnessed Agentic RL
Agent Lightning v1.0 reports a 14.6-point SWE-bench Verified gain for Qwen3.5-9B using 6K training examples.
The paper frames “harnessed agentic RL” as training where the deploy-time agent harness controls tools, context, and interaction flow while the trainer sees LLM request-response sequences. It says that setup creates practical issues around retokenization, sample merging, advantage calculation, loss normalization, and scheduling. Agent Lightning v1.0 is presented as a roughly 3,500-line framework for testing those problems across instruction-following, search, and coding agents. The authors say they are releasing the workflow and training scripts for reproducible coding-agent RL. HF Daily Papers' note
The paper frames “harnessed agentic RL” as training where the deploy-time agent harness controls tools, context, and interaction flow while the trainer sees LLM request-response sequences. It says that setup creates practical issues around retokenization, sample merging, advantage calculation, loss normalization, and scheduling. Agent Lightning v1.0 is presented as a roughly 3,500-line framework for testing those problems across instruction-following, search, and coding agents. The authors say they are releasing the workflow and training scripts for reproducible coding-agent RL. HF Daily Papers' note
score 5