Megadose Built for builders and researchers.

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

· HF Daily Papers ·
LEGO-RL reports SWE-bench Verified gains across three coding-agent harnesses while keeping rollout-training probability correlation above 0.99.

The paper says native agent harnesses create RL training problems through crashes, reward hacking, and train-inference mismatch. LEGO-RL addresses that with in-process LLM proxying, sandbox orchestration, and training observability tools. In tests, Qwen3.5-35B-A3B improved from 64.0% to 70.4% on OpenHands SDK, 62.4% to 68.2% on Claude Code, and 57.2% to 66.6% on OpenCode. HF Daily Papers' note

score 5

Categories: Research