OpenForgeRL: Train Harness-native Agents in Any Environment
OpenForgeRL lets researchers train agents inside the same inference harnesses they use at deployment.
The framework uses a proxy to capture harness model calls as RL training data, while Kubernetes runs each rollout in its own remote container. The paper validates it across claw/tool agents and GUI browser or computer-use agents. Reported results beat similar-size open baselines on nearly all tested benchmarks, with GUI agents matching or passing some larger models. The authors also find RL improves reliability behaviors like self-verification and tool coverage, while error recovery remains weak. HF Daily Papers' note
The framework uses a proxy to capture harness model calls as RL training data, while Kubernetes runs each rollout in its own remote container. The paper validates it across claw/tool agents and GUI browser or computer-use agents. Reported results beat similar-size open baselines on nearly all tested benchmarks, with GUI agents matching or passing some larger models. The authors also find RL improves reliability behaviors like self-verification and tool coverage, while error recovery remains weak. HF Daily Papers' note
score 6