ClawGym II: Exploring Black-Box RL on Agent Harness
The paper reports black-box RL gains for agent harnesses without opening up the harness internals.
ClawGym II wraps task environments and harnesses in temporary sandboxes for concurrent rollouts, then captures model calls through a serving proxy. The authors rebuild multi-turn trajectories as prefix trees and adapt PPO and GRPO around that structure. With Qwen3-30A3B, they report Pass@1 gains of 9.98 points through OpenClaw and 14.81 points through Claude Code on ClawGym-Bench, stable over 200-400 optimization steps. They also report consistent gains on JobBench and OfficeQA. ArXiv · AI/CL/LG's note
ClawGym II wraps task environments and harnesses in temporary sandboxes for concurrent rollouts, then captures model calls through a serving proxy. The authors rebuild multi-turn trajectories as prefix trees and adapt PPO and GRPO around that structure. With Qwen3-30A3B, they report Pass@1 gains of 9.98 points through OpenClaw and 14.81 points through Claude Code on ClawGym-Bench, stable over 200-400 optimization steps. They also report consistent gains on JobBench and OfficeQA. ArXiv · AI/CL/LG's note
score 5