ClawGym II: Exploring Black-Box RL on Agent Harness
The paper reports black-box RL gains for agent harnesses without needing to open up the harness internals.
The authors build a sandboxed rollout system and a serving proxy that captures model calls at the model boundary. They reconstruct multi-turn trajectories as prefix trees, then adapt PPO and GRPO to train over that structure. With Qwen3-30A3B, the framework raises Pass@1 on ClawGym-Bench by 9.98 points through OpenClaw and 14.81 points through Claude Code, staying stable across 200-400 optimization steps. They also report gains on JobBench and OfficeQA. HF Daily Papers' note
The authors build a sandboxed rollout system and a serving proxy that captures model calls at the model boundary. They reconstruct multi-turn trajectories as prefix trees, then adapt PPO and GRPO to train over that structure. With Qwen3-30A3B, the framework raises Pass@1 on ClawGym-Bench by 9.98 points through OpenClaw and 14.81 points through Claude Code, staying stable across 200-400 optimization steps. They also report gains on JobBench and OfficeQA. HF Daily Papers' note
score 5