RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
The benchmark says hybrid computer-use agents still struggle to reproduce working behavior, not just screens.
RecreationWorld tests agents that must inspect a running reference app, rebuild it, and verify the result across Ubuntu, macOS, Windows, Android, and the web. The authors introduce RecreationBench, 250 held-out tasks with programmatic and visual checks validated against the reference. GPT-6 Astra leads with 58.1% overall, but clears every programmatic test on only 2.8% of tasks. The paper says agents match static interface structure better than interactions or computed outputs.
ArXiv · AI/CL/LG's note
RecreationWorld tests agents that must inspect a running reference app, rebuild it, and verify the result across Ubuntu, macOS, Windows, Android, and the web. The authors introduce RecreationBench, 250 held-out tasks with programmatic and visual checks validated against the reference. GPT-6 Astra leads with 58.1% overall, but clears every programmatic test on only 2.8% of tasks. The paper says agents match static interface structure better than interactions or computed outputs.
ArXiv · AI/CL/LG's note
score 6