Megadose AI progress, ranked and analyzed.

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

· ArXiv · AI/CL/LG ·
The benchmark says hybrid computer-use agents still struggle to reproduce working behavior, not just screens.

RecreationWorld tests agents that must inspect a running reference app, rebuild it, and verify the result across Ubuntu, macOS, Windows, Android, and the web. The authors introduce RecreationBench, 250 held-out tasks with programmatic and visual checks validated against the reference. GPT-6 Astra leads with 58.1% overall, but clears every programmatic test on only 2.8% of tasks. The paper says agents match static interface structure better than interactions or computed outputs.

ArXiv · AI/CL/LG's note

score 6

Categories: Research