RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
RobotWorld tests whether multimodal agents can turn instructions and observations into physical robot execution, and finds the handoff is still unreliable.
The benchmark covers 84 simulated tasks across manipulation, locomotion, driving, aerial control, and related robot interfaces. The paper says current agents can build useful perception and control workflows, including segmentation, calibration, spatial estimation, and dynamics-based computation. Those pieces often fail to compose into finished behavior: agents lose object state, repeat ineffective actions, recover late, or mark incomplete tasks as done. The authors also report model-dependent differences, with Astra stronger on spatial and constrained-contact goals and Opus 5.5 stronger on balance and timed interaction. ArXiv · AI/CL/LG's note
The benchmark covers 84 simulated tasks across manipulation, locomotion, driving, aerial control, and related robot interfaces. The paper says current agents can build useful perception and control workflows, including segmentation, calibration, spatial estimation, and dynamics-based computation. Those pieces often fail to compose into finished behavior: agents lose object state, repeat ineffective actions, recover late, or mark incomplete tasks as done. The authors also report model-dependent differences, with Astra stronger on spatial and constrained-contact goals and Opus 5.5 stronger on balance and timed interaction. ArXiv · AI/CL/LG's note
score 6