RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
RobotWorld tests whether digital agents can turn multimodal reasoning into physical robot execution, and finds the transfer is still uneven.
The benchmark covers 84 simulated tasks across manipulation, locomotion, driving, aerial control, and related robot interfaces. The paper says current agents can build complex perception and control workflows, including segmentation, calibration, spatial estimation, and dynamics-based computation. But those pieces often fail to compose into completed tasks: agents lose object state, repeat ineffective actions, recover too late, or stop before the job is done. Model differences show up by task type, with Astra stronger on spatial and constrained-contact goals and Opus 5.5 stronger on balance and timed-interaction goals. HF Daily Papers' note
The benchmark covers 84 simulated tasks across manipulation, locomotion, driving, aerial control, and related robot interfaces. The paper says current agents can build complex perception and control workflows, including segmentation, calibration, spatial estimation, and dynamics-based computation. But those pieces often fail to compose into completed tasks: agents lose object state, repeat ineffective actions, recover too late, or stop before the job is done. Model differences show up by task type, with Astra stronger on spatial and constrained-contact goals and Opus 5.5 stronger on balance and timed-interaction goals. HF Daily Papers' note
score 5