RoboQuest: Generalist Physical Agents that Search, Inspect and Test
The benchmark’s best-tested agent solved only 23% of episodes.
RoboQuest tests whether embodied agents can search, inspect, and try tools before committing to a task. The paper says failures were usually not simple execution failures: agents could perform many underlying actions when given the hidden information. The harder gap was exploration, with agents often deciding too soon, failing to manage disturbances, and struggling to learn by trial and error. ArXiv · AI/CL/LG's note
RoboQuest tests whether embodied agents can search, inspect, and try tools before committing to a task. The paper says failures were usually not simple execution failures: agents could perform many underlying actions when given the hidden information. The harder gap was exploration, with agents often deciding too soon, failing to manage disturbances, and struggling to learn by trial and error. ArXiv · AI/CL/LG's note
score 6