VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
The benchmark finds strong spatial recognition still breaks down when models have to act.
VA-Bench tests MLLMs on an observe-reason-act-revise loop using RGB demonstrations, active camera choices, Cartesian commands, and execution feedback. The best model reached 100.0% target localization and 78.9% spatial-relations performance in the annotated run, but averaged only 53.93% task success across three runs. Active camera control helped sharply in one matched comparison, lifting success from 27.86% to 57.50%. Held-out geometry cut success by more than 30 points, and no model finished a strict long-horizon episode. HF Daily Papers' note
VA-Bench tests MLLMs on an observe-reason-act-revise loop using RGB demonstrations, active camera choices, Cartesian commands, and execution feedback. The best model reached 100.0% target localization and 78.9% spatial-relations performance in the annotated run, but averaged only 53.93% task success across three runs. Active camera control helped sharply in one matched comparison, lifting success from 27.86% to 57.50%. Held-out geometry cut success by more than 30 points, and no model finished a strict long-horizon episode. HF Daily Papers' note
score 5