GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
The benchmark finds a wide human-model gap on global spatial reasoning from long videos.
GST-Bench tests whether VLMs can infer unseen viewpoints and align egocentric video observations with top-down scene maps. It is built from 6,790 minutes of synthetic video with human-verified VQA questions. Across 22 state-of-the-art VLMs, the best zero-shot score is 42.68, versus 79.08 for humans. A local version suggests the weakness is not local spatial perception, but consolidating long-horizon observations into a consistent global scene. HF Daily Papers' note
GST-Bench tests whether VLMs can infer unseen viewpoints and align egocentric video observations with top-down scene maps. It is built from 6,790 minutes of synthetic video with human-verified VQA questions. Across 22 state-of-the-art VLMs, the best zero-shot score is 42.68, versus 79.08 for humans. A local version suggests the weakness is not local spatial perception, but consolidating long-horizon observations into a consistent global scene. HF Daily Papers' note
score 5