Megadose AI progress, ranked and analyzed.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

· HF Daily Papers ·
The benchmark finds a wide human-model gap on global spatial reasoning from long videos.

GST-Bench tests whether VLMs can infer unseen viewpoints and align egocentric video observations with top-down scene maps. It is built from 6,790 minutes of synthetic video with human-verified VQA questions. Across 22 state-of-the-art VLMs, the best zero-shot score is 42.68, versus 79.08 for humans. A local version suggests the weakness is not local spatial perception, but consolidating long-horizon observations into a consistent global scene. HF Daily Papers' note

score 5

Categories: Research