Megadose Built for builders and researchers.

SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models

· ArXiv · AI/CL/LG ·
The benchmark finds current vision-language models far behind humans at predicting unseen spatial outcomes.

SpaceCast-Bench tests whether models can observe a scene, account for a change, and infer the resulting spatial relationships. Its 3,862 questions cover 182 real-world scenes across static perception, local prediction, and global prediction. In the paper’s evaluation of 21 models, the best model scored 58.0%, compared with 87.2% for humans, while spatially specialized models stayed near random chance. The authors also report that bridge views and explicit 3D evidence help more reliably than generated outcome images or videos. ArXiv · AI/CL/LG's note

score 5

Categories: Research