SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
SpaceCast-Bench measures whether vision-language models can predict spatial outcomes they cannot directly see.
The benchmark uses 3,862 questions from 182 real-world scenes across static perception, local prediction, and global prediction. In the paper’s evaluation of 21 models, the best model scored 58.0%, well below the reported 87.2% human performance. The authors say bridge views help models integrate separated observations, while explicit 3D evidence helps more reliably than generated outcome images or videos. Fine-tuning on their generated data raised Qwen3-VL-4B from 34.0% to 65.7%. HF Daily Papers' note
The benchmark uses 3,862 questions from 182 real-world scenes across static perception, local prediction, and global prediction. In the paper’s evaluation of 21 models, the best model scored 58.0%, well below the reported 87.2% human performance. The authors say bridge views help models integrate separated observations, while explicit 3D evidence helps more reliably than generated outcome images or videos. Fine-tuning on their generated data raised Qwen3-VL-4B from 34.0% to 65.7%. HF Daily Papers' note
score 4