SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
SpatialSpeak trains VLMs to reconstruct local geometry and global scene context before asking them to reason through spatial answers.
The paper frames that reconstruction as question answering, so geometric estimates and final reasoning use the same text-output interface. Its second stage has the model state question-relevant geometry, judge reliability, and refine answers with visual compensation when needed. On ReVSI, the reconstruction pretraining raises the benefit from spatial CoT training from 2.6 to 6.9 points. The authors report state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, including a ReVSI score of 62.8. HF Daily Papers' note
The paper frames that reconstruction as question answering, so geometric estimates and final reasoning use the same text-output interface. Its second stage has the model state question-relevant geometry, judge reliability, and refine answers with visual compensation when needed. On ReVSI, the reconstruction pretraining raises the benefit from spatial CoT training from 2.6 to 6.9 points. The authors report state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, including a ReVSI score of 62.8. HF Daily Papers' note
score 5