Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
Lightweight visual scaffolds raised VLM spatial-planning accuracy by as much as 34 percentage points.
The paper tests grid-based spatial reasoning in SPaRC while keeping the reasoning problem fixed and changing how the visual input is presented. The authors report that clearer visual structure reduces grounding-related errors across multiple VLMs. The scaffolds also made GRPO-based training more useful, adding up to 4.6 accuracy points where the original visual input produced near-zero gains. Rule reasoning remained comparatively hard. ArXiv · AI/CL/LG's note
The paper tests grid-based spatial reasoning in SPaRC while keeping the reasoning problem fixed and changing how the visual input is presented. The authors report that clearer visual structure reduces grounding-related errors across multiple VLMs. The scaffolds also made GRPO-based training more useful, adding up to 4.6 accuracy points where the original visual input produced near-zero gains. Rule reasoning remained comparatively hard. ArXiv · AI/CL/LG's note
score 5