Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
FactoSR trains vision-language models to check space as separate XY, depth, and time constraints instead of one collapsed visual problem.
The paper argues VLMs struggle because they learn from 2D projections while spatial reasoning needs 3D geometry and temporal continuity. Its framework, FactoSR, uses reinforcement learning around three verifiable sub-objectives: planar correspondence, depth consistency, and temporal reversibility. The authors report gains of 5.9% on VSI-Bench and 4.5% on All-Angles-Bench. The paper is marked accepted by ECCV 2026. HF Daily Papers' note
The paper argues VLMs struggle because they learn from 2D projections while spatial reasoning needs 3D geometry and temporal continuity. Its framework, FactoSR, uses reinforcement learning around three verifiable sub-objectives: planar correspondence, depth consistency, and temporal reversibility. The authors report gains of 5.9% on VSI-Bench and 4.5% on All-Angles-Bench. The paper is marked accepted by ECCV 2026. HF Daily Papers' note
score 5