ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
ConsiSpace adds geometry-consistent memory and post-SFT consistency rewards to stabilize video spatial reasoning.
The paper targets MLLMs that lose reliable spatial evidence across long videos and changing viewpoints. Its framework stores implicit evidence tokens alongside explicit geometric cues, then trains with answer-, metric-, and topology-consistency rewards after supervised fine-tuning. The authors report gains on VSI-Bench, OSI-Bench, and MMSI-Video-Bench, with an average improvement of 12.6 points over the strongest baselines. HF Daily Papers' note
The paper targets MLLMs that lose reliable spatial evidence across long videos and changing viewpoints. Its framework stores implicit evidence tokens alongside explicit geometric cues, then trains with answer-, metric-, and topology-consistency rewards after supervised fine-tuning. The authors report gains on VSI-Bench, OSI-Bench, and MMSI-Video-Bench, with an average improvement of 12.6 points over the strongest baselines. HF Daily Papers' note
score 4