SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward
SCOUT trains VLMs to reason through 3D spatial steps, then rewards the process instead of only the final answer.
The paper says current VLMs struggle with robust spatial reasoning, partly because RL methods assign credit poorly across intermediate steps. SCOUT adds a structured chain-of-thought setup for 3D environmental perception and a multi-objective process-reward RL method. The authors also built SCOUT-24k, a synthetic structured spatial reasoning dataset. They report SCOUT-3B beating baselines, while SCOUT-7B outperforms GPT-4o by 4.28% and generalizes from single-image training to multi-image and video cases. ArXiv · AI/CL/LG's note
The paper says current VLMs struggle with robust spatial reasoning, partly because RL methods assign credit poorly across intermediate steps. SCOUT adds a structured chain-of-thought setup for 3D environmental perception and a multi-objective process-reward RL method. The authors also built SCOUT-24k, a synthetic structured spatial reasoning dataset. They report SCOUT-3B beating baselines, while SCOUT-7B outperforms GPT-4o by 4.28% and generalizes from single-image training to multi-image and video cases. ArXiv · AI/CL/LG's note
score 5