SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
SpatialCLI trains VLMs to use specialist spatial tools, then keep much of the gain after the tools are removed.
The paper frames the gap as one between broad task reasoning and fine visual perception. Its three stages are Call, Learn, and Internalize: expose tools, improve tool use with SFT and RL, then verbalize successful trajectories so the model absorbs the capability. The authors also introduce SpatialCLI-Bench, a 516-example benchmark covering localization, segmentation, depth, and pose. On MindCube, they report Qwen3-VL-8B-Instruct rising from 29.3% to 84.6% with tools, and retaining 73.8% without them after internalization. HF Daily Papers' note
The paper frames the gap as one between broad task reasoning and fine visual perception. Its three stages are Call, Learn, and Internalize: expose tools, improve tool use with SFT and RL, then verbalize successful trajectories so the model absorbs the capability. The authors also introduce SpatialCLI-Bench, a 516-example benchmark covering localization, segmentation, depth, and pose. On MindCube, they report Qwen3-VL-8B-Instruct rising from 29.3% to 84.6% with tools, and retaining 73.8% without them after internalization. HF Daily Papers' note
score 5