WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
WM-VLM adds a visual-state branch to a pretrained VLM and reports gains of up to 39.25 points on mental rotation tasks.
The paper tests whether vision-language models can reason through generated intermediate images, not just text. Its two-stage training first teaches the model to predict the next visual state, then to use that state in answering spatial questions. The authors use programmatically built 2D and 3D tasks with verifiable intermediate states. Ablations found performance dropped sharply when those generated states were removed or corrupted. HF Daily Papers' note
The paper tests whether vision-language models can reason through generated intermediate images, not just text. Its two-stage training first teaches the model to predict the next visual state, then to use that state in answering spatial questions. The authors use programmatically built 2D and 3D tasks with verifiable intermediate states. Ablations found performance dropped sharply when those generated states were removed or corrupted. HF Daily Papers' note
score 4