Megadose AI progress, ranked and analyzed.

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

· HF Daily Papers ·
The paper tests whether multimodal models actually depend on the visual scratch work they generate.

See2Think combines a 1,200-problem benchmark with a trace format that records thoughts, visual actions, rendered states, and later reasoning. The authors find performance depends heavily on the model and inference setup, with no single setting winning across tasks. Models often choose relevant visual operations, but rendering those intermediate states faithfully is the main bottleneck. When task-relevant visual feedback is corrupted, accuracy drops by more than 10 points, suggesting the models do use those states in controlled cases. HF Daily Papers' note

score 5

Categories: Research