Megadose AI progress, ranked and analyzed.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

· HF Daily Papers ·
The paper argues that these models often have the visual facts encoded, but fail when deciding whether to use them.

The authors test this with image reconstruction probes and a new benchmark, WhatIfVis, covering spatial-temporal, color, count, size, and weight questions. They find coarse visual attributes can be recovered from final-layer image tokens in frozen MLLMs, pointing the failure downstream of perception. Vanilla models remain unstable even when told to use or ignore the image, while supervised fine-tuning and a learned steering vector improve controllability. Source: HF Daily Papers' note

score 5

Categories: Research