Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
The paper argues that these models often have the visual facts encoded, but fail when deciding whether to use them.
The authors test this with image reconstruction probes and a new benchmark, WhatIfVis, covering spatial-temporal, color, count, size, and weight questions. They find coarse visual attributes can be recovered from final-layer image tokens in frozen MLLMs, pointing the failure downstream of perception. Vanilla models remain unstable even when told to use or ignore the image, while supervised fine-tuning and a learned steering vector improve controllability. Source: HF Daily Papers' note
The authors test this with image reconstruction probes and a new benchmark, WhatIfVis, covering spatial-temporal, color, count, size, and weight questions. They find coarse visual attributes can be recovered from final-layer image tokens in frozen MLLMs, pointing the failure downstream of perception. Vanilla models remain unstable even when told to use or ignore the image, while supervised fine-tuning and a learned steering vector improve controllability. Source: HF Daily Papers' note
score 5