Reason Through the Latent! Making Latent Visual Reasoning Necessary
CVRR forces the answer to depend on recurrent latent state, not leftover multimodal context.
The paper argues that having visual information inside hidden states is not enough to prove a model is reasoning through them. Its CVRR setup removes visual states and the original multimodal KV cache before decoding, leaving only the final recurrent state to carry image-conditioned information. On V*, MMVP, BLINK, and MME-RealWorld-Lite, CVRR keeps strong performance under that constraint while other latent reasoners do not match it. The authors also report causal interventions showing predictions remain sensitive to recurrent content and that persistent visual evidence changes the recurrent path. Source: HF Daily Papers' note
The paper argues that having visual information inside hidden states is not enough to prove a model is reasoning through them. Its CVRR setup removes visual states and the original multimodal KV cache before decoding, leaving only the final recurrent state to carry image-conditioned information. On V*, MMVP, BLINK, and MME-RealWorld-Lite, CVRR keeps strong performance under that constraint while other latent reasoners do not match it. The authors also report causal interventions showing predictions remain sensitive to recurrent content and that persistent visual evidence changes the recurrent path. Source: HF Daily Papers' note
score 4