Megadose AI progress, ranked and analyzed.

The Alignment Illusion in Multimodal Large Language Models

· ArXiv · AI/CL/LG ·
Standard alignment scores can look intact even when the visual stream has been wrecked.

The paper tests 13 multimodal models and finds that replacing visual tokens with Gaussian noise sharply hurts task accuracy. Four common scalar similarity measures do not reliably distinguish the corrupted stream from the original one. The authors argue the apparent alignment is partly induced by the shared language-model pathway, not by true content-level visual-text interaction. They propose a principal-angle gap metric that tracks accuracy more consistently under controlled corruption. ArXiv · AI/CL/LG's note

score 5

Categories: Research