Megadose AI progress, ranked and analyzed.

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them

· HF Daily Papers ·
The paper finds that merely adding an image can shift VLM judge labels, even when the image is irrelevant to the task.

The authors introduce MIST, a 200-sentence stress test where labels should be decided from text alone. Across 13 VLM judges, aligned images changed 20.5% of labels and misleading images changed 19.4%. Those changes usually did not track what the image depicted, and human agreement stayed unchanged across image conditions. The paper argues that substitutability tests can describe an evaluation setup as much as the model being tested. HF Daily Papers' note

score 4

Categories: Research