Visual Credit Audit for Multimodal Spatial Reasoning
A new audit finds some spatial-benchmark wins are right answers without real image credit.
The paper introduces Visual Credit Audit, a way to separate correctness from whether the image actually supported a multimodal model’s yes/no spatial decision. Across four open MLLMs and two spatial benchmarks, 12.73% to 26.25% of decisions were correct but uncredited. Image permutation sharply reduced dependence-credited correctness, while relation-specific tests showed that blank or text-only controls miss whether models respond to the relevant visual relation. ArXiv · AI/CL/LG's note
The paper introduces Visual Credit Audit, a way to separate correctness from whether the image actually supported a multimodal model’s yes/no spatial decision. Across four open MLLMs and two spatial benchmarks, 12.73% to 26.25% of decisions were correct but uncredited. Image permutation sharply reduced dependence-credited correctness, while relation-specific tests showed that blank or text-only controls miss whether models respond to the relevant visual relation. ArXiv · AI/CL/LG's note
score 4