SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
Injected scene text pushed every tested multimodal model toward the wrong geographic target.
SIGNPOST-Bench tests how MLLMs handle images when visual evidence and localized text disagree. Its 25,555 image variants come from 5,111 counterfactual groups across four datasets. In adversarial variants, median localization error rose from 282 km to 1,347 km. Across models, 6.5% to 20.1% of geocodable adversarial predictions landed within 50 km of the injected target. HF Daily Papers' note
SIGNPOST-Bench tests how MLLMs handle images when visual evidence and localized text disagree. Its 25,555 image variants come from 5,111 counterfactual groups across four datasets. In adversarial variants, median localization error rose from 282 km to 1,347 km. Across models, 6.5% to 20.1% of geocodable adversarial predictions landed within 50 km of the injected target. HF Daily Papers' note
score 5