Megadose AI progress, ranked and analyzed.

Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

· ArXiv · AI/CL/LG ·
Longer prompts made corrupted-image answers more stable in the paper’s tests.

The authors argue that question wording changes cross-modal attention in VLMs, effectively filtering image patches by spatial frequency. Simple verbose padding broadened that filter, while semantically finer questions narrowed it and made models more sensitive to matching corruptions. On Qwen3-VL and LLaVA-OneVision over GQA and CLEVR, verbose paraphrases reduced drift variance by 70–81% on 8B models and improved accuracy under corruption. ArXiv · AI/CL/LG's note

score 5

Categories: Research