PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
PragMatch tests whether vision-language models detect sarcasm’s pragmatic clash, not just image-text mismatch.
The paper introduces a 3,000-pair benchmark built from MMSD2.0, mixing original sarcastic examples with constructed literal and hard-negative pairs. Its experiments find model predictions shift when lexical, OCR-derived, or stylistic surface cues are masked or injected, even when the image-text relationship itself is unchanged. The authors frame that as evidence that current LVLMs remain vulnerable to shortcut learning in multimodal sarcasm detection. ArXiv · AI/CL/LG's note
The paper introduces a 3,000-pair benchmark built from MMSD2.0, mixing original sarcastic examples with constructed literal and hard-negative pairs. Its experiments find model predictions shift when lexical, OCR-derived, or stylistic surface cues are masked or injected, even when the image-text relationship itself is unchanged. The authors frame that as evidence that current LVLMs remain vulnerable to shortcut learning in multimodal sarcasm detection. ArXiv · AI/CL/LG's note
score 4