ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
The paper proposes a benchmark for testing whether vision-language models are pulled off course by irrelevant or false context.
ENTRAP-VL is a manually curated set of 1,500 items across eight categories. It separates textual entrainment from visual entrainment, with eight text-context conditions and three visual-context conditions. The authors argue that multimodal models need a purpose-built probe because misleading context can come through either the image or the words. They say the work is an instrument and taxonomy, not a measurement of any specific model. HF Daily Papers' note
ENTRAP-VL is a manually curated set of 1,500 items across eight categories. It separates textual entrainment from visual entrainment, with eight text-context conditions and three visual-context conditions. The authors argue that multimodal models need a purpose-built probe because misleading context can come through either the image or the words. They say the work is an instrument and taxonomy, not a measurement of any specific model. HF Daily Papers' note
score 4