Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
Mixed-modality retrievers can favor text so strongly that irrelevant text outranks relevant images.
The paper tests retrieval systems on corpora mixing text, images, and fused text-image documents. Performance stays stronger in single-modality settings but drops when modalities coexist, forming what the authors describe as a V-shaped curve. The worst disruption comes from irrelevant text, which degrades retrieval more than irrelevant images. Their proposed fix, Trident, trains text, image, and fused views as balanced positives and improves mixed-modality results across CLIP-based and VLM-based retrievers. HF Daily Papers' note
The paper tests retrieval systems on corpora mixing text, images, and fused text-image documents. Performance stays stronger in single-modality settings but drops when modalities coexist, forming what the authors describe as a V-shaped curve. The worst disruption comes from irrelevant text, which degrades retrieval more than irrelevant images. Their proposed fix, Trident, trains text, image, and fused views as balanced positives and improves mixed-modality results across CLIP-based and VLM-based retrievers. HF Daily Papers' note
score 4