Megadose AI progress, ranked and analyzed.

Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

· HF Daily Papers ·
The benchmark separates modality bias from evidence-form bias by forcing conflicts across vision, audio, and text.

Tri-PvP contains 8,000 tri-modal conflict samples, with vision and audio represented either as perceptual signals or declarative propositions. Across five omni-modal models, the authors report a robust visual bias in most settings. They also find an asymmetry: models lean more toward perceptual evidence in vision, but propositional evidence in audio. Layer probing and contrastive decoding suggest the bias appears early in representations and is only partly reduced by surface-level fixes. HF Daily Papers' note

score 5

Categories: Research