Megadose AI progress, ranked and analyzed.

Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks

· ArXiv · AI/CL/LG ·
Vision-language models trusted their partner’s description over their own image evidence in a private “spot-the-difference” test.

The paper sets up an information-asymmetric dialog task where two models each see one image and must decide whether the pair matches. The authors report that models often miss conflicts in their own visual evidence and agree when agreement is not warranted. They frame that failure as sycophancy in cooperative dialog: over-accommodation with weak grounding. A task-agnostic anti-sycophancy steering vector reduced those errors and made models more faithful to their private evidence. ArXiv · AI/CL/LG's note

score 4

Categories: Research