Megadose Built for builders and researchers.

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

· ArXiv · AI/CL/LG ·
The paper isolates a failure mode where vision-language models see the right objects and attributes but attach an attribute to the wrong nearby instance.

The authors define the problem as Dense Same-Class Attribute Misbinding, or DSCAM. They introduce InstaBind-Lite, a 524-image benchmark with boxed same-class instances and deterministic questions designed to identify where the mistaken attribute came from. In tests across five open-source and two commercial/API models, open-source systems averaged a 19.84% misbinding rate, while API systems averaged 7.55%. Most identifiable transfers came from adjacent instances, and localization or instance-first prompting helped only some models. ArXiv · AI/CL/LG's note

score 4

Categories: Research