Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
Quoting evidence, then retrieving its location, beat direct bounding-box attribution across six open vision-language models.
The paper tests whether attribution failures in visual document understanding are partly caused by forcing models to answer with coordinates. On a verified bilingual CiteVQA subset, evidence recall rose from at most 8 points with coordinates to 26-47 with the quote-and-retrieve interface, while hallucination rates roughly halved. Answer quality changed little. The authors then used the pipeline as a training scaffold, improving an 8B model’s strict attributed accuracy from 22.4 to 33.8 without region-level labels. HF Daily Papers' note
The paper tests whether attribution failures in visual document understanding are partly caused by forcing models to answer with coordinates. On a verified bilingual CiteVQA subset, evidence recall rose from at most 8 points with coordinates to 26-47 with the quote-and-retrieve interface, while hallucination rates roughly halved. Answer quality changed little. The authors then used the pipeline as a training scaffold, improving an 8B model’s strict attributed accuracy from 22.4 to 33.8 without region-level labels. HF Daily Papers' note
score 4