Megadose AI progress, ranked and analyzed.

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

· ArXiv · AI/CL/LG ·
Quoting document evidence beat coordinate boxes for attribution across six open vision-language models.

The paper tests a quote-and-retrieve interface on a verified bilingual CiteVQA subset, using verbatim text evidence instead of model-drawn bounding boxes. Evidence recall rose from at most 8 points to 26-47, while hallucination rates roughly halved with little change in answer quality. The authors then use the same pipeline as a training scaffold, raising an 8B model’s strict attributed accuracy from 22.4 to 33.8 without region-level labels. ArXiv · AI/CL/LG's note

score 4

Categories: Research