Megadose AI progress, ranked and analyzed.

Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation

· ArXiv · AI/CL/LG ·
SCoRE keeps only cited visual evidence during search, then reloads and orders the original images before answering.

The paper targets visual RAG over document pages where the needed evidence may be tiny, scattered, or easy to lose in an agent’s exploration trail. Its method maintains a textual ledger of query-relevant observations with source pointers, instead of carrying raw trajectories or compressed memories into the final answer. At the end, it reconstructs a bounded visual context from the referenced images and ties claims back to indexed image evidence. Training combines filtered trajectory distillation with reinforcement learning rewards for evidence coverage, compact consolidation, and answer correctness. ArXiv · AI/CL/LG's note

score 4

Categories: Research