InSight-doc: Agentic Visual Perception for Long-Document Understanding
InSight-doc spends high resolution only where the document evidence demands it.
The paper describes an agentic visual perception framework that begins with low-resolution pages, then zooms into selected regions for finer evidence. Its training data includes 17.9K supervised examples with region-level zoom trajectories and 19.2K hard reinforcement-learning examples. The authors report 4.3 to 16.4 point gains over baseline document VQA results, with more than 40% lower hallucination on long documents and 41% to 68% lower inference latency. HF Daily Papers' note
The paper describes an agentic visual perception framework that begins with low-resolution pages, then zooms into selected regions for finer evidence. Its training data includes 17.9K supervised examples with region-level zoom trajectories and 19.2K hard reinforcement-learning examples. The authors report 4.3 to 16.4 point gains over baseline document VQA results, with more than 40% lower hallucination on long documents and 41% to 68% lower inference latency. HF Daily Papers' note
score 5