EviRover: Reinforcing Agentic Perception Beyond a Glance
EviRover is trained to gather extra evidence when one image pass is not enough.
The paper frames this as “perception under insufficient evidence,” where fine visual detail or current knowledge can decide the answer. The authors introduce two training sets, EviRover-SFT-5K and EviRover-RL-12K, plus EviLens, a 688-instance human-verified benchmark across five perception categories. A 4B EviRover model beats its backbone by 30 points on EviLens and shows transfer gains on WebEyes, standard perception tests, and BrowseComp-VL. Code, models, and data are released. HF Daily Papers' note
The paper frames this as “perception under insufficient evidence,” where fine visual detail or current knowledge can decide the answer. The authors introduce two training sets, EviRover-SFT-5K and EviRover-RL-12K, plus EviLens, a 688-instance human-verified benchmark across five perception categories. A 4B EviRover model beats its backbone by 30 points on EviLens and shows transfer gains on WebEyes, standard perception tests, and BrowseComp-VL. Code, models, and data are released. HF Daily Papers' note
score 5