EviRover: Reinforcing Agentic Perception Beyond a Glance
The paper frames hard visual queries as an evidence-gathering task, not a single-image prediction.
EviRover is trained to resolve “perception under insufficient evidence” through interaction, using supervised fine-tuning and agentic reinforcement learning. The authors built two training sets, EviRover-SFT-5K and EviRover-RL-12K, plus EviLens, a 688-example human-verified benchmark across five perception categories. They report that a 4B EviRover beats its backbone by 30 points on EviLens and improves BrowseComp-VL by 15 points. Code, models, and data are released. ArXiv · AI/CL/LG's note
EviRover is trained to resolve “perception under insufficient evidence” through interaction, using supervised fine-tuning and agentic reinforcement learning. The authors built two training sets, EviRover-SFT-5K and EviRover-RL-12K, plus EviLens, a 688-example human-verified benchmark across five perception categories. They report that a 4B EviRover beats its backbone by 30 points on EviLens and improves BrowseComp-VL by 15 points. Code, models, and data are released. ArXiv · AI/CL/LG's note
score 5