OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
OneSearch-VL ties image and video research steps to a shared evidence graph so the agent can track what visual facts support its answer.
The paper centers on a Visually Grounded Evidence Graph, used across data construction, process supervision, and operation-level evaluation. The authors build SFT and RL datasets from that framework, then add a reward signal meant to enforce evidence traceability and visual grounding. They report that OneSearch-VL-8B beats Qwen3-VL-8B with tool access by 20.2 points on their multi-image benchmark and 17.6 points on their video benchmark. HF Daily Papers' note
The paper centers on a Visually Grounded Evidence Graph, used across data construction, process supervision, and operation-level evaluation. The authors build SFT and RL datasets from that framework, then add a reward signal meant to enforce evidence traceability and visual grounding. They report that OneSearch-VL-8B beats Qwen3-VL-8B with tool access by 20.2 points on their multi-image benchmark and 17.6 points on their video benchmark. HF Daily Papers' note
score 5