DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
DeepVoyager-VL makes visual evidence part of the search process, not just the prompt or final answer.
The paper proposes a long-horizon multimodal search framework where an agent can actively acquire and load images during intermediate reasoning. Its training data is synthesized from a multimodal event graph designed to create problems with visual dependencies across longer chains. The authors say the models are fine-tuned on that data without reinforcement learning, then tested across ten multimodal search benchmarks. HF Daily Papers' note
The paper proposes a long-horizon multimodal search framework where an agent can actively acquire and load images during intermediate reasoning. Its training data is synthesized from a multimodal event graph designed to create problems with visual dependencies across longer chains. The authors say the models are fine-tuned on that data without reinforcement learning, then tested across ten multimodal search benchmarks. HF Daily Papers' note
score 5