Megadose AI progress, ranked and analyzed.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

· HF Daily Papers ·
The paper claims a 35B Video-DR model beats Claude, GPT-5, and Gemini on a new video deep-research benchmark.

Video-DeepResearch is built to force multimodal agents to ground answers in video frames before turning to web retrieval. The authors say current systems lean too hard on text search or internal memory, so their pipeline unlocks tools in stages to reduce those shortcuts. They also introduce Video-DR-Bench, a 200-question multi-hop video QA benchmark built with human-AI collaboration. Their reported top model reaches 64.0% average accuracy, ahead of Claude-4.5-Sonnet at 59.0%, Gemini 2.5 Pro at 57.5%, and GPT-5 at 52.5%. Source: HF Daily Papers' note.

score 5

Categories: Research