VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
The benchmark is built to test video models as tool-using agents, and frontier systems still fall below 60% accuracy.
VideoGAIA turns video understanding into a multi-turn process where models must inspect video, call external tools, collect added evidence, and combine it across turns. It includes 271 model-human co-designed tasks across complex real-world scenarios. Each item was checked by three human experts for correctness and difficulty. The authors argue that older single-turn video QA benchmarks are saturating, with leading models near 90% on Video-MME. HF Daily Papers' note
VideoGAIA turns video understanding into a multi-turn process where models must inspect video, call external tools, collect added evidence, and combine it across turns. It includes 271 model-human co-designed tasks across complex real-world scenarios. Each item was checked by three human experts for correctness and difficulty. The authors argue that older single-turn video QA benchmarks are saturating, with leading models near 90% on Video-MME. HF Daily Papers' note
score 5