FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
The benchmark finds streaming VLMs still miss fast-changing video events even when sampling gets denser.
FastBench tests high-dynamic real-world video streams with 306 QA pairs spanning eight domains and annotated evidence intervals. Its pipeline filters out questions answerable at 2 FPS, then verifies answers with trajectory tools and human inspection. Gemini-3.5-Flash tops the reported results at 50.7%, while Qwen3-VL-8B improves from 32.9% at 2 FPS to 44.6% at 24 FPS before gains saturate. The authors’ ProactiveFrame baseline beats sparse uniform sampling, but remains far below oracle-guided focusing, suggesting models struggle to decide when they need finer temporal detail from the stream alone. ArXiv · AI/CL/LG's note
FastBench tests high-dynamic real-world video streams with 306 QA pairs spanning eight domains and annotated evidence intervals. Its pipeline filters out questions answerable at 2 FPS, then verifies answers with trajectory tools and human inspection. Gemini-3.5-Flash tops the reported results at 50.7%, while Qwen3-VL-8B improves from 32.9% at 2 FPS to 44.6% at 24 FPS before gains saturate. The authors’ ProactiveFrame baseline beats sparse uniform sampling, but remains far below oracle-guided focusing, suggesting models struggle to decide when they need finer temporal detail from the stream alone. ArXiv · AI/CL/LG's note
score 5