StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
StreamArena tests agents on full-length video streams where short-clip benchmarks hide memory failures.
The benchmark uses 243 videos averaging 88.8 minutes, with 3,646 open-ended QA pairs. It targets real-time perception, historical recall, proactive interaction, and multimodal tool use. The paper says recent-frame methods miss distant events, text summaries drop visual evidence, and compressed visual memory loses detail. Its proposed StreamMind architecture separates low-latency frontend monitoring from backend persistent multimodal memory, outperforming listed streaming baselines. HF Daily Papers' note
The benchmark uses 243 videos averaging 88.8 minutes, with 3,646 open-ended QA pairs. It targets real-time perception, historical recall, proactive interaction, and multimodal tool use. The paper says recent-frame methods miss distant events, text summaries drop visual evidence, and compressed visual memory loses detail. Its proposed StreamMind architecture separates low-latency frontend monitoring from backend persistent multimodal memory, outperforming listed streaming baselines. HF Daily Papers' note
score 5