Megadose AI progress, ranked and analyzed.

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

· ArXiv · AI/CL/LG ·
EcoFrame uses a model’s own uncertainty and attention to decide whether more video frames are worth processing.

The paper presents a training-free scheduler for long-video VLM tasks, aimed at avoiding fixed one-shot frame selection and expensive agent-style search. It expands the frame budget only when output entropy suggests the current evidence is not enough. It also uses frame-level attention to search more densely in likely relevant time spans while keeping broader coverage when attention is diffuse. In reported tests on Video-MME, LongVideoBench, and MLVU, EcoFrame improves the accuracy-efficiency tradeoff, including 64.4 average accuracy on Qwen2.5-VL and speedups over AKS, BOLT, and A.I.R. ArXiv · AI/CL/LG's note

score 5

Categories: Research