Megadose AI progress, ranked and analyzed.

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

· HF Daily Papers ·
Frame selection, not frame count alone, is the biggest lever in long-video MLLMs.

The paper tests visual-token allocation under one controlled harness, varying selection, compression, and reinvestment separately. Eight query-selected frames beat sixteen uniformly spaced frames by 6.9 points on LongVideoBench’s hour-long bin. Halving each frame’s spatial budget cost at most 0.44 points, but the gain came when those saved tokens were spent on more compressed frames. The study also reports an AKS baseline bug and benchmark-harness gaps, arguing that selector comparisons need tighter controls. HF Daily Papers' note

score 4

Categories: Research