Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency
LoHi trades some per-frame resolution for denser video coverage, then adds selected high-resolution frames where detail matters.
The paper says this beats sparse native-resolution sampling at the same token budget across multiple VLMs and long-video benchmarks. Its high-resolution frames are chosen either from codec I-frame metadata or from CLIP-based relevance and diversity. The authors report a 10.6-point average accuracy gain over the native-resolution baseline and up to 7x lower front-end decoding latency on hour-long videos. HF Daily Papers' note
The paper says this beats sparse native-resolution sampling at the same token budget across multiple VLMs and long-video benchmarks. Its high-resolution frames are chosen either from codec I-frame metadata or from CLIP-based relevance and diversity. The authors report a 10.6-point average accuracy gain over the native-resolution baseline and up to 7x lower front-end decoding latency on hour-long videos. HF Daily Papers' note
score 4