Megadose AI progress, ranked and analyzed.

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

· HF Daily Papers ·
The survey maps where video LLM inference costs pile up, from frames and encoders to token reduction and decoding.

The paper covers efficiency methods for visual and audiovisual VideoLLMs that report concrete savings in parameters, FLOPs, latency, memory, or token counts. It sorts those methods by pipeline stage, including frame sampling, modality encoding, connector-level compression, and LLM prefill/decoding. The authors separate shared-protocol accuracy-cost comparisons from weaker cross-paper evidence. They also flag gaps in audiovisual efficiency work and standardized evaluation. HF Daily Papers' note

score 4

Categories: Research