Megadose AI progress, ranked and analyzed.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

· HF Daily Papers ·
OmniPack claims near-original omni-modal model performance at a fraction of the compute.

The paper proposes a training-free token compression framework for audio-visual LLMs. It compresses structurally redundant visual and audio tokens before the LLM, then refines task-relevant representations inside the LLM using text guidance and audio-visual collaboration. Across five benchmarks and three Omni-LLM backbones, the authors report the best performance-efficiency trade-off against existing methods. On Qwen2.5-Omni-7B, it keeps 98.0% of original performance while cutting FLOPs to 16.7%, and 92.9% performance at 6.8% FLOPs. HF Daily Papers' note

score 5

Categories: Research