Megadose Built for builders and researchers.

V-CoLA: Vision Token Compression with Linear Attention

· HF Daily Papers ·
V-CoLA keeps linear-attention VLMs close to baseline while cutting vision tokens in half.

The paper says older token-compression methods built for softmax attention degrade on newer hybrid linear-attention models. V-CoLA is training-free and uses a uniqueness-aware importance score plus adaptive token merging. In the reported experiments, it preserves 99.5% of original performance with 50% of vision tokens, and more than 88% with 12.5%. The authors also report 1.86x to 6.15x prefill speedups. HF Daily Papers' note

score 4

Categories: Research