Beyond Selection: Token Parameterization for Extreme Visual Token Compression
Braco compresses visual tokens by re-parameterizing them, not just pruning them.
The paper frames extreme visual-token compression as a split problem: preserve a useful subspace, then organize coordinates so the model can still learn and align vision with language. Its proposed coder, Braco, combines transform-basis truncation, coordinate embeddings, budget-aware orthogonal re-parameterization, and lightweight spatial residual tokens. In experiments, it sets a strong accuracy-efficiency frontier at 23x to 64x compression and remains competitive at 144x. The authors report 95.2% accuracy with 84.2% to 86.7% lower prefill FLOPs versus the uncompressed upper bound, plus up to about 36% end-to-end speedup over prior methods. HF Daily Papers' note
The paper frames extreme visual-token compression as a split problem: preserve a useful subspace, then organize coordinates so the model can still learn and align vision with language. Its proposed coder, Braco, combines transform-basis truncation, coordinate embeddings, budget-aware orthogonal re-parameterization, and lightweight spatial residual tokens. In experiments, it sets a strong accuracy-efficiency frontier at 23x to 64x compression and remains competitive at 144x. The authors report 95.2% accuracy with 84.2% to 86.7% lower prefill FLOPs versus the uncompressed upper bound, plus up to about 36% end-to-end speedup over prior methods. HF Daily Papers' note
score 4