SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models
SinkPruner targets high-norm visual tokens that other pruning methods tend to keep, arguing they are redundant rather than useful.
The paper presents a training-free pruning framework for multimodal model inference. Its visual sanitizer removes high-norm redundancies, then a text-guided pruner keeps tokens aligned with the query. The authors report tests across twelve image-language and four video-language benchmarks. Under 89% token reduction, SinkPruner preserves 96.5% of LLaVA-1.5 performance and 91.8% of Qwen2.5-VL performance. ArXiv · AI/CL/LG's note
The paper presents a training-free pruning framework for multimodal model inference. Its visual sanitizer removes high-norm redundancies, then a text-guided pruner keeps tokens aligned with the query. The authors report tests across twelve image-language and four video-language benchmarks. Under 89% token reduction, SinkPruner preserves 96.5% of LLaVA-1.5 performance and 91.8% of Qwen2.5-VL performance. ArXiv · AI/CL/LG's note
score 4