Megadose AI progress, ranked and analyzed.

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

· HF Daily Papers ·
CoVeR keeps 3D scene coverage high while cutting visual tokens to about 8%.

The paper presents a deterministic, training-free token selector for multi-view 3D reasoning with vision-language models. It uses token coordinates, not learned importance scores, to avoid overselecting duplicate regions and leaving parts of the scene uncovered. In experiments, it beats prior state-of-the-art methods across three 3D reasoning benchmarks and works as a plug-in module across four VLMs. With about 8% of visual tokens, it retains 93.5% of full-token performance and is 3.9 points ahead of the prior SOTA on average. HF Daily Papers' note

score 4

Categories: Research