Megadose Built for builders and researchers.

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

· ArXiv · AI/CL/LG ·
VisionWeave cuts visual token use by learning which image regions need fine detail and which can be compressed.

The paper reports an end-to-end method for multimodal LLMs that mixes fine and coarse visual representations through a gated spatial pooler and granularity router. On a Qwen3.8-27B-based model, it saves 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks. The authors say token-pruning baselines at a fixed 50% savings target preserve only 88% performance. Deployed on SGLang, VisionWeave delivers 2.3x throughput, with mean TTFT down 54.4% and mean TPOT down 60.6%. ArXiv · AI/CL/LG's note

score 6

Categories: Research