Megadose AI progress, ranked and analyzed.

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

· ArXiv · AI/CL/LG ·
A single learned embedding is used to pull only query-relevant visual tokens from long image or video context.

ReToken is trained as a retrieval target over a pre-filled visual KV cache, avoiding the cost of processing every visual token at once. In the paper’s tests, it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points on Visual Haystacks. It also transfers zero-shot to long video on LVBench, where Qwen3VL-8B gains 8.0 points. The authors say training and long-video inference fit on one H100. ArXiv · AI/CL/LG's note

score 5

Categories: Research