RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models
RVSD claims a single-pass, training-free decoding method to curb visual hallucinations in vision-language models.
The paper says the approach combines token sparsification with Semantic-Space Visual Retrieval inside one decoding pass. Its token selection step is meant to remove redundant visual tokens while keeping critical visual information. The retrieval mechanism supplies visual compensation on demand through a shared semantic space. The authors report state-of-the-art hallucination mitigation, including under long-context generation settings. ArXiv · AI/CL/LG's note
The paper says the approach combines token sparsification with Semantic-Space Visual Retrieval inside one decoding pass. Its token selection step is meant to remove redundant visual tokens while keeping critical visual information. The retrieval mechanism supplies visual compensation on demand through a shared semantic space. The authors report state-of-the-art hallucination mitigation, including under long-context generation settings. ArXiv · AI/CL/LG's note
score 4