ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models
ReWEIGH penalizes visually unsupported token candidates during decoding without retraining the model.
The method pools rank-based visual evidence across image positions, then checks each candidate token against a token-specific reference from unlabeled images. It caches that evidence during prefill and applies only a bounded penalty when support is weak. The paper reports up to a 21.3% reduction in hallucinated object mentions on four 7B backbones, with average added latency of 1.33% per token after caching. The authors say the reductions also hold across six architecture families up to 32B parameters. ArXiv · AI/CL/LG's note
The method pools rank-based visual evidence across image positions, then checks each candidate token against a token-specific reference from unlabeled images. It caches that evidence during prefill and applies only a bounded penalty when support is weak. The paper reports up to a 21.3% reduction in hallucinated object mentions on four 7B backbones, with average added latency of 1.33% per token after caching. The authors say the reductions also hold across six architecture families up to 32B parameters. ArXiv · AI/CL/LG's note
score 5