Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
Re$^3$Cap uses multimodal retrieval as a reasoning signal to refine captions without extra annotations.
The paper frames retrieval as a way to help large vision-language models catch hallucinations and missing details in image captions. Its system pairs a Caption Refinement Suggester with a Caption Quality Assessor to produce more accurate, detailed descriptions. The authors say experiments show Re$^3$Cap beating supervised fine-tuning, and report an 8.64% average gain over GRPO on relation reasoning in COCO-LN500. Accepted to EMNLP 2026 Main Conference. ArXiv · AI/CL/LG's note
The paper frames retrieval as a way to help large vision-language models catch hallucinations and missing details in image captions. Its system pairs a Caption Refinement Suggester with a Caption Quality Assessor to produce more accurate, detailed descriptions. The authors say experiments show Re$^3$Cap beating supervised fine-tuning, and report an 8.64% average gain over GRPO on relation reasoning in COCO-LN500. Accepted to EMNLP 2026 Main Conference. ArXiv · AI/CL/LG's note
score 4