RefCaptioner: Multi-Reference Image-Grounded Video Captioning
The paper defines a new video-captioning task where captions must tie phrases to multiple reference images.
RefCaptioner is a two-stage post-training framework built for factual video descriptions with explicit reference grounding. The authors say it improves reference selection, phrase-level binding, distractor rejection, and consistency across references while keeping standard captioning ability. They also introduce a training corpus of 20,000 videos with 171,354 reference images, plus MRVBench for evaluating the task. Experiments report the best overall performance among open-source models, with human evaluators preferring its captions.
HF Daily Papers' note
RefCaptioner is a two-stage post-training framework built for factual video descriptions with explicit reference grounding. The authors say it improves reference selection, phrase-level binding, distractor rejection, and consistency across references while keeping standard captioning ability. They also introduce a training corpus of 20,000 videos with 171,354 reference images, plus MRVBench for evaluating the task. Experiments report the best overall performance among open-source models, with human evaluators preferring its captions.
HF Daily Papers' note
score 4