Megadose AI progress, ranked and analyzed.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

· HF Daily Papers ·
The paper defines a new video-captioning task where captions must tie phrases to multiple reference images.

RefCaptioner is a two-stage post-training framework built for factual video descriptions with explicit reference grounding. The authors say it improves reference selection, phrase-level binding, distractor rejection, and consistency across references while keeping standard captioning ability. They also introduce a training corpus of 20,000 videos with 171,354 reference images, plus MRVBench for evaluating the task. Experiments report the best overall performance among open-source models, with human evaluators preferring its captions.

HF Daily Papers' note

score 4

Categories: Research