PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
The paper frames caption grounding as choosing the right masks for each phrase, then tests that against a new panoptic benchmark.
PANORAMA pairs dense image captions with pixel-level masks for foreground objects and background regions. The authors introduce PanoCaps, with human annotations, entity-level image-text alignments, and a gPQ metric for judging both wording and mask quality. Their model uses phrase-conditioned mask proposals and learns which masks match each referring phrase, including cases where one phrase maps to multiple instances. It reports the best overall grounding on PanoCaps and competitive results on other pixel-level grounding tasks. HF Daily Papers' note
PANORAMA pairs dense image captions with pixel-level masks for foreground objects and background regions. The authors introduce PanoCaps, with human annotations, entity-level image-text alignments, and a gPQ metric for judging both wording and mask quality. Their model uses phrase-conditioned mask proposals and learns which masks match each referring phrase, including cases where one phrase maps to multiple instances. It reports the best overall grounding on PanoCaps and competitive results on other pixel-level grounding tasks. HF Daily Papers' note
score 5