PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
PANORAMA is built to make image captions line up with actual pixels, across both objects and background regions.
The paper introduces PanoCaps, a human-annotated benchmark for dense captions with near-complete pixel coverage and entity-level image-text alignments. Its model treats grounding as a mask proposal selection problem, using phrase-conditioned mask candidates and learning which masks match each caption phrase. The authors report that PANORAMA leads overall grounding results on PanoCaps and matches or beats specialized models on several pixel-level grounding tasks. ArXiv · AI/CL/LG's note
The paper introduces PanoCaps, a human-annotated benchmark for dense captions with near-complete pixel coverage and entity-level image-text alignments. Its model treats grounding as a mask proposal selection problem, using phrase-conditioned mask candidates and learning which masks match each caption phrase. The authors report that PANORAMA leads overall grounding results on PanoCaps and matches or beats specialized models on several pixel-level grounding tasks. ArXiv · AI/CL/LG's note
score 5