Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Symbal is built to find recurring caption errors tied to specific visual features, even without access to the captioning model.
The paper defines the task as systematic misalignment detection for MLLM-generated image captions. Symbal uses a two-stage setup with off-the-shelf foundation models to identify those patterns and summarize them in natural language. The authors also introduce SymbalBench, with 1.7 million image-text pairs across natural and medical images and 420 annotated datasets. On that benchmark, Symbal identifies systematic misalignments in 63.8% of datasets, nearly four times the closest baseline. ArXiv · AI/CL/LG's note
The paper defines the task as systematic misalignment detection for MLLM-generated image captions. Symbal uses a two-stage setup with off-the-shelf foundation models to identify those patterns and summarize them in natural language. The authors also introduce SymbalBench, with 1.7 million image-text pairs across natural and medical images and 420 annotated datasets. On that benchmark, Symbal identifies systematic misalignments in 63.8% of datasets, nearly four times the closest baseline. ArXiv · AI/CL/LG's note
score 4