CAPEval: A Decoupled Caption Evaluation across Understanding and Generation
Caption quality is split into coverage for understanding and precision for generation.
CAPEval evaluates captions by separating how much ground-truth visual content they cover from how often their stated claims are correct. The benchmark uses human-written captions and human-verified atomic checklist items. In experiments across 10 captioners and four model families, coverage correlated more strongly with multimodal understanding, while precision better predicted text-to-image generation performance. HF Daily Papers' note
CAPEval evaluates captions by separating how much ground-truth visual content they cover from how often their stated claims are correct. The benchmark uses human-written captions and human-verified atomic checklist items. In experiments across 10 captioners and four model families, coverage correlated more strongly with multimodal understanding, while precision better predicted text-to-image generation performance. HF Daily Papers' note
score 4