DEPICT: Scoring Text-to-Image Alignment by Answer Agreement
DEPICT scores image-prompt alignment by comparing what a model answers from the image with what the caption alone implies.
The paper says this avoids the fixed-YES setup used by some decomposed evaluation methods, which can punish faithful images when the expected answer is not yes. Its agreement rule raises negation accuracy from 19% to 88% in the authors’ report. DEPICT also combines that decomposed agreement score with a holistic score to restore context. The authors evaluate it across five benchmarks and eleven backbones, finding it beats other training-free metrics and outperforms fine-tuned evaluators on two of three human-correlation benchmarks. HF Daily Papers' note
The paper says this avoids the fixed-YES setup used by some decomposed evaluation methods, which can punish faithful images when the expected answer is not yes. Its agreement rule raises negation accuracy from 19% to 88% in the authors’ report. DEPICT also combines that decomposed agreement score with a holistic score to restore context. The authors evaluate it across five benchmarks and eleven backbones, finding it beats other training-free metrics and outperforms fine-tuned evaluators on two of three human-correlation benchmarks. HF Daily Papers' note
score 5