Megadose AI progress, ranked and analyzed.

VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification

· HF Daily Papers ·
VDiff-Bench tests whether vision-language models can spot small changes between two near-matching images, and many still fail on low-level differences.

The benchmark has 1,756 four-choice questions across 10 change types, including position, color, texture, OCR/text, illumination, and noise/resolution. Its distractors are designed to be close to the true answer, so models have to identify the exact difference rather than notice a broad mismatch. In experiments on 11 open- and closed-source MLLMs, performance was uneven, with especially weak results on subtle noise and texture changes. The paper says some 7-8B open models scored 52.5-70.6% on semantic changes but only 8.7-33.3% on low-level ones. HF Daily Papers' note

score 5

Categories: Research