PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
No tested frontier MLLM reached 60% accuracy on isolated visual perception tasks.
PerceptionBench builds 3,000 verified short-answer questions around ten atomic perceptual capabilities. The authors say the benchmark is meant to separate perception failures from reasoning or domain-knowledge failures. Across sixteen frontier MLLMs, perception-related hallucination was the weakest capability on average, and similar total scores hid very different capability profiles. HF Daily Papers' note
PerceptionBench builds 3,000 verified short-answer questions around ten atomic perceptual capabilities. The authors say the benchmark is meant to separate perception failures from reasoning or domain-knowledge failures. Across sixteen frontier MLLMs, perception-related hallucination was the weakest capability on average, and similar total scores hid very different capability profiles. HF Daily Papers' note
score 5