An Exam for Active Observers
ActiveVision tests whether multimodal models can look again, and the reported results are stark.
The paper introduces a 17-task benchmark built to require repeated visual perception rather than one static image description. The authors report that GPT-5.5 solves 10.6% of items at its highest exposed reasoning-effort tier, while Claude Fable 5 solves 3.5%. Three human participants average 96.1%. The gap largely remains even when models write and run vision code, because detecting code failures requires the same active perception they lack. HF Daily Papers' note
The paper introduces a 17-task benchmark built to require repeated visual perception rather than one static image description. The authors report that GPT-5.5 solves 10.6% of items at its highest exposed reasoning-effort tier, while Claude Fable 5 solves 3.5%. Three human participants average 96.1%. The gap largely remains even when models write and run vision code, because detecting code failures requires the same active perception they lack. HF Daily Papers' note
score 5