Megadose AI progress, ranked and analyzed.

An Exam for Active Observers

· HF Daily Papers ·
ActiveVision tests whether multimodal models can look again, and the reported results are stark.

The paper introduces a 17-task benchmark built to require repeated visual perception rather than one static image description. The authors report that GPT-5.5 solves 10.6% of items at its highest exposed reasoning-effort tier, while Claude Fable 5 solves 3.5%. Three human participants average 96.1%. The gap largely remains even when models write and run vision code, because detecting code failures requires the same active perception they lack. HF Daily Papers' note

score 5

Categories: Research