Megadose AI progress, ranked and analyzed.

An Exam for Active Observers

· ArXiv · AI/CL/LG ·
The benchmark’s top model solved only 10.6% of items, while three human participants averaged 96.1%.

The paper introduces ActiveVision, a 17-task benchmark meant to test whether multimodal models can actively redirect visual attention across repeated observations. GPT-5.5 at the highest exposed reasoning-effort tier scored best among the evaluated models, but still failed completely on 11 tasks. Claude Fable 5 scored 3.5%, despite strong results on reasoning and coding leaderboards. The authors say the gap often remained even when models wrote and ran their own vision code, because detecting code failures required the same active perception they lacked. ArXiv · AI/CL/LG's note

score 6

Categories: Research