An Exam for Active Observers
Why AI vision systems fail at looking carefully, the way humans do
Today's most advanced AI image-understanding systems—including GPT-4.5 and Claude—cannot perform active observation, the repeated, purposeful looking that humans use to solve visual tasks. When tested on 17 new benchmark tasks designed to require this skill, the best model solved only 10.6% of items, while average humans scored 96.1%, suggesting a fundamental gap in how these systems perceive images.
AI systems that cannot look carefully will fail at tasks requiring sustained visual attention—medical diagnosis, quality inspection, scientific analysis, and navigation in complex scenes. Even when given the ability to write their own code to re-examine images, current models produce unreliable results and cannot catch their own mistakes, pointing to a core architectural flaw that researchers must now address.