Discover the Best AI Tools
We test the top AI tools for writing, video, images & music — so you don't have to.
Explore AI ToolsAI Wasn't Dumb — It Was Blind: MIT's VISTA Gave Claude Opus a Perfect ARC-AGI-3 Score
For months, the story around AI reasoning benchmarks has been simple: models keep getting smarter, scores keep climbing. A new paper from MIT flips that story on its head. The models weren't dumb, the researchers argue — they were blind.
The experiment
On October 1, a team at MIT — Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, and Kaiming He — uploaded a paper to arXiv describing VISTA, a "visual harness" for AI agents. Instead of feeding Claude Opus 5.0 the official ARC-AGI-3 interface — a 64x64 grid of numbers representing each game state — they gave it plain 512x512 screenshots of the actual games, plus a lossless visual memory: every frame the game returns is archived in its original form, and the model can flip back through any of them mid-reasoning, zoom into corners, or read exact pixel values. Reviewing old frames doesn't count as a step — only real actions in the game do. No retraining, no new model weights. Just better eyes and a photo album.
The numbers
The results are startling. Claude Opus 5.0's Relative Human Action Efficiency score on ARC-AGI-3 jumped from 40.68 to a perfect 100.00. It cleared all 25 public games and 183 levels using 57.4% fewer actions than first-time human players. GPT-5.6 Sol scored 98.27 under the same harness. Even a 320-billion-parameter open-weight model that managed only 1.89 with the official interface was lifted to 66.93. And the perception bottleneck is real: simply swapping the numeric grid for screenshots took GPT-5.6 Sol from 13.33 to 47.32 with nothing else changed — while using fewer tokens (30.7 million per game versus 71.9 million for text).
Why it matters
ARC-AGI-3 is an interactive benchmark: agents enter unfamiliar game-like worlds with no instructions and must explore, infer hidden rules, and adapt. It's considered the toughest general-agentic-intelligence test still unsaturated. The VISTA result suggests the industry's obsession with bigger models and longer context windows may be aimed at the wrong bottleneck. Notably, more wasn't better here: stretching context from 200K to 780K tokens actually dropped the score from 99 to 93.9. The authors' conclusion is blunt — stop assuming the model lacks intelligence and start asking what it's missing from its eyes and memory.
The honest caveats
The team is upfront about one big asterisk: their models were trained after the public games were released, so they may have seen similar puzzles before. True generalization will be tested on ARC-AGI-3's private set. And this doesn't mean reasoning is “solved” — it means the interface matters as much as the brain.
AiPost's take
This is the second perfect ARC-AGI-3 score in two months — Nvidia's AVO system did it in August with a totally different architecture. The pattern is becoming hard to ignore: the system around the model is unlocking more capability than the model itself. If you're building AI agents, the lesson is practical and cheap — give your model screenshots and a memory before you pay for a bigger one.

Comments
Post a Comment