Discover the Best AI Tools

We test the top AI tools for writing, video, images & music — so you don't have to.

Explore AI Tools

Anthropic's AI Found 129,000 Software Bugs — Now It's Arming More Security Teams

Anthropic just made a big bet on AI-powered defense. It's expanding access to its most powerful AI models — with fewer safeguards — to more vetted cybersecurity teams. The reason: its AI already found over 129,000 real software vulnerabilities this year. What Project Glasswing uncovered Partners under Anthropic's Glasswing initiative found at least 129,000 verified software vulnerabilities between April and July 2026, more than 33,000 of them rated critical or high severity. Anthropic's own open-source scanning added 5,500 more between April and October. The company says the true number is likely at least five times higher, since the data covers only a limited group of partners. The revamped Cyber Verification Program Anthropic announced Tuesday a revamped Cyber Verification Program (CVP) combining two programs it ran for six months. Glasswing gave critical-software organizations access to Claude Mythos, its most cyber-capable model family, while the original CVP gave vette...

AI Wasn't Dumb — It Was Blind: MIT's VISTA Gave Claude Opus a Perfect ARC-AGI-3 Score

MIT VISTA illustration — an AI robot studying pixel-game puzzles through a magnifying glass, titled AI Wasn't Dumb — It Was Blind

 

For months, the story around AI reasoning benchmarks has been simple: models keep getting smarter, scores keep climbing. A new paper from MIT flips that story on its head. The models weren't dumb, the researchers argue — they were blind.




The experiment

On October 1, a team at MIT — Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, and Kaiming He — uploaded a paper to arXiv describing VISTA, a "visual harness" for AI agents. Instead of feeding Claude Opus 5.0 the official ARC-AGI-3 interface — a 64x64 grid of numbers representing each game state — they gave it plain 512x512 screenshots of the actual games, plus a lossless visual memory: every frame the game returns is archived in its original form, and the model can flip back through any of them mid-reasoning, zoom into corners, or read exact pixel values. Reviewing old frames doesn't count as a step — only real actions in the game do. No retraining, no new model weights. Just better eyes and a photo album.

The numbers

The results are startling. Claude Opus 5.0's Relative Human Action Efficiency score on ARC-AGI-3 jumped from 40.68 to a perfect 100.00. It cleared all 25 public games and 183 levels using 57.4% fewer actions than first-time human players. GPT-5.6 Sol scored 98.27 under the same harness. Even a 320-billion-parameter open-weight model that managed only 1.89 with the official interface was lifted to 66.93. And the perception bottleneck is real: simply swapping the numeric grid for screenshots took GPT-5.6 Sol from 13.33 to 47.32 with nothing else changed — while using fewer tokens (30.7 million per game versus 71.9 million for text).

Why it matters

ARC-AGI-3 is an interactive benchmark: agents enter unfamiliar game-like worlds with no instructions and must explore, infer hidden rules, and adapt. It's considered the toughest general-agentic-intelligence test still unsaturated. The VISTA result suggests the industry's obsession with bigger models and longer context windows may be aimed at the wrong bottleneck. Notably, more wasn't better here: stretching context from 200K to 780K tokens actually dropped the score from 99 to 93.9. The authors' conclusion is blunt — stop assuming the model lacks intelligence and start asking what it's missing from its eyes and memory.

The honest caveats

The team is upfront about one big asterisk: their models were trained after the public games were released, so they may have seen similar puzzles before. True generalization will be tested on ARC-AGI-3's private set. And this doesn't mean reasoning is “solved” — it means the interface matters as much as the brain.

AiPost's take

This is the second perfect ARC-AGI-3 score in two months — Nvidia's AVO system did it in August with a totally different architecture. The pattern is becoming hard to ignore: the system around the model is unlocking more capability than the model itself. If you're building AI agents, the lesson is practical and cheap — give your model screenshots and a memory before you pay for a bigger one.





Comments

Popular posts from this blog

Top 5 Best AI Tools in 2026 (I Tested Them All)

SpaceXAI's Grok 4.7 Is a Bigger, Smarter Model at the Same Price

YouTube Just Rolled Out a Batch of AI Tools for Creators — Here's What They Do