OpenAI Says It Caught a Massive AI Heist: 15,000 Accounts Were Trying to Steal Its Models' Hidden Reasoning
We test the top AI tools for writing, video, images & music — so you don't have to.
Explore AI Tools


Google has a new flagship AI model, and it goes by a new name: Argon. Announced on Wednesday, September 30, Argon anchors the Gemini 4 generation of models — a new top tier sitting above the old "Pro" line, built for what Google calls "deep reasoning across complex, long-horizon workflows," from coding and knowledge work to cybersecurity defense and creative writing. The naming scheme is new, and it arrives after Google cancelled Gemini 3.5 Pro to focus development on 3.8 Flash.
The biggest concrete change is output length: Argon can generate up to 1 million tokens in a single response, up from 64,000. That is the model built for long, multi-step work — entire codebases rewritten, sprawling documents digested, and extended agentic runs completed without stopping and restarting.
On Google's own reported numbers, Argon leads outright on 12 of 18 disclosed benchmarks and ties on one. The standouts: DeepSWE v1.1 at 77.9% — a state of the art, ahead of Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%); CWE-bench v1 at 68% (tied for first); Zapier's AutomationBench at 51.3% (#1); LVBench long-video understanding at 91.7% (state of the art); Harvey's Legal Agent Benchmark at 19.6% versus GPT-6 Astra's 5.4%; and Vals Finance Agent v2 at 65.4%.
But honesty demands the caveats: the race is still close. Argon trails GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and Claude Opus 5.5 on Terminal-bench 4.0 (57.4% vs 66.4%). These are Google-reported figures, and independent evaluations are still to come.
In an unusual move, Google is giving vetted cyber defenders first access through the Fairwind Program — "without cyber guardrails" — so they can use Argon's full cybersecurity capabilities. Google is also participating in the Trump administration's voluntary pre-release model access process. The approach already has a proof point: Wiz says it used Argon through its Scan for Good program to find a critical vulnerability in healthcare software that earlier frontier models missed.
Thousands of Google employees are already using Argon, and the internal numbers are striking — all reported by Google itself. Agentic runs freed over 300 TiB of data-center memory, moved more than 800,000 lines of Fuchsia's Zircon kernel from C/C++ to Rust, and improved a quantum computing algorithm by 40% beyond a published baseline within minutes.
Access is limited for now: cyber defenders first, then Google AI Ultra subscribers and paid API customers. Google has given no public timeline for broader release. Pricing is introductory — $2 per million input tokens and $10 per million output tokens (cached inputs 95% off) — rising to $4/$20 after the introductory period.
Google spent much of this year playing catch-up after a blistering DevDay season from OpenAI (Dots, GPT-6.1 Sol). Argon is its answer — and the defenders-first release is the real signal. By handing its most powerful model to cybersecurity first, Google is saying AI's next battleground isn't chatbots. It's the security of the software the world runs on.
Comments
Post a Comment