Google's Gemini 4 Argon stretches a single trajectory to 1 million output tokens — 16x the old 64K ceiling — scores 77.9% on DeepSWE v1.1, and starts at $2 per million input tokens, but only vetted cyber defenders get it before the public does.
The most capable model Google has built this year is the one you are not allowed to use yet.
On 30 September 2026, Google announced Gemini 4 Argon, the first frontier model of the Gemini 4 generation — confirmed the same evening by Google DeepMind — and then handed it only to a curated group of cyber defenders instead of to developers, enterprises or consumers. The rollout runs through Google DeepMind's Fairwind Program, which already covers more than 650 partner organisations globally and gates access behind due-diligence checks, phishing-resistant MFA and a ban on redistributing model access.
The headline is endurance, not just accuracy. Argon raises the model's output limit to one million tokens in a single trajectory, up from 64,000 in the previous generation. Google's framing is that when a model is allowed to think and generate for hundreds of thousands of tokens in one go, it can hold a hard problem open rather than running out of runway mid-solution. That matters more than it sounds: long-horizon agent work — a migration, an audit, a multi-document investigation — fails most often because the run truncates, not because the model got the answer wrong.
The scores Google published are leading where Google says they count. Argon sets a new state of the art on DeepSWE v1.1 at 77.9%, a benchmark built around real-world, long-horizon software engineering inside actual codebases. It ranks first on Zapier's AutomationBench, which measures end-to-end execution of business workflows across core functions, at 51.3%. It leads the Vals Index, a benchmark that weights finance, coding, legal and tax work by each sector's contribution to US GDP, and it posts 91.7% on LVBench for long-video understanding. On CWE-bench v1, which tests whether a model can actually remediate known security vulnerabilities, it ties for first at 68%.
On the security side, the model ships unlocked — to a short list. Google says trusted defenders and its own internal teams will receive Argon with its cyber guardrails removed so they can use its full vulnerability-finding capability. Wiz is already running it through Scan for Good, and in an early demonstration the model surfaced a critical flaw exposing personal data in healthcare software used by hospitals worldwide — a bug Google says earlier frontier models had walked past. Fairwind partners are restricted to defensive and research purposes, with background checks on the organisations themselves.
Inside Google, the productivity claims are concrete. Google reports Argon agents freeing more than 300 TiB of data-centre memory once fully rolled out, beating a published quantum-subroutine baseline by 40% in minutes, and driving C-to-Rust migrations that reach over 800,000 lines in the Fuchsia Zircon kernel. On libgav1, its video-decoding library, agents replaced 32,000 lines of SIMD code and delivered a memory-safe decoder running 2.7x faster than the hand-written Rust port it started from.
Pricing and the safety sequence. Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, with cached input at 95% off; once the intro period ends, the price becomes $4 and $20. Google describes the release as deliberately phased: it is participating in the US government's voluntary pre-release access process, is hardening safeguards across four fronts — misuse, indirect prompt injection, misalignment monitoring on the model's chain of thought, and sandbox isolation — and will widen access to paid API customers and Google AI Ultra subscribers first, then to developers, enterprises and consumers. No date has been given for that last step.
A model you cannot access is not a tool, it is a forecast. We build AI-assisted workflows — document pipelines, support automation, internal assistants — and our job is to keep those systems on models that are actually available, actually priced for your volume, and actually verified against your own task set before anything ships. When Argon opens up, we will benchmark it against what you run today, tell you honestly whether the jump is worth the token price, and wire the migration with usage dashboards so the claimed saving shows up as a number you can check at month end.
If you are already paying for an AI model in production, the useful hour this week is deciding what you would switch to when the next frontier model lands. Talk to us and we will tell you straight whether the answer is "upgrade", "wait", or "your current model is fine — fix the workflow instead".
Sources
Send us your project outline or chat directly on WhatsApp.