Google DeepMind launches Gemini 4 Argon: 1M-token outputs, cyber defenders first, and half of Claude's price
Google DeepMind's first Gemini 4 model can generate up to a million tokens in a single response, is rolling out first to cyber defenders through the Fairwind Program, and undercuts its rivals on price. The benchmarks look strong — with a few conspicuous exceptions.

Google DeepMind has officially entered the Gemini 4 era. On September 30, the lab announced Gemini 4 Argon, its new frontier model, aimed at the hardest class of work AI is asked to do today: long, multi-step problems that require sustained reasoning rather than a single clever answer. The announcement, written by SVP of Google DeepMind and Chief AI Architect Koray Kavukcuoglu, frames Argon as the first model of the Gemini 4 generation, built from the ground up for software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense.
The headline specification is the output window: up to one million tokens generated in a single response, up from 64K on the previous Gemini generation. There is a catch, though. Almost nobody can use it yet. Argon is rolling out first to a vetted group of cyber defenders through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers queued behind them — and no date announced for general availability.
A million tokens in one shot#
Output length is an underappreciated frontier. Most people focus on the input context window — how much a model can read — but for agents doing real work, the binding constraint is often how much a model can write before its thought collapses. Today's frontier APIs all cap a single response far lower: Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra each top out at 128K output tokens. Argon's 1M limit is roughly eight times that.
Google's argument is that a model with room to generate hundreds of thousands of tokens in one trajectory can reason more deeply — refactoring a large codebase, drafting a comprehensive legal brief, or auditing a sprawling system without splitting the job across turns and losing the thread. The cost is real: a full million output tokens runs $10 at introductory pricing, $20 after. Cached input, meanwhile, gets a 95% discount. Notably, Google has not disclosed Argon's input context window.
Cyber defenders get it first#
The strangest part of the launch is who is excluded from it. Developers, businesses, and consumers all wait — but cybersecurity teams get priority. Google trained Argon to autonomously find, validate, and patch critical software vulnerabilities, and trusted defenders (plus internal Google teams) are receiving the model without its cyber guardrails, so they can use its full defensive capability.
The first public proof point comes from Wiz, which is using Argon through its Scan for Good initiative. According to Google, the model found a critical vulnerability in healthcare software used by hospitals worldwide — one that previous frontier models had missed. Google is also participating in the U.S. government's voluntary pre-release model-access process, and says it will tighten safeguards in four areas — misuse defenses, prompt-injection resistance, chain-of-thought monitoring, and sealed training sandboxes — before the broader release. The Fairwind-first rollout looks like a bet that the first flagship models to matter in practice will be the ones fighting off attacks, not the ones answering questions.

The benchmark mix: strong, with gaps#
Google compared Argon against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1, and says it leads outright on 12 of 18 benchmarks, tying for first on one more. The headline numbers: 77.9% on DeepSWE v1.1, a long-horizon software-engineering test (Opus 5.5 scored 74.2%, GPT-6 Astra 74.1%); first place on the Vals Index at 68.9%; 51.3% on Zapier's AutomationBench for end-to-end business execution; 19.6% on the Harvey legal-agent benchmark against 5.4% for GPT-6 Astra; and a new state of the art of 91.7% on LVBench, long-video understanding. On vulnerability remediation it ties for first at 68% on CWE-bench v1. Independent tester Artificial Analysis scores Argon 53 on its Intelligence Index — tied with GPT-6 Astra for the top spot.
But the sweep is not clean. Argon trails GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%), lags Claude Opus 5.5 on Terminal-Bench 4.0 (57.4% vs 66.4%), and loses to Astra on computer-use test OSWorld-2.0 (69.2% vs 72.6%). And there is an early credibility cloud: Bloomberg reported that some Google employees with direct access to the model say it performs worse on real coding tasks than the benchmarks suggest — a characterization Google told Bloomberg is inaccurate. As always, treat vendor-run benchmark tables as the opening bid, not the verdict.

Half the price of Claude#
Argon launches at an introductory $2 per million input tokens and $10 per million output, confirmed by Logan Kilpatrick — half of Claude Opus 5.5's rates, and roughly a fifth of GPT-6 Astra's. After the introductory window, pricing rises to $4 input and $20 output. Aggressive introductory pricing on frontier models is becoming the standard playbook: buy market share and developer lock-in before the window closes, then normalize. The question for anyone building on Argon is when that window shuts — Google has not said how long the introductory rates last.
What to watch#
First, the access timeline. The gap between the Fairwind launch and paid API availability will tell you how seriously Google takes its own safety checklist — and how much revenue it is willing to defer to get it right. Second, whether the benchmark gaps close: FrontierSWE and Terminal-Bench are exactly the tasks Argon is marketed for, and trailing there invites skepticism. Third, the internal-use evidence. Google says thousands of its own engineers are already using Argon, citing memory optimizations that freed over 300 TiB in its data centers, a 32,000-line SIMD-code replacement in the libgav1 Rust port that made the decoder 2.7x faster, and an 800,000-line migration of the Fuchsia Zircon kernel from C/C++ to Rust. If those are real, they are a stronger signal than any benchmark table.
And a small victory lap for continuity: this is the launch that DeepMind's chief promised just days ago, when Gemini 4 entered post-training. The timeline held. What remains to be seen is whether the model lives up to it.
Sources
- MarkTechPost — “Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense” (September 30, 2026)
- AI Weekly — “Google's Gemini 4 Argon Rolls Out to Cyber Defenders First” (September 30, 2026)
- Digest AI — “Google releases Gemini 4 Argon with 1M token output limit” (September 30, 2026)