AI Frontier Post
AI News

Gemini 4 Argon: Google's new frontier model tops the coding benchmarks — and goes to cyber defenders first

On Wednesday, Google announced Gemini 4 Argon, its new frontier AI model built for deep reasoning across long, complex workflows — real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. It beats OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 on the headline coding benchmark. But you can't use it yet: Argon is going first to a small group of trusted cyber defenders.

A glowing spark motif rising from a cyber shield against a dark blue background
Google's Gemini 4 Argon: frontier performance on coding, knowledge work, and cybersecurity defense — cyber defenders get access first. AI-generated illustration for AI Frontier Post.

Google's frontier-model counterattack is here. Announcing Gemini 4 Argon on Wednesday, the company called it a new era of frontier intelligence — a model built to sustain deep reasoning across complex, long-horizon workflows without losing track of earlier steps. And Google brought numbers: Argon takes the top score on the long-horizon coding benchmark DeepSWE v1.1, ranks first on Zapier's AutomationBench for end-to-end business execution, and sets the state of the art on LVBench for long-video understanding.

The headline, though, is who gets it first: not developers, not ChatGPT-style consumers — cyber defenders. Google is rolling Argon out initially through its Fairwind Program, a restricted-access program for trusted defenders including governments and critical-infrastructure operators, while it participates in the U.S. government's voluntary pre-release model-access process. Broader access will start with paid API customers and Google AI Ultra subscribers, 9to5Google reports — with no public timetable attached.

The benchmarks#

Google says Argon scores 77.9% on DeepSWE v1.1, a benchmark measuring performance on real-world, long-horizon software-engineering tasks — ahead of Claude Opus 5.5 at 74.2% and OpenAI's GPT-6 Astra at 74.1%, per 9to5Google's read of the announcement. Beyond code, Argon leads the Vals Index, which weights economic impact across finance, coding, legal, and tax work by their contribution to U.S. GDP; it also leads on Vals Finance Agent v2 (multi-step financial research) and Harvey's Legal Agent Benchmark (legal research and drafting).

On Zapier's AutomationBench, which tests end-to-end execution across core business functions, Argon ranks #1 with 51.3%. And on LVBench, measuring long-video understanding, it reaches 91.7% — state of the art. The underlying theme, Google says, is stamina: Argon is designed to follow much larger and more complicated problems through many steps without dropping the thread.

A holographic security shield projected over a nighttime AI operations center
Cyber defenders get Argon first: trusted partners in the Fairwind Program receive it without cyber guardrails. AI-generated illustration for AI Frontier Post.

Cyber defenders get it first#

The cybersecurity framing is deliberate — and unusually candid. Google trained Argon to be highly capable at cybersecurity defense: it can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and Google's own internal teams, the company says it will release the model without cyber guardrails so they can use its full frontier-level defensive capabilities.

There is already a proof point. Cloud-security company Wiz is using Argon through its Scan for Good initiative, which protects critical public infrastructure for free. In an early demonstration, Google says Argon uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide — a risk previous frontier models had missed.

On CWE-bench v1, which evaluates a model's ability to remediate security vulnerabilities, Argon ties for first place with a 68% score. On Google's internal vulnerability benchmark, it uncovered exposures across complex codebases spanning 20 programming languages; on Wiz's black-box penetration-testing benchmark, it outperformed the previous generation (3.8 Flash Cyber) at discovering attack surfaces and producing proof-of-concept evidence.

1 million tokens of headroom#

To support those long trajectories, Google is expanding Argon's output token limit to 1 million tokens — up from the previous 64K. Google's pitch: when the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning and can solve tough problems in one go.

Google is also reporting serious internal returns. Thousands of Googlers are already using Argon, and the company says it beat a published baseline by 40% in quantum algorithmic optimization within minutes; a team of Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations across Google's data centers, freeing up over 300 TiB of memory, with an estimated 500 TiB to 1 PiB in total savings; and Argon agents are working on migrating C/C++ codebases to Rust — scaling up to 800K+ lines for the Fuchsia Zircon kernel, with rigorous auditing before anything reaches production.

An endless luminous ribbon of reasoning steps stretching to the horizon
The stamina pitch: Argon's 1-million-token output limit is designed to let a single reasoning trajectory run hundreds of thousands of tokens. AI-generated illustration for AI Frontier Post.

The phased rollout — and the price#

Google says safely releasing frontier capabilities at this level requires a phased approach, and it's being explicit about what that means. The model is engaged in the U.S. government's voluntary pre-release process, and Google says it will keep gathering feedback from early testers as it iterates on guardrails. The safeguards work covers four areas, The Deep View reports: refusing harmful requests (including CBRN misuse) while preserving legitimate dual-use research; prompt-injection robustness, where Google claims a leading score on Gray Swan's Indirect Prompt Injection benchmark; monitoring of Argon's chain-of-thought for misalignment; and hardened, sealed sandbox environments for high-risk training and evaluation.

On pricing, Google is going aggressive on launch: an introductory $2 per million input tokens and $10 per million output tokens, with cached input at 95% off — matching OpenAI's GPT-6.1 Sol rate card almost exactly, Unite.AI notes. That's a clear signal about where Google wants to compete: not just on benchmarks, but on the economics of long-running agentic work.

The open question is timing. Google says Argon will be available to developers, enterprises, and consumers "as soon as possible" — but the Fairwind-first strategy, the government pre-release process, and the lack of any date make it clear that "announced" and "shippable" are two very different words this time. For now, Argon exists, it scores, and the cyber defenders have it. Everyone else is on the waitlist.