← Back to blog

Gemini 4 Argon: Fairwind first, API key later

On September 30, 2026, Google DeepMind announced Gemini 4 Argon. The first people outside Google who can run it are trusted cyber defenders in the Fairwind Program. Developers, enterprises, and consumers are explicitly later.

That access order is the story. The 1 million output-token ceiling is the engineering change. The 77.9% DeepSWE number is Google’s own chart.

If you are not in Fairwind, you do not have a model decision yet. You have a pricing footnote and a wait.

What shipped

Koray Kavukcuoglu’s post calls Argon the opening model of the Gemini 4 generation. Google is expanding the output limit from 64K tokens to 1 million tokens in a single call, so a long agent trajectory can stay in one inference pass instead of being stitched across turns.

Introductory API pricing, when the API exists, is $2 per million input tokens and $10 per million output tokens. Cached input is 95% off the input price. A footnote says that after the introductory period the rate becomes $4 and $20. Google does not give the date of that step, or the date paid API customers and Google AI Ultra subscribers actually get the model.

Google says it is in the U.S. government’s voluntary pre-release access process and will iterate on guardrails before a wider release. For trusted defenders and Google’s own teams, the post says Argon will ship without cyber guardrails so those teams can use the full defensive capability. That is a scoped exception, not the public default.

SiliconANGLE reports that Fairwind, opened September 3 with Gemini 3.8 Flash Cyber, has signed up more than 650 organizations, including CrowdStrike and Palo Alto Networks. Google’s own post does not publish that count. Treat 650 as press reporting of the program, not as a number in the announcement.

The numbers Google is willing to print

On the announcement, Argon is “a new state of the art” on DeepSWE v1.1 at 77.9% for long-horizon software engineering. 9to5Google’s reading of Google’s comparison chart puts Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1% on that same test. Those are still Google’s bars until an independent lab runs the identical harness.

On CWE-bench v1, Google says Argon ties for first at 68% on vulnerability remediation. On Gray Swan’s indirect prompt-injection benchmark, Google says Argon is leading and does not print the attack-success rate in the post. On Zapier’s AutomationBench, Google claims first place at 51.3%. On LVBench, long-video understanding, Google claims 91.7%.

Google also says Argon leads the Vals Index and shows strong results on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. The post does not put the peer scores for those last two in the prose. Do not invent a sweep from a chart you cannot read as a table.

Internal proof is still Google grading Google

The operational claims are large, and they are internal:

  • Argon agents working on C/C++ to Rust migrations, up to 800,000-plus lines for the Fuchsia Zircon kernel, still going through automated and manual audit before production.
  • A Rust port of libgav1 whose SIMD path Argon reworked into a memory-safe decoder Google says runs 2.7× the prior Rust port, with identical video output, closer to the optimized C++.
  • Fleet memory work Google says freed more than 300 TiB once rolled out, with an estimate of 500 TiB to 1 PiB in total savings.
  • Wiz using Argon in Scan for Good, including a claim that it found a critical exposure in healthcare software that earlier frontier models missed.

Useful as evidence that Google is already running the model on its own estate. Not a customer reference you can call.

What to do on Monday

  • Fairwind partners: start designing agent loops against a 1M output ceiling. That changes handoff and truncation, independent of who wins DeepSWE.
  • Everyone else: do not budget the $2/$10 introductory rate. The published follow-on rate is $4/$20, and the switch date is missing.
  • Treat the unguardrailed build as a defender-program fact, not as the model you will get on the public API.
  • Wait for a non-Google DeepSWE run before you repeat 77.9% as a settled ranking.

Sources

← Back to blog