On September 30, 2026, Google announced Gemini 4 Argon, a frontier reasoning model whose headline change is an output limit of 1 million tokens, up from 64K, according to Google's announcement. It is the model our earlier post on Gemini 4 entering post-training was anticipating. Gemini 4 Argon is not yet generally available: it is rolling out first to cyber defenders through Google's Fairwind program.
Gemini
AnthropicGemini 4 Argon: what Google announced
Google says Argon is rolling out to cyber defenders via the Fairwind program, with broader availability planned for "developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers." A U.S. government pre-release access process is also under way. DataCamp adds that access is currently limited to trusted cyber defenders and Google internal teams, with no published model ID or free tier as of September 30.
Google also describes safety measures: internal activation monitoring for misuse, adversarial training against prompt injection, monitoring of chain-of-thought for misalignment, and sandboxed environments with agent control protocols. This staged, defender-first release follows the pattern we covered in Gemini 3.8 Flash Cyber, where cyber-capable models reach vetted users first.
Gemini 4 Argon pricing and the 1M output limit
Google lists introductory pricing of $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. Standard pricing afterward is $4 and $20. DataCamp notes the length of the introductory period has not been published.

The output limit changes the cost math more than the per-token rate does. By simple arithmetic on the list prices, one response that fills the full 1M-token limit costs about $10 at the introductory rate and $20 at the standard rate. A 64K-token response costs roughly $0.64 and $1.28. Cached input at the introductory rate works out to about $0.10 per million tokens. Whether long outputs are useful depends on the task: whole-codebase rewrites, long reports and bulk document generation fit, while most chat replies will never approach the limit.
Benchmarks: where Argon leads and where it trails
The figures below are launch-day numbers tabulated by OfficeChai and DataCamp. DataCamp counts Argon leading 13 of the 19 benchmarks Google published against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. OfficeChai counts 14, so the exact tally depends on how ties are scored. These are vendor-reported results, not independent tests.
| Benchmark | Gemini 4 Argon | Best listed rival |
|---|---|---|
| DeepSWE v1.1 | 77.9% | Opus 5.5 74.2%, Astra 74.1%, Fable 5.1 67.4% |
| GraphWalks (256K-1M) | 84.2% | Astra 71.8% |
| Vals Index | 68.9% | Opus 5.5 67.0% |
| AutomationBench | 51.3% | Opus 5.5 42.5% |
| Vals Finance Agent v2 | 65.4% | Fable 5.1 58.9%, Opus 5.5 58.6%, Astra 53.5% |
| OSWorld-2.0 | 69.2% | Astra 72.6% |
| Terminal-Bench 4.0 | 57.4% | Opus 5.5 66.4% |
| PostTrainBench | 45.3% | Opus 5.5 49.3% |
Gemini 4 Argon vs the best rival score, in points
Green bars are benchmarks where Argon is ahead of the best listed rival; gray bars are where it trails. Wins cluster in knowledge work, long context and science, losses in terminal and computer-use work.
Source: Google launch figures as tabulated by OfficeChai (Oct 1, 2026). A subset of the 19 benchmarks; vendor-reported, not independently verified.
The pattern is consistent. Argon leads on knowledge work, very long context and science questions, and trails on terminal-driven and computer-use tasks, one of the two software-engineering suites (FrontierSWE v2, 55.0% against Astra's 65.5%) and ML engineering. On CWE-bench v1, a vulnerability benchmark, it ties Astra at 68%, which fits the cyber-defender rollout.
A model can lead on a software-engineering benchmark and trail on another, so a single ranking hides more than it shows.
Reading of the launch tables above
How to read the results
- Source of figures. The scores are Google's, relayed by press. Independent evaluations usually follow within weeks and may shift the picture.
- Task shape matters. DeepSWE rewards repository-level fixes, while Terminal-Bench and OSWorld reward long sequences of tool actions. Leading one says little about the other.
- Long context is a real differentiator. An 84.2% GraphWalks score against 71.8% for Astra, at 256K to 1M tokens, is the widest gap among the software and context tests.
- Access is the constraint. Until paid API access opens, most teams cannot test Argon on their own workloads.
Choosing between frontier models
With Argon, GPT-6 Astra and Anthropic's Opus 5.5 each leading different rows, the practical answer is to match the model to the workload and re-test as releases land. That is the case for model-agnostic tooling, which is how teams on platforms like Metir tend to work: keep prompts and evals portable so a new model can be trialled without a rewrite.
Sources:
- Gemini 4 Argon - Google
- Gemini 4 Argon: Benchmarks, Pricing, and Access - DataCamp
- Gemini 4 Argon benchmarks - OfficeChai
- Google unveils Gemini 4 Argon with strong scores and limited access - The Rundown AI
Image credits
Header image: the Googleplex building in Mountain View, California, by Asoundd, Wikimedia Commons, CC BY-SA 4.0 (source); a file photo, not an image of the Gemini 4 Argon launch. In-body image: aerial view of the Google campus by Austin McKinley, Wikimedia Commons, CC BY 3.0 (source).
