Quick Answer:
Gemini 4 Argon is Google's most powerful model to date, announced on 30 September 2026. It raises the output limit to 1 million tokens, leads Google's published comparison on knowledge work, agentic coding and long-context tests (for example 77.9% on DeepSWE v1.1), and launches first to trusted cyber defenders through the Fairwind Programme. Introductory API pricing is $2 input / $10 output per million tokens. Every number below is vendor-reported; independent replication is still to come.
For most of 2026 the story of Google's flagship model has been a story of waiting. Gemini 3.8 shipped, the rumour mill spun, and the loudest frontier launches came from Anthropic and OpenAI. On 30 September, Google finally put a number on the board: Gemini 4 Argon.
The headline claims are bold. Google says Argon beats GPT-6 Astra and Anthropic's Claude Fable 5.1 and Opus 5.5 across most of the benchmarks it chose to publish, can autonomously find and patch serious software vulnerabilities, and can write a million tokens in a single answer. This review goes back to Google's own announcement, reads the full comparison table line by line, and separates what the data supports from what the marketing implies.
AI Revolution X summarises the Gemini 4 Argon launch, its cyber capabilities and the headline benchmark claims.
Executive Summary
What happened: on 30 September 2026 Google DeepMind announced Gemini 4 Argon, described by Google as its most powerful model yet and, in the words of its announcement, "built to sustain deep reasoning across complex, long-horizon workflows". The launch was covered the same day by TechCrunch, and Google's own post is on the Google blog.
- Output limit: up to 1 million tokens in a single response, up from 64,000.
- Coding: 77.9% on DeepSWE v1.1, ahead of GPT-6 Astra (74.1%), Claude Opus 5.5 (74.2%) and Claude Fable 5.1 (67.4%) in Google's table.
- Knowledge work: 68.9% on the Vals Index and 51.3% on AutomationBench, both first place in Google's comparison.
- Cyber: tied for first on CWE-bench v1 at 68%, and released to trusted defenders first through the Fairwind Programme.
- Prompt-injection resistance: a 0.7% attack success rate at 15 attempts on Gray Swan's IPI benchmark, the lowest on Google's chart.
- Price: $2 per million input tokens and $10 per million output tokens during an introductory period, rising to $4 and $20.
Our view: this is a genuine frontier release and the strongest evidence yet that Google has closed the gap. It is not a clean sweep. Argon loses to GPT-6 Astra on FrontierSWE v2, Terminal-Bench Science and OSWorld-2.0, and to Claude Opus 5.5 on Terminal-bench 4.0 and PostTrainBench, and every figure comes from Google. Treat the table as a strong claim to be tested, not a verdict.
Why Argon Arrived Late
Gemini 4 was one of the most anticipated and most delayed launches of the year. We tracked the speculation in our earlier piece on the Gemini 4 release date and expected specs, and the question of timing dominated coverage for months. The AI Revolution X video above cites Axios as describing the launch as arriving after an unusually long flagship delay. We could not retrieve the Axios article itself, so we treat that framing as secondary reporting.
What Google's own announcement does tell us is how it wants Argon positioned. The model is not pitched as a chatbot upgrade. It is pitched as a worker: something that can run for a long time, hold a large problem in its head and produce a very large amount of output without a human restarting it. That positioning matters because it follows the direction the whole industry has taken since the autumn, from single answers to long-running agents. If you want the competitive backdrop, our guide to the best AI models right now and the frontier AI benchmarks roundup for September 2026 show where each lab stood the week before Argon landed.
There is also a release-process point worth noting. Google says it is engaged with the US government's voluntary pre-release model access process while it expands availability. In practice that means the most capable version of the model reaches a small, vetted group first, and wider access follows. We saw a similar staged pattern from OpenAI with its Astra line, which we covered in our GPT-6 Astra review and in the piece on Astra and the critical cyber threshold.
Capabilities Deep Dive
1 million tokens of output
The single most concrete technical change is the output ceiling. Google says Argon can now generate up to 1 million tokens in a single response, up from 64,000. Its explanation is that a model that can "think deeply and generate hundreds of thousands of tokens in a single trajectory" gains "a new level of depth in reasoning to solve tough problems in one go". That is a statement about reasoning budget as much as output length: the model has room to work through a problem step by step without being forced to compress its thinking.
Google's announcement does not state the input context window, so we will not either. What it does show, through the GraphWalks long-context test in the comparison table, is retrieval-style reasoning over very long inputs, which we cover in the benchmarks section.
Long-horizon agentic work
Google describes Argon as built for deep reasoning across long-horizon workflows, and the benchmark selection reflects that: DeepSWE for long-running software engineering, AutomationBench for business workflow automation, Vals Finance Agent for multi-step financial research, and Harvey's Legal Agent Benchmark for legal research and drafting. In each of those, Argon is first in Google's comparison. These are the kinds of tasks where an agent must plan, use tools, recover from mistakes and finish, not simply answer.
Multimodal understanding
Argon posts 91.7% on LVBench, a long-video understanding test, against 87.5% for GPT-6 Astra, 83.7% for Claude Opus 5.5 and 79.7% for Claude Fable 5.1. On Chartography, a chart-reading benchmark, it scores 71.6% to Astra's 71.0%, a very narrow margin. The video result is the clearer win and plays to a traditional Gemini strength.
Computer use
Computer use is where the picture is mixed. On Agent's Last Exam, measured as a pass rate, Argon scores 39.5% against 34.2% for Astra and 38.2% for Opus 5.5. On the offline subset of OSWorld-2.0, using a partial score, Argon scores 69.2% and GPT-6 Astra scores 72.6%. The table shows no figures for the Claude models on OSWorld-2.0, or for Fable 5.1 on Agent's Last Exam, so we cannot compare them there.
Wes Roth asks how much the early Gemini 4 Argon results really tell us, alongside OpenAI's DevDay announcements.
Benchmarks: What Google Published
Google published a single comparison table against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, with methodology documented on DeepMind's evaluations page. We reproduce Google's image below exactly as published, and then summarise it in plain text so the numbers are accessible to search engines and screen readers.

| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| Harvey's Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% |
| RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% |
| GraphWalks, up to 128k (BFS F1) | 99.7% | 98.7% | 91.4% | 90.6% |
| GraphWalks, 256k to 1M (BFS F1) | 84.2% | 71.8% | 65.0% | 66.8% |
| Agent's Last Exam (pass rate) | 39.5% | 34.2% | not shown | 38.2% |
| OSWorld-2.0 (offline subset, partial score) | 69.2% | 72.6% | not shown | not shown |
| Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Where Argon clearly wins
Count the rows: of the 19 benchmark lines in Google's table, Argon is first or tied for first on 14. The widest gaps are in the work-automation group. On AutomationBench it leads Opus 5.5 by 8.8 points and Astra by 9.9. On Vals Finance Agent v2 it leads the next best (Fable 5.1, 58.9%) by 6.5 points. On Harvey's Legal Agent Benchmark it scores 19.6% against single-digit scores for everyone else, though the absolute level shows how hard that test remains: even the leader fails four in five tasks.

Coding is the other headline. DeepSWE v1.1 measures long-horizon software engineering, and Argon's 77.9% is 3.7 points above Opus 5.5 and 3.8 above Astra. Vibe Code Bench is closer: 91.9% against 90.3% for both Claude models.

Where Argon does not win
Five lines in the table are not Argon's. GPT-6 Astra is ahead on FrontierSWE v2 (65.5% vs 55.0%, a 10.5-point gap), Terminal-Bench Science 0.1 (68.1% vs 57.6%) and OSWorld-2.0 (72.6% vs 69.2%). Claude Opus 5.5 is ahead on Terminal-bench 4.0 (66.4% vs 57.4%) and PostTrainBench (49.3% vs 45.3%). Terminal-bench is a command-line agent test, so if your work is shell-heavy, Opus 5.5 still has a credible claim. CWE-bench v1 is a three-way tie at 68.0% between Argon and Astra, with Opus 5.5 one point behind.
Honest caveats
- Vendor-reported. Google chose the benchmarks, the comparison models and the settings. Independent replication usually narrows leads.
- Harness differences. Google's CWE-bench chart notes that different models were run through different agent harnesses (for example Argon through Antigravity, Astra through Codex and Opus 5.5 through Claude Code), so scores compare model-plus-harness, not models alone.
- Missing comparisons. Not every model has a score on every row, which limits like-for-like reading on computer use.
- Saturation and novelty. Several of these benchmarks are new, so their difficulty and reliability are less established than older tests. We have explained how to read these tests in our September 2026 frontier benchmarks guide.
Cybersecurity and the Fairwind Programme
The most striking part of the launch is not the leaderboard but the release strategy. Google says Argon can autonomously find, validate and patch critical software vulnerabilities. Rather than shipping that capability to everyone, it is rolling the model out first to "a set of trusted cyber defenders through our Fairwind Program", and, per the announcement, released to them without the cyber guardrails that a general-purpose model would carry. The logic is dual-use: the same skill that lets a defender patch a flaw lets an attacker exploit it, so access is gated by who you are rather than only by what you ask.

The CWE-bench chart is more nuanced than the headline suggests. Argon is tied at 68% with GPT-6 Astra and with Grok 4.7, and Claude Opus 5.5 sits one point behind at 67%. Google describes this as tied for first rather than a clear win, which is the honest reading. Further down the chart, Claude Fable 5.1 scores 58%, Muse Spark 1.3 and DeepSeek-V4.1-Flash 55%, and GPT-6 Sol 52%.
Wiz and Scan for Good
Google names Wiz as a launch partner. Wiz is using Argon in its Scan for Good initiative, which offers free vulnerability detection for critical public infrastructure. According to Google, the model uncovered a critical vulnerability in healthcare software that exposed sensitive personal information and that previous frontier models had missed. That is a single anecdote, and Google does not publish the number of findings, false positives or the software involved, so treat it as illustrative rather than statistical.
Google also says Argon beats its own Gemini 3.8 Flash Cyber model on an internal vulnerability benchmark covering 20 programming languages and on Wiz's black-box penetration-testing benchmark. We did not find absolute scores for those tests in the text, so we do not quote any.
What defenders and security teams should do
- If you run critical infrastructure or maintain widely used open-source code, ask Google about Fairwind eligibility rather than waiting for general availability.
- Assume attackers will have comparable models within months; patch-cycle speed matters more than any single tool.
- Keep human review of every automated patch. A model that writes fixes can also introduce regressions.
- Read our coverage of how OpenAI handled the same dilemma in the Astra critical-threshold decision.
Safety, Prompt Injection and Misalignment
Google describes four layers of protection under its Frontier Safety Framework: misuse defence, prompt-injection robustness, misalignment monitoring and hardened sandboxes. The misuse layer is designed to refuse harmful requests while preserving legitimate dual-use research, and includes internal activation monitoring and manual and automated red-teaming. The misalignment layer monitors the model's chain of thought. The sandbox layer isolates the environments in which the model runs code.

The prompt-injection result is the most useful safety number for anyone building agents. On Gray Swan's IPI benchmark, which measures how often an attacker succeeds in hijacking a model through instructions hidden in content it reads, Argon's attack success rate at 15 attempts is 0.7%. Claude Opus 5.5 and Claude Fable 5.1 are both at 1.0%, GPT-6 Astra is at 8.5% and GPT-6 Sol at 10.1%. The gap between the top group and the rest is large, and for agent builders it matters more than a point of coding accuracy, because an agent with email, browser or code access is only as safe as its resistance to injected instructions.
Two cautions. First, a 0.7% rate is still not zero, and in a deployed agent with high-value permissions, even a small rate across thousands of tasks is a real risk. Second, this is Google's chart, using Gray Swan's benchmark; we recommend running your own injection tests against your own tools. Our guide to AI browser automation agents covers practical defences.
WorldofAI runs early hands-on tests of Gemini 4 Argon across coding and general tasks.
Real-World Use at Google
Benchmarks are a proxy. Google also offers three internal examples of Argon at work for thousands of its staff, which we report as claims rather than verified results:
- Quantum algorithm optimisation: Google says Argon beat published baselines by 40%.
- Memory efficiency: a change that freed 300 TiB of memory, with total estimated savings of 500 TiB to 1 PiB.
- Language migration: large-scale conversion of C and C++ code to Rust across more than 800,000 lines.
The migration example is the most transferable. Moving a large legacy codebase to a memory-safe language is exactly the kind of long, repetitive, verification-heavy job that a million-token output window and strong long-horizon reasoning should help with. If those numbers hold up externally, they point to a future in which refactoring projects that once took teams of engineers a year become supervised projects that take weeks. Our Google Antigravity plan review covers the agentic coding environment Google lists beside Argon on its CWE-bench chart, shows how the surrounding tooling is evolving.
Pricing and Availability
| Tier | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Introductory | approx. £1.50 ($2) | approx. £7.50 ($10) |
| Standard (after introductory period) | approx. £3.00 ($4) | approx. £15.00 ($20) |
| Cached input | 95% discount on the input price | |
Pound figures are our approximate conversions; Google quotes prices in dollars. Availability today is limited to the Fairwind cyber-defender group. Google lists paid API customers and Google AI Ultra subscribers as the next audiences, without giving dates. TechCrunch's report did not include pricing, so the price details come from Google's own post.
The 1M output limit has a cost implication that the headline price hides: a single maximal response would cost roughly $10 at the introductory output rate (about £7.50), and $20 at the standard rate. Long agent runs that call the model repeatedly will add up, so use the 95% cached-input discount where you can. See our AI API pricing comparison for how this lines up with rival models.
How It Compares
Against GPT-6 Astra: Argon leads on most knowledge-work, coding and long-context lines, but Astra is clearly stronger on FrontierSWE v2, Terminal-Bench Science and OSWorld-2.0, so for deep scientific or desktop-automation tasks it remains a serious option. Read our GPT-6 Astra review for the OpenAI view.
Against Claude Opus 5.5: Argon is first or tied on 14 of 19 lines overall, but Opus 5.5 leads on Terminal-bench 4.0 and PostTrainBench and is neck and neck on prompt-injection resistance (1.0% vs 0.7%). Our Claude Opus 5.5 review and the head-to-head in Claude Opus 5.5 vs GPT-6 Astra give the background.
Against Claude Fable 5.1: Fable 5.1 trails Argon on almost every line in Google's table. See our Fable 5.1 review for the strengths that a single comparison table does not capture.
Inside Google's own stack: Google has not said whether Argon will power its other agentic products, such as Gemini Spark, Chrome Conductor or the Workspace agentic features, or when.
Limitations and Caveats
- Access is narrow. Today only vetted cyber defenders get the model. Most developers cannot yet test any claim here.
- All numbers are Google's. No independent leaderboard has yet confirmed them. Wait for third-party evaluations.
- Input context is unstated. The announcement stresses a 1M output limit but we found no stated input window.
- No system card figures reviewed. We did not find a separate system card with detailed safety evaluations; our safety section reflects the blog post only.
- Pricing will change. The $2/$10 rate is introductory and doubles.
- Cost of long outputs. Million-token responses can be expensive and slow, and quality over such lengths needs real-world verification.
- Absolute scores remain low on hard tests. 19.6% on Harvey's legal benchmark and 39.5% on Agent's Last Exam mean human oversight is still essential.
Who Should Use It
- Security teams and critical-infrastructure operators: apply for Fairwind access and run Argon against your own code under supervision.
- Engineering teams planning large migrations: prepare a pilot for when API access opens, especially for C and C++ to Rust work.
- Agent builders: prioritise Argon in your evaluation set because of its prompt-injection result, but run your own attacks.
- Finance, legal and operations teams: the Vals and Harvey results suggest it is worth benchmarking, but keep a human reviewer in the loop.
- Everyone else: wait for general availability and independent tests before switching a production workflow.
The Bottom Line
Gemini 4 Argon is the first Google model in a long while that can credibly claim to lead the frontier on a broad set of tests. The case is strongest in long-horizon coding, enterprise workflow automation, long-context reasoning, video understanding and prompt-injection resistance. It is weaker, in Google's own data, on command-line agent work, some science tasks and desktop automation.
The careful position is to take the launch seriously without taking the leaderboard as settled. The staged Fairwind release is a sensible response to a model with strong cyber capability, and it signals that the next phase of competition will be about trust and access as much as raw scores. We will update this article when API access opens, when independent benchmarks land and when Google publishes further safety documentation.
Last updated: 02/10/2026. Sourced from Google's Gemini 4 Argon announcement (30/09/2026), Google's published benchmark charts and TechCrunch. All benchmark figures are vendor-reported.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






