Quick Answer:
Gemini 4 Carbon is an unreleased internal Google checkpoint. Business Insider reports that employees testing it on an internal coding platform rate it as comparable to Claude Opus 5.5 for coding, while still asking for more testing. Carbon is separate from Gemini 4 Argon, which Google announced on 30/09/2026 and has so far released only to selected cyber defenders. There are no published benchmarks, price or release date for Carbon, and the "recursive self-improvement" talk around it is unverified speculation.
Google has not even finished rolling out its newest flagship, and the leaks are already about the one after it. Gemini 4 Argon is still limited to trusted cyber teams, yet reporting says staff are testing a checkpoint called Carbon that they compare to the best coding model on the market.
This guide separates the three layers of the story: what Google has announced about Argon, what is reported about Carbon from a single Business Insider piece, and what creators and commentators are adding on top. We have not tested any of these models, and where a claim rests on anonymous sources we say so.
AI Revolution X walks through the Carbon report, Argon in Antigravity and the recursive self-improvement rumours.
Executive Summary
- The report: Business Insider says it saw documents and screenshots showing Google employees testing several Gemini 4 checkpoints, including Carbon, which was made available on Google's internal coding platform, reported as Jetski. One employee said Carbon matches Anthropic's Opus 5.5 for coding, but wanted more testing.
- Argon's status: announced 30/09/2026, with access limited to selected cybersecurity teams. Paying API customers and Google AI Ultra subscribers are next, with no date.
- Barium-B: TestingCatalog's summary of the report says Barium-B is the checkpoint selected for the public Argon release. We could not confirm this in a second source.
- Internal doubts: Bloomberg separately reported employee scepticism about Argon's real-world coding. Google pushed back and pointed to comments from DeepMind's Koray Kavukcuoglu.
- Antigravity hints: interface strings spotted in an Insiders build list Argon with 256K, 512K and 900K context options.
- Unverified: that Carbon beats Argon, that it is a big jump, that it signals recursive self-improvement, or that it will ship under that name.
Our view: Carbon is a real codename attached to a thin but credible report. It matters mainly as a signal that Google is iterating on coding quality, the area where Argon's reception has been mixed, ahead of a competitor release cycle led by Claude Opus 5.5.
What We Actually Know About Carbon
Everything public about Carbon traces back to one Business Insider story from 09/10/2026, summarised by outlets such as TestingCatalog, dnyuz and Placera. Business Insider reports that it reviewed documents and screenshots showing employees testing internal Gemini 4 versions, and that Carbon was made available on an internal coding platform. TestingCatalog describes the platform as Jetski, Google's internal name associated with Antigravity.
The key quote is secondhand and anonymous: one employee said Carbon is comparable to Opus 5.5 for coding, while saying they wanted to test it further. That is a feeling from one tester, not a benchmark. It does not say Carbon is better than Opus 5.5, and it does not say on which tasks or how it was judged. Early feedback on Argon checkpoints was described as generally positive, with one employee saying they lagged on some coding tasks.
What we do not know is longer than what we do. There is no parameter count, no context length, no pricing, no safety evaluation and no release date. It is unclear whether Carbon would ship as an Argon update or as a separate Gemini 4 release. Even the name is unstable: internal codenames change, and Google's previous cycle saw a planned Gemini 3.5 Pro never ship while cheaper Flash models arrived instead, as covered in our Gemini 3.5 Pro write-up.
One further claim is worth handling with care. TestingCatalog reports that the Business Insider piece says Barium-B is the checkpoint chosen for the public Argon release. If true, Argon as shipped is not the strongest internal checkpoint, which would explain why some employees see Carbon as ahead. It would also mean that Google is shipping a model it considers safe or stable before one it considers stronger, a pattern consistent with its limited, cyber-first rollout.
Gemini 4 Argon: The Model Carbon Would Follow
To judge Carbon you have to understand what Argon already is. Google announced Gemini 4 Argon on 30/09/2026 as its first flagship in almost a year. Koray Kavukcuoglu, Google's chief AI architect, said it is "fundamentally changing the way we work and build at Google", and Google says thousands of employees already use it internally for coding, research and writing.
- Output limit: up to 1 million tokens per response, up from 64,000.
- Introductory pricing: $2 per million input tokens and $10 per million output tokens, which Google positions as half the price of Claude Opus 5.5.
- Access: selected cybersecurity teams through Google's Fairwind programme, released without the usual cyber guardrails. Wiz used it to find a critical vulnerability in hospital software that earlier models missed.
- Oversight: Google is taking part in the US government's voluntary pre-release review.
- Engineering showcase: Google cites more than 300 tebibytes of freed data-centre memory, a C/C++ to Rust migration, and a video decoder running 2.7 times faster than the earlier Rust version.
For a deeper walk through the launch, see our dedicated Gemini 4 Argon launch review and the earlier Gemini 4 release date and specs guide.
Argon vs Opus 5.5: Google's Own Benchmarks
Carbon has no numbers, so the best available yardstick is Argon's published table, because the Carbon comparison is with Opus 5.5. Google's own table compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5.

| Benchmark | Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| Vals Index | 68.9% | 67.0% | 63.1% |
| DeepSWE v1.1 | 77.9% | 74.2% | 74.1% |
| FrontierSWE v2 | 55.0% | 62.3% | 65.5% |
| Vibe Code Bench | 91.9% | 90.3% | 89.6% |
| Terminal-bench 4.0 | 57.4% | 66.4% | 58.2% |
| PostTrainBench | 45.3% | 49.3% | 44.3% |
| AutomationBench | 51.3% | 42.5% | 41.4% |
| CWE-bench v1 | 68.0% | 67.0% | 68.0% |
| Agent's Last Exam | 39.5% | 38.2% | 34.2% |
Read as a coding comparison the table is genuinely mixed. Argon wins DeepSWE v1.1 and Vibe Code Bench, loses FrontierSWE v2 and Terminal-bench 4.0 to rivals, and loses PostTrainBench to Opus 5.5. Trending Topics notes that Artificial Analysis measured Opus 5.5 at 59.6% on Terminal-Bench 4.0 with a different setup, against the 66.4% in Google's table, a reminder of how much harness choice moves these numbers. Google leads alone in 13 of 19 benchmarks and shares first place in one, but the rows where it trails are mostly the long-horizon, terminal-heavy coding tasks.
That pattern lines up with the anecdote. If a tester says Carbon is "comparable to Opus 5.5 for coding", the most natural reading is that Carbon closes the gap on exactly the agentic and terminal tasks where Argon is weakest. It is also a reading we cannot prove. For a full head-to-head with numbers from the other side, see Claude Opus 5.5 vs GPT-6 Astra.
The Bloomberg Doubts and Google's Reply
The Carbon leak landed days after a different story. Bloomberg reported that some Google employees doubt Argon's real-world coding, saying Gemini 4 scores well on benchmarks but struggles with some practical tasks. Google denied this and pointed to Kavukcuoglu's comments. A Taiwanese outlet reported that Alphabet's share gain on launch day narrowed from over 2% to about 0.5% after the Bloomberg story, which gives a sense of how closely the market reads these leaks.
It is notable that both stories centre on coding. That is the commercial prize for frontier labs, because coding agents drive developer spend and enterprise adoption. If Argon's weaknesses are real, a stronger coding checkpoint two steps behind the flagship is exactly what you would expect Google to be working on. If they are not, Carbon is simply another checkpoint in a normal training run. The reporting cannot distinguish between these.
Context also helps. Trending Topics notes that Google's promised Gemini 3.5 Pro never shipped, and that on the Artificial Analysis Intelligence Index its best model recently trailed the leader by 17 points. Argon is therefore Google's attempt to reset the narrative, and leaks about an even better checkpoint invite the question of why the public model is not the better one.
Antigravity, Context Tiers and the Rollout Clues
TestingCatalog spotted changes in Antigravity that add Gemini 4 Argon to the model selector with three context options: 256K as the default, 512K (about 1.3 times the quota per turn) and 900K (about 1.8 times). These resemble the existing Gemini Flash options. The source is interface strings in an Insiders build, not an announcement, and it is unclear how context size relates to reasoning effort.
Alongside this, Gemini on the web now offers low, medium and high thinking effort settings across available models, and an Antigravity agent previously called Concierge was recently renamed Chief of Staff. TestingCatalog (09/10/2026) says Argon's wider rollout could begin as soon as the following week, but Google has not confirmed a date.
At its Gemini at Work 2026 event Google also gave more detail on Argon, including comments from CEO Sundar Pichai on capability, safety precautions and the upcoming release, according to creator coverage by WorldofAI. The same coverage called the news on subscription tiers disappointing. We could not independently confirm those tier details, so check Google's announcements before changing any plan. For the productivity side of the event, see our guide to Gemini in Workspace agentic features.
The Recursive Self-Improvement Rumour
Some creator coverage links Carbon to speculation that Google has achieved recursive self-improvement (RSI), meaning AI materially improving the process that builds the next AI. The talk cites a Chinese-language report we could not verify, and the Business Insider reporting we can verify says nothing of the kind. A comparison to Opus 5.5 on coding is not evidence of RSI; strong coding models are a precondition, not proof.
This is worth saying because RSI is now a live safety topic. We have covered the serious versions: OpenAI's An Alien Mind essay on RSI and Weco's claims in AIDE² recursive self-improvement, fact-checked. In each case the useful questions are concrete: what part of the research loop is automated, what is measured, and who verified it? Carbon's reporting answers none of them.
The sensible stance is to hold two ideas at once. Labs are using their models heavily in their own engineering (Google says thousands of employees use Argon internally), and the pace of checkpoint iteration is high. But a codename leak is not a capability milestone.
What Developers Should Do Now
- Do not plan around Carbon. There is no API, price or date. Build for Argon and the models you can access today.
- Test coding claims on your own repositories. The gap between benchmark tables and daily use is exactly what the Bloomberg report is about.
- Watch the Antigravity changelog. Context options and model names appear there before announcements.
- Compare cost per solved task. Argon's $2/$10 introductory pricing is half of Opus 5.5's according to Google, but a model that needs more retries can cost more in practice.
- Keep a second provider. Rapid checkpoint turnover is a reason to avoid hard-coding a single vendor.
If you use Claude Code or Antigravity daily, run a fixed set of ten real tasks on each model when access opens and record success, retries and tokens. That gives you better evidence than any leak.
How the Contenders Line Up
| Model | Status (11/10/2026) | Coding evidence | Price signal |
|---|---|---|---|
| Gemini 4 Carbon | Internal checkpoint, unannounced | One anonymous "comparable to Opus 5.5" comment | None |
| Gemini 4 Argon | Announced 30/09/2026, cyber partners only | Google table: mixed vs Opus 5.5 and GPT-6 Astra | $2 / $10 introductory |
| Claude Opus 5.5 | Available | Leads Terminal-bench 4.0 and PostTrainBench in Google's table | About double Argon per Google |
| GPT-6 Astra | Available | Leads FrontierSWE v2 and OSWorld-2.0 in Google's table | See our Astra review |
The one thing the table makes clear is how little is settled. Three labs have models within a few points of each other on most coding benchmarks, and the winner depends on the benchmark and the harness. See our GPT-6 Astra review for the OpenAI side.
Timeline: Ten Days of Gemini 4 News
- 30/09/2026: Google announces Gemini 4 Argon, limited to selected cyber teams through Fairwind, with introductory pricing of $2 / $10 per million tokens.
- Around launch (30/09 to 03/10/2026): Bloomberg reports employee doubts about real-world coding performance; Google denies and cites Kavukcuoglu.
- Around the same period: Google holds its Gemini at Work 2026 event, with further comments on Argon's capabilities, safety and release, according to creator coverage.
- 09/10/2026: Business Insider reports on Carbon, Barium-B and other checkpoints; TestingCatalog reports Argon in the Antigravity model selector.
- 10/10 to 11/10/2026: creators amplify the story, adding RSI speculation and a comparison to Claude Fable 6 and GPT-7 rumours that we could not source.
The pace is the story. Between a flagship announcement and a leak about its successor checkpoint there were barely ten days. That is consistent with how frontier labs now train: many checkpoints from one run are evaluated internally, and the one that ships is chosen for a mixture of capability, safety and cost. For readers it means a leaked codename should be read as "a candidate exists", not "a product is coming".
Why Labs Test Several Checkpoints at Once
The Business Insider report describes employees testing multiple Gemini 4 versions, with names such as Argon, Carbon and Barium-B. That is normal practice. A large training run produces intermediate snapshots, and post-training (reinforcement learning on coding, safety tuning, preference training) can be repeated with different recipes. Each variant has a different profile: one may be stronger at terminal tasks, another more cautious, another cheaper to serve.
Choosing the public release is a trade-off. A model for a trusted-tester programme with the cyber guardrails removed, like Argon in Fairwind, needs extra scrutiny, and Google says it is taking part in the US government's voluntary pre-release review. A stronger checkpoint could take longer to clear. That is a plausible, though unconfirmed, reason why Carbon could trail Argon to market even if testers prefer it.
The practical lesson for developers is to treat model names as moving targets. The model you evaluate in an Antigravity preview, the one you call through an API and the one in a subscription tier may differ, so pin versions and re-run evaluations when names change.
The Safety and Cyber Angle
Argon's limited release is itself a safety decision. Google says Fairwind users get access without the usual cyber guardrails so they can find, validate and patch vulnerabilities, and cites Wiz finding a critical flaw in hospital software. Coding strength and cyber capability are two sides of the same skill, which is why a coding-focused successor like Carbon would raise the same questions again.
We have seen the same dual-use tension on the Anthropic side, for example in our coverage of Project Glasswing. If Carbon is as strong at coding as one tester suggests, expect the same gated approach: defender-first access, government pre-release review and staged availability, rather than an overnight public launch.
The Bottom Line
Gemini 4 Carbon is a real codename, tested by real employees, reported by a credible outlet, and described in one anonymous quote. That is enough to take seriously as a signal that Google is pushing hard on coding quality, and not enough to base a decision on. Argon, the model you may actually get, is strong on Google's own table and weaker on terminal-heavy agentic coding. If Carbon fixes that, it will matter; until Google publishes figures, it is a rumour with a good source. We will update this guide when Google confirms a name, price or date.
Sources
- dnyuz: Business Insider summary of Google employees testing Gemini 4 models and Placera: Alphabet, Google's new AI model tested by employees (BI) (09/10/2026): Carbon testing and the Opus 5.5 comparison.
- TestingCatalog: Gemini 4 Argon hints emerge as Google tests Carbon checkpoint (09/10/2026): Antigravity context options, Barium-B and Chief of Staff rename.
- Trending Topics: Gemini 4 Argon, Google, Anthropic and OpenAI: Argon launch details, pricing, Fairwind rollout and Google's benchmark table.
- NewsBytes: Gemini 4 Argon faces internal scepticism over coding performance: the Bloomberg-sourced doubts and Google's response.
- AI Revolution X, WorldofAI (10/10/2026) and WorldofAI (09/10/2026) on YouTube: creator coverage of Carbon, Argon and Gemini at Work.
- Images: hero collage and benchmark table via Trending Topics; the table is Google DeepMind's own graphic. Vendor-produced, not independent measurements.
Last updated: 11/10/2026. Based on press reports of a single Business Insider story and Google's published Argon figures. We have not tested Carbon or Argon ourselves.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.





