AI Tools Review
Gemini 4 Carbon Leak: What We Know So Far

Insights

Gemini 4 Carbon Leak: What We Know So Far

AI Tools Review Editorial Team11 October 2026

    Quick Answer:

    Gemini 4 Carbon is an unreleased internal Google checkpoint. Business Insider reports that employees testing it on an internal coding platform rate it as comparable to Claude Opus 5.5 for coding, while still asking for more testing. Carbon is separate from Gemini 4 Argon, which Google announced on 30/09/2026 and has so far released only to selected cyber defenders. There are no published benchmarks, price or release date for Carbon, and the "recursive self-improvement" talk around it is unverified speculation.

    Google has not even finished rolling out its newest flagship, and the leaks are already about the one after it. Gemini 4 Argon is still limited to trusted cyber teams, yet reporting says staff are testing a checkpoint called Carbon that they compare to the best coding model on the market.

    This guide separates the three layers of the story: what Google has announced about Argon, what is reported about Carbon from a single Business Insider piece, and what creators and commentators are adding on top. We have not tested any of these models, and where a claim rests on anonymous sources we say so.

    AI Revolution X walks through the Carbon report, Argon in Antigravity and the recursive self-improvement rumours.

    Executive Summary

    • The report: Business Insider says it saw documents and screenshots showing Google employees testing several Gemini 4 checkpoints, including Carbon, which was made available on Google's internal coding platform, reported as Jetski. One employee said Carbon matches Anthropic's Opus 5.5 for coding, but wanted more testing.
    • Argon's status: announced 30/09/2026, with access limited to selected cybersecurity teams. Paying API customers and Google AI Ultra subscribers are next, with no date.
    • Barium-B: TestingCatalog's summary of the report says Barium-B is the checkpoint selected for the public Argon release. We could not confirm this in a second source.
    • Internal doubts: Bloomberg separately reported employee scepticism about Argon's real-world coding. Google pushed back and pointed to comments from DeepMind's Koray Kavukcuoglu.
    • Antigravity hints: interface strings spotted in an Insiders build list Argon with 256K, 512K and 900K context options.
    • Unverified: that Carbon beats Argon, that it is a big jump, that it signals recursive self-improvement, or that it will ship under that name.

    Our view: Carbon is a real codename attached to a thin but credible report. It matters mainly as a signal that Google is iterating on coding quality, the area where Argon's reception has been mixed, ahead of a competitor release cycle led by Claude Opus 5.5.

    What We Actually Know About Carbon

    Everything public about Carbon traces back to one Business Insider story from 09/10/2026, summarised by outlets such as TestingCatalog, dnyuz and Placera. Business Insider reports that it reviewed documents and screenshots showing employees testing internal Gemini 4 versions, and that Carbon was made available on an internal coding platform. TestingCatalog describes the platform as Jetski, Google's internal name associated with Antigravity.

    The key quote is secondhand and anonymous: one employee said Carbon is comparable to Opus 5.5 for coding, while saying they wanted to test it further. That is a feeling from one tester, not a benchmark. It does not say Carbon is better than Opus 5.5, and it does not say on which tasks or how it was judged. Early feedback on Argon checkpoints was described as generally positive, with one employee saying they lagged on some coding tasks.

    What we do not know is longer than what we do. There is no parameter count, no context length, no pricing, no safety evaluation and no release date. It is unclear whether Carbon would ship as an Argon update or as a separate Gemini 4 release. Even the name is unstable: internal codenames change, and Google's previous cycle saw a planned Gemini 3.5 Pro never ship while cheaper Flash models arrived instead, as covered in our Gemini 3.5 Pro write-up.

    One further claim is worth handling with care. TestingCatalog reports that the Business Insider piece says Barium-B is the checkpoint chosen for the public Argon release. If true, Argon as shipped is not the strongest internal checkpoint, which would explain why some employees see Carbon as ahead. It would also mean that Google is shipping a model it considers safe or stable before one it considers stronger, a pattern consistent with its limited, cyber-first rollout.

    Gemini 4 Argon: The Model Carbon Would Follow

    To judge Carbon you have to understand what Argon already is. Google announced Gemini 4 Argon on 30/09/2026 as its first flagship in almost a year. Koray Kavukcuoglu, Google's chief AI architect, said it is "fundamentally changing the way we work and build at Google", and Google says thousands of employees already use it internally for coding, research and writing.

    • Output limit: up to 1 million tokens per response, up from 64,000.
    • Introductory pricing: $2 per million input tokens and $10 per million output tokens, which Google positions as half the price of Claude Opus 5.5.
    • Access: selected cybersecurity teams through Google's Fairwind programme, released without the usual cyber guardrails. Wiz used it to find a critical vulnerability in hospital software that earlier models missed.
    • Oversight: Google is taking part in the US government's voluntary pre-release review.
    • Engineering showcase: Google cites more than 300 tebibytes of freed data-centre memory, a C/C++ to Rust migration, and a video decoder running 2.7 times faster than the earlier Rust version.

    For a deeper walk through the launch, see our dedicated Gemini 4 Argon launch review and the earlier Gemini 4 release date and specs guide.

    Argon vs Opus 5.5: Google's Own Benchmarks

    Carbon has no numbers, so the best available yardstick is Argon's published table, because the Carbon comparison is with Opus 5.5. Google's own table compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5.

    Google DeepMind benchmark table comparing Gemini 4 Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 across knowledge work, agentic coding, ML engineering, science, long context, computer use, multimodal and cybersecurity benchmarks
    Google DeepMind's published benchmark table for Gemini 4 Argon against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 (Gemini 4 Argon column highlighted; methodology on deepmind.google). These are Google's own figures, as reproduced by Trending Topics.
    BenchmarkGemini 4 ArgonClaude Opus 5.5GPT-6 Astra
    Vals Index68.9%67.0%63.1%
    DeepSWE v1.177.9%74.2%74.1%
    FrontierSWE v255.0%62.3%65.5%
    Vibe Code Bench91.9%90.3%89.6%
    Terminal-bench 4.057.4%66.4%58.2%
    PostTrainBench45.3%49.3%44.3%
    AutomationBench51.3%42.5%41.4%
    CWE-bench v168.0%67.0%68.0%
    Agent's Last Exam39.5%38.2%34.2%

    Read as a coding comparison the table is genuinely mixed. Argon wins DeepSWE v1.1 and Vibe Code Bench, loses FrontierSWE v2 and Terminal-bench 4.0 to rivals, and loses PostTrainBench to Opus 5.5. Trending Topics notes that Artificial Analysis measured Opus 5.5 at 59.6% on Terminal-Bench 4.0 with a different setup, against the 66.4% in Google's table, a reminder of how much harness choice moves these numbers. Google leads alone in 13 of 19 benchmarks and shares first place in one, but the rows where it trails are mostly the long-horizon, terminal-heavy coding tasks.

    That pattern lines up with the anecdote. If a tester says Carbon is "comparable to Opus 5.5 for coding", the most natural reading is that Carbon closes the gap on exactly the agentic and terminal tasks where Argon is weakest. It is also a reading we cannot prove. For a full head-to-head with numbers from the other side, see Claude Opus 5.5 vs GPT-6 Astra.

    The Bloomberg Doubts and Google's Reply

    The Carbon leak landed days after a different story. Bloomberg reported that some Google employees doubt Argon's real-world coding, saying Gemini 4 scores well on benchmarks but struggles with some practical tasks. Google denied this and pointed to Kavukcuoglu's comments. A Taiwanese outlet reported that Alphabet's share gain on launch day narrowed from over 2% to about 0.5% after the Bloomberg story, which gives a sense of how closely the market reads these leaks.

    It is notable that both stories centre on coding. That is the commercial prize for frontier labs, because coding agents drive developer spend and enterprise adoption. If Argon's weaknesses are real, a stronger coding checkpoint two steps behind the flagship is exactly what you would expect Google to be working on. If they are not, Carbon is simply another checkpoint in a normal training run. The reporting cannot distinguish between these.

    Context also helps. Trending Topics notes that Google's promised Gemini 3.5 Pro never shipped, and that on the Artificial Analysis Intelligence Index its best model recently trailed the leader by 17 points. Argon is therefore Google's attempt to reset the narrative, and leaks about an even better checkpoint invite the question of why the public model is not the better one.

    Antigravity, Context Tiers and the Rollout Clues

    TestingCatalog spotted changes in Antigravity that add Gemini 4 Argon to the model selector with three context options: 256K as the default, 512K (about 1.3 times the quota per turn) and 900K (about 1.8 times). These resemble the existing Gemini Flash options. The source is interface strings in an Insiders build, not an announcement, and it is unclear how context size relates to reasoning effort.

    Alongside this, Gemini on the web now offers low, medium and high thinking effort settings across available models, and an Antigravity agent previously called Concierge was recently renamed Chief of Staff. TestingCatalog (09/10/2026) says Argon's wider rollout could begin as soon as the following week, but Google has not confirmed a date.

    At its Gemini at Work 2026 event Google also gave more detail on Argon, including comments from CEO Sundar Pichai on capability, safety precautions and the upcoming release, according to creator coverage by WorldofAI. The same coverage called the news on subscription tiers disappointing. We could not independently confirm those tier details, so check Google's announcements before changing any plan. For the productivity side of the event, see our guide to Gemini in Workspace agentic features.

    The Recursive Self-Improvement Rumour

    Some creator coverage links Carbon to speculation that Google has achieved recursive self-improvement (RSI), meaning AI materially improving the process that builds the next AI. The talk cites a Chinese-language report we could not verify, and the Business Insider reporting we can verify says nothing of the kind. A comparison to Opus 5.5 on coding is not evidence of RSI; strong coding models are a precondition, not proof.

    This is worth saying because RSI is now a live safety topic. We have covered the serious versions: OpenAI's An Alien Mind essay on RSI and Weco's claims in AIDE² recursive self-improvement, fact-checked. In each case the useful questions are concrete: what part of the research loop is automated, what is measured, and who verified it? Carbon's reporting answers none of them.

    The sensible stance is to hold two ideas at once. Labs are using their models heavily in their own engineering (Google says thousands of employees use Argon internally), and the pace of checkpoint iteration is high. But a codename leak is not a capability milestone.

    What Developers Should Do Now

    • Do not plan around Carbon. There is no API, price or date. Build for Argon and the models you can access today.
    • Test coding claims on your own repositories. The gap between benchmark tables and daily use is exactly what the Bloomberg report is about.
    • Watch the Antigravity changelog. Context options and model names appear there before announcements.
    • Compare cost per solved task. Argon's $2/$10 introductory pricing is half of Opus 5.5's according to Google, but a model that needs more retries can cost more in practice.
    • Keep a second provider. Rapid checkpoint turnover is a reason to avoid hard-coding a single vendor.

    If you use Claude Code or Antigravity daily, run a fixed set of ten real tasks on each model when access opens and record success, retries and tokens. That gives you better evidence than any leak.

    How the Contenders Line Up

    ModelStatus (11/10/2026)Coding evidencePrice signal
    Gemini 4 CarbonInternal checkpoint, unannouncedOne anonymous "comparable to Opus 5.5" commentNone
    Gemini 4 ArgonAnnounced 30/09/2026, cyber partners onlyGoogle table: mixed vs Opus 5.5 and GPT-6 Astra$2 / $10 introductory
    Claude Opus 5.5AvailableLeads Terminal-bench 4.0 and PostTrainBench in Google's tableAbout double Argon per Google
    GPT-6 AstraAvailableLeads FrontierSWE v2 and OSWorld-2.0 in Google's tableSee our Astra review

    The one thing the table makes clear is how little is settled. Three labs have models within a few points of each other on most coding benchmarks, and the winner depends on the benchmark and the harness. See our GPT-6 Astra review for the OpenAI side.

    Timeline: Ten Days of Gemini 4 News

    • 30/09/2026: Google announces Gemini 4 Argon, limited to selected cyber teams through Fairwind, with introductory pricing of $2 / $10 per million tokens.
    • Around launch (30/09 to 03/10/2026): Bloomberg reports employee doubts about real-world coding performance; Google denies and cites Kavukcuoglu.
    • Around the same period: Google holds its Gemini at Work 2026 event, with further comments on Argon's capabilities, safety and release, according to creator coverage.
    • 09/10/2026: Business Insider reports on Carbon, Barium-B and other checkpoints; TestingCatalog reports Argon in the Antigravity model selector.
    • 10/10 to 11/10/2026: creators amplify the story, adding RSI speculation and a comparison to Claude Fable 6 and GPT-7 rumours that we could not source.

    The pace is the story. Between a flagship announcement and a leak about its successor checkpoint there were barely ten days. That is consistent with how frontier labs now train: many checkpoints from one run are evaluated internally, and the one that ships is chosen for a mixture of capability, safety and cost. For readers it means a leaked codename should be read as "a candidate exists", not "a product is coming".

    Why Labs Test Several Checkpoints at Once

    The Business Insider report describes employees testing multiple Gemini 4 versions, with names such as Argon, Carbon and Barium-B. That is normal practice. A large training run produces intermediate snapshots, and post-training (reinforcement learning on coding, safety tuning, preference training) can be repeated with different recipes. Each variant has a different profile: one may be stronger at terminal tasks, another more cautious, another cheaper to serve.

    Choosing the public release is a trade-off. A model for a trusted-tester programme with the cyber guardrails removed, like Argon in Fairwind, needs extra scrutiny, and Google says it is taking part in the US government's voluntary pre-release review. A stronger checkpoint could take longer to clear. That is a plausible, though unconfirmed, reason why Carbon could trail Argon to market even if testers prefer it.

    The practical lesson for developers is to treat model names as moving targets. The model you evaluate in an Antigravity preview, the one you call through an API and the one in a subscription tier may differ, so pin versions and re-run evaluations when names change.

    The Safety and Cyber Angle

    Argon's limited release is itself a safety decision. Google says Fairwind users get access without the usual cyber guardrails so they can find, validate and patch vulnerabilities, and cites Wiz finding a critical flaw in hospital software. Coding strength and cyber capability are two sides of the same skill, which is why a coding-focused successor like Carbon would raise the same questions again.

    We have seen the same dual-use tension on the Anthropic side, for example in our coverage of Project Glasswing. If Carbon is as strong at coding as one tester suggests, expect the same gated approach: defender-first access, government pre-release review and staged availability, rather than an overnight public launch.

    The Bottom Line

    Gemini 4 Carbon is a real codename, tested by real employees, reported by a credible outlet, and described in one anonymous quote. That is enough to take seriously as a signal that Google is pushing hard on coding quality, and not enough to base a decision on. Argon, the model you may actually get, is strong on Google's own table and weaker on terminal-heavy agentic coding. If Carbon fixes that, it will matter; until Google publishes figures, it is a rumour with a good source. We will update this guide when Google confirms a name, price or date.

    Sources

    Last updated: 11/10/2026. Based on press reports of a single Business Insider story and Google's published Argon figures. We have not tested Carbon or Argon ourselves.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    What is Gemini 4 Carbon?
    Carbon is the codename of an internal Gemini 4 checkpoint that, according to a Business Insider report published around 09/10/2026, Google employees have been testing on an internal coding platform. One employee said it was comparable to Anthropic's Claude Opus 5.5 for coding, while asking for more testing. Google has not announced Carbon, and no benchmarks, pricing or release date exist for it.
    Is Gemini 4 Carbon better than Gemini 4 Argon?
    Nobody outside Google can say. Reporting based on Business Insider says Argon is the public release and that Carbon is a separate checkpoint under test, with early feedback that it is comparable to Opus 5.5 for coding. Because there are no published figures for Carbon, any claim that it beats Argon is speculation. It is also unclear whether Carbon will become an Argon update or a separate Gemini 4 release.
    When will Gemini 4 Argon be generally available?
    Google announced Gemini 4 Argon on 30/09/2026 and initially limited it to selected cybersecurity partners through its Fairwind programme. Paying API customers and Google AI Ultra subscribers are listed as next, with Google saying only "as soon as possible". TestingCatalog reported on 09/10/2026 that the wider rollout could start as soon as the following week, but Google has not confirmed a date.
    Does Gemini 4 Argon beat Claude Opus 5.5?
    On Google's own published table Argon leads Opus 5.5 on most rows, for example Vals Index 68.9% against 67.0% and DeepSWE v1.1 77.9% against 74.2%, but Opus 5.5 leads on Terminal-Bench 4.0 (66.4% against 57.4%), PostTrainBench (49.3% against 45.3%) and FrontierSWE v2 (62.3% against 55.0%). All of these are Google's numbers, with no independent replication yet, and Bloomberg has reported internal doubts about real-world coding.
    Did Google achieve recursive self-improvement with Carbon?
    There is no evidence of that. The idea comes from speculation that a big jump between checkpoints might mean AI is helping to build AI, and creators have repeated it, citing a Chinese-language report we could not verify. The only sourced facts are an anonymous employee comparison with Opus 5.5 on coding and a codename. Treat recursive self-improvement claims as unverified rumour until Google publishes something.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.