AI Tools Review
Claude Opus 4.6 in Antigravity: What Changes

Insights

Claude Opus 4.6 in Antigravity: What Changes

AI Tools Review Editorial Team06 February 2026

    Quick Answer:

    The February 2026 update to Google Antigravity powered by Claude Opus 4.6 increased the context window from 200k to 1 million tokens, introduced native Agent Teams for complex coordination, and added an Adaptive Thinking engine that plans code refactors before execution. (Anthropic has since shipped Opus 4.7 and 4.8: this piece covers the 4.6-era integration.)

    Google Antigravity, the "Agent-First" IDE that has redefined our development workflows, is currently powered by Claude Opus 4.5. It's powerful, capable, and daily-driving production code for thousands of engineers. But with Anthropic's release of Opus 4.6 this week, Antigravity is about to get a massive engine upgrade. Here is why the combination of Google's IDE and Anthropic's new model is the most anticipated event in dev tools this year.

    From 200K to 1 Million Tokens: The Repo-Level Context

    The current iteration of Antigravity runs on Opus 4.5, which offers a respectable 200k token context window. In practice, this means Antigravity can "see" a mid-sized library or a specific module's worth of documentation and code. It's excellent for task-specific work but often hits a ceiling when refactoring large, interconnected architectures.

    Opus 4.6 brings a 1-million-token context window with "needle-in-a-haystack" retrieval accuracy of 76% (compared to Sonnet 4.5's 18.5%). For Antigravity users, this is transformative.

    • Full-Repo Awareness: Instead of RAG-based context shredding, Antigravity will be able to load entire medium-sized repositories into active memory.
    • Dependency Resolution: The agent can trace a breaking change from a frontend component through the API layer down to the database schema without losing the thread.
    • Legacy Migration: Uploading megabytes of old documentation and legacy code becomes feasible, allowing the agent to plan migrations with full historical context.

    The Small Print on That 1M Window

    Before anyone reorganises their workflow around whole-repository prompts, it is worth reading Anthropic's own framing carefully, because the headline number comes with three qualifications that matter a great deal in a paid IDE.

    First, Opus 4.6 launched on 05/02/2026 with a default context window of 200,000 tokens. The 1-million-token window is an opt-in beta on the Claude Developer Platform rather than the standing default, which means a client such as Antigravity has to deliberately request it.

    Second, it is not free. Anthropic's base rate for Opus 4.6 is $5 per million input tokens and $25 per million output tokens, unchanged from previous Opus releases. But once a prompt crosses the 200,000-token line while using the 1M beta, a premium tier applies at $10 per million input and $37.50 per million output. That is double on the way in and half as much again on the way out. Loading an entire repository into every turn of an agentic loop is therefore a decision with a bill attached, not a free upgrade.

    The practical read

    Large context is best understood as a tool for a specific class of problem, principally migrations, cross-cutting refactors and dependency archaeology, rather than as a replacement for retrieval. For day-to-day feature work, a well-scoped 200k context is usually both faster and cheaper, and the agent is less likely to get distracted by code that has nothing to do with the task.

    Third, output is capped separately. Opus 4.6 supports up to 128,000 output tokens, and Anthropic also shipped a context compaction beta that automatically summarises older turns so long conversations do not simply fall off the end. Anthropic additionally offers a US-only inference option at 1.1x token pricing, which matters if you are working under data-residency constraints. For a fuller treatment of the model itself, see our Opus 4.6 deep dive.

    Supercharging the Agent Manager

    Antigravity's standout feature has always been its Agent Manager, the ability to spin up multiple agents. However, under Opus 4.5, these agents were "smart islands." They worked reasonably well individually but often struggled with complex, coordinated interactions.

    Claude Opus 4.6 Agent Teams Architecture

    Opus 4.6 native Agent Teams capability maps perfectly to Antigravity's architecture.

    Opus 4.6 introduces native Agent Teams capability. In benchmarking, this allowed a "Conductor" agent to orchestrate sub-agents to resolve GitHub issues across multiple repositories autonomously.

    When this lands in Antigravity, we expect the "Agent Manager" to evolve from a multi-tab view into a true engineering management interface. You won't just be chatting with five agents; you'll be defining a high-level goal ("Refactor the auth service"), and the primary Opus 4.6 instance will spin up sub-agents to handle documentation, testing, implementation, and review, all autonomously coordinated.

    Adaptive Thinking in the IDE

    Coding isn't just generating text; it's planning. Opus 4.6's Adaptive Thinking engine is the missing piece for reliable autonomous development.

    Currently, when you ask Antigravity to "fix this bug," Opus 4.5 often jumps straight to a solution. Sometimes it works; sometimes it misses a subtle edge case. Opus 4.6 can detect complexity and automatically pause to "think", simulating a developer staring at the screen and tracing the logic *before* typing a single character.

    Current State (Opus 4.5)

    Linear execution. Great for known patterns, struggles with novel architectural bugs or multi-file race conditions.

    Future State (Opus 4.6)

    Adaptive "Thought Loops." The IDE will pause, verify its assumptions, plan the refactor, and only then execute the edit.

    Effort Levels: The Dial That Actually Matters

    Adaptive Thinking is not a switch you flip. In the Claude API, Opus 4.6 exposes an effort parameter with four settings: low, medium, high (the default) and max. At high effort the model applies extended thinking adaptively, reaching for deeper reasoning only when the problem appears to warrant it. At max, it will spend considerably more tokens before it commits to an answer.

    This is the dial that determines whether an agentic IDE feels brilliant or infuriating, and it is worth understanding how much the published benchmark numbers depend on it. Anthropic's Opus 4.6 system card reports a Terminal-Bench 2.0 score of 65.4% under its max effort configuration, and an ARC-AGI-2 run performed at max effort with a 120,000-token thinking budget. On SWE-bench Verified the model reaches 81.42% with prompt modification. Those are not the numbers you get from a default-configured request, and the gap between a benchmark configuration and a production one is exactly where developer disappointment tends to come from.

    Use low or medium for

    Mechanical edits, renames, boilerplate, test scaffolding, and anything where you already know the answer and just want it typed quickly. Latency is the product here.

    Use high or max for

    Race conditions, architectural decisions, migration planning, and any bug you have already failed to fix once. Correctness is the product, and the extra tokens are cheaper than another afternoon.

    Anthropic also reported that Opus 4.6 leads on GDPval-AA, a third-party evaluation of economically valuable knowledge work, outperforming GPT-5.2 by roughly 144 Elo points and its own predecessor by around 190, alongside a 53.0% score on Humanity's Last Exam. Those figures describe general capability rather than IDE behaviour, but they set the ceiling on what any harness built on top of the model can achieve.

    The Security Audit You Didn't Ask For

    One of the most startling stats from the Opus 4.6 release was its performance in cybersecurity. In blind tests, it found over 500 zero-day vulnerabilities in open-source software without being explicitly trained to do so.

    Opus 4.6 Cybersecurity Performance

    Integrating this into Google Antigravity means your IDE becomes an active security partner. It's not just a linter; it's a penetration tester living in your sidebar. We anticipate Antigravity will leverage this to offer "Security Scan" skills that go far beyond standard static analysis, identifying logic flaws and potential exploit vectors as you write the code.

    When Can We Expect It?

    Google has historically been quick to update Antigravity's model backend. With the API for Opus 4.6 already generally available, an update to Antigravity v1.15.0 or v1.16.0 seems imminent.

    The shift from 4.5 to 4.6 isn't just a version bump; it's a fundamental change in the "autonomy" of the IDE. If Opus 4.5 gave us a very smart junior developer, Opus 4.6 is giving us a senior engineer who can manage their own time, check their own work, and coordinate a small team.

    Start cleaning up your backlogs. When this update drops, you're going to want to give it some real work.

    Update: What Actually Shipped

    This piece was written in February 2026 as a preview. Seven months on, it is worth recording what the integration actually turned into, because the reality is both less dramatic and more useful than the anticipation.

    The Claude models did land in Antigravity's model picker, and they sit alongside Google's own rather than replacing them. The free public preview has shipped with a selection spanning two Gemini 3.1 Pro variants, Gemini 3 Flash, Claude Sonnet 4.6, Claude Sonnet 4.6 with Thinking, Claude Opus 4.6 with Thinking, and the open-weights GPT-OSS 120B. Both Claude models default to Thinking mode. Notably, billing flows through Google's infrastructure rather than requiring you to supply an Anthropic API key, which removes a genuine adoption barrier for teams already on Google Cloud.

    The caveat nobody previewed

    Antigravity caps model output at 64,000 tokens across the board, even though Opus natively supports 128,000. If you are asking for a very large generated file or an exhaustive refactor in one pass, the harness, not the model, is your constraint. This has been a recurring complaint on Google's own AI developer forum.

    The Agent Manager did evolve roughly along the lines this article anticipated, though with a firmer ceiling than the phrase "agent teams" implies. It functions as a mission-control surface where you spawn and supervise agents running asynchronously, with up to five parallel agents each operating in a separate workspace. You review their artifacts, which include implementation plans, task lists, screenshots, browser recordings and completion walkthroughs, approve pending actions, and give feedback without leaving the one view. Agents can plan and execute across editor, terminal and browser, so a single instruction can produce code, launch the app, and verify the new component in a real browser without you intervening.

    Google then pushed the platform considerably further at I/O on 19/05/2026 with Antigravity 2.0, which added a standalone desktop application, a CLI, an SDK, an enterprise tier, dynamic subagents and Gemini 3.5 Flash as the default model. That release did more to change day-to-day workflows than any model swap did.

    One prediction did not land. Anthropic has since shipped Claude Opus 4.7 and Opus 4.8, but Antigravity has not tracked those releases as quickly as this article assumed it would; as of writing, Opus 4.6 remains the Opus option in the picker. The lesson for anyone planning around a third-party IDE is straightforward: model availability inside a harness lags the API by months, not days, and you should budget accordingly. For the Google side of the same story, our guide to Antigravity Skills covers how the platform's own extensibility layer has matured over the same period.

    Frequently Asked Questions

    What did the Claude Opus 4.6 update bring to Google Antigravity?
    The February 2026 update to Google Antigravity powered by Claude Opus 4.6 increased the context window from 200k to 1 million tokens, introduced native Agent Teams for complex coordination, and added an Adaptive Thinking engine that plans code refactors before execution. Anthropic has since shipped Opus 4.7 and 4.8, so this covers the 4.6-era integration.
    How big was the context window jump from Opus 4.5 to Opus 4.6?
    Opus 4.5 offered a 200k token context window, enough to see a mid-sized library or a specific module's worth of code. Opus 4.6 brought a 1-million-token context window with needle-in-a-haystack retrieval accuracy of 76 per cent, compared to Sonnet 4.5's 18.5 per cent, enabling full-repo awareness instead of RAG-based context shredding.
    What are Agent Teams in Claude Opus 4.6?
    Agent Teams was a native Opus 4.6 capability that, in benchmarking, allowed a Conductor agent to orchestrate sub-agents to resolve GitHub issues across multiple repositories autonomously. In Antigravity this meant the Agent Manager could evolve from a multi-tab view into a true engineering management interface, with sub-agents handling documentation, testing, implementation and review from a single high-level goal.
    How good was Opus 4.6 at finding security vulnerabilities?
    In blind tests, Opus 4.6 found over 500 zero-day vulnerabilities in open-source software without being explicitly trained to do so. Integrated into Antigravity, this made the IDE an active security partner able to identify logic flaws and potential exploit vectors as code was written, going well beyond standard static analysis.
    What is Adaptive Thinking in the Antigravity IDE?
    Adaptive Thinking was Opus 4.6's planning engine for reliable autonomous development. Where Opus 4.5 often jumped straight to a solution and could miss subtle edge cases, Opus 4.6 could detect complexity and automatically pause to think, verifying assumptions and planning the refactor before executing any edit.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.