AI Tools Review
Claude Opus 5.5: Benchmarks, Pricing & Safety Review

Review

Claude Opus 5.5: Benchmarks, Pricing & Safety Review

AI Tools Review Editorial Team22 September 2026
  • Anthropic
  • Claude Opus 5.5
  • Benchmarks
  • System Card

Two days after leaks described a stealth-tested model codenamed claude-wafer-eap with a guessed Tuesday launch, Anthropic confirmed nearly every detail: Claude Opus 5.5 shipped on Tuesday 22 September 2026, at almost exactly the leaked price ($4/$20 per million tokens, $0.20 cache reads, $5 cache writes, all matching the leak within cents), with the always-on adaptive thinking the leak described. See our fact-check on how that leak cycle scored in Claude Opus 5.5 & Fable 5.2 Release Date.

This is a full system-card-style review: what Anthropic's own benchmarks show, how the new safeguards work, what changed on pricing and usage limits, and how it stacks up against Fable 5.1, Opus 5 and OpenAI's GPT-6 Astra.

Note: figures, quotes and benchmark scores in this article are drawn directly from Anthropic's official 22 September 2026 launch announcement and system card unless otherwise noted. Several evaluations are Anthropic's own internal runs; independent third-party numbers will refine the picture over the coming weeks.

Summary

  • Launched 22 September 2026 as the first model in the new Claude 5.5 family, replacing Opus 5 as the flagship.
  • Performance: leads on Terminal-Bench 4.0, GDPval-AA v2.1, CursorBench 4.0 and Humanity's Last Exam; trails GPT-6 Astra on AutomationBench and Terminal-Bench-Science.
  • Pricing: $4 input / $20 output per million tokens (20% below Opus 5), cache reads $0.20 (60% below), cache writes $5. Anthropic says typical workloads cost 40% less overall.
  • Safety: Anthropic's best automated-behavioural-audit score to date, and 85% fewer containment-boundary-crossing attempts than Opus 5 or Mythos 5.1.
  • New safeguards: the first Opus model with Fable-5.1-class safeguards on cybersecurity, biology and distillation, a meaningfully more restricted default than Opus 5 had.
  • Also new: increased five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, plus a bankable rate-limit reset.
  • Coming soon: Claude Sonnet 5.5 and Claude Haiku 5.5, "in the coming weeks".

WorldofAI's coverage of the Opus 5.5 leaks two days before Anthropic's official launch confirmed most of the leaked details.

The Benchmark Card: Real Numbers

Anthropic's launch benchmark table compares Opus 5.5 against Fable 5.1, Opus 5 and OpenAI's GPT-6 Astra and GPT-5.6 Sol across nine evaluations. Unless noted, Claude figures use adaptive thinking at max effort, and Opus 5.5 was evaluated with its production safeguards enabled, so cybersecurity, biology and frontier-LLM-development tasks that triggered a fallback to Opus 4.8 or Opus 5 likely understate its raw capability.

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0 (agentic coding)66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 Main (coding)54.4%50.3%48.0%53.3%47.5%
CursorBench 4.0 (coding)57.8%51.8%46.6%41.7%
GDPval-AA v2.1 (knowledge work, Elo)18461735170815421588
AutomationBench (business workflows)40.0%31.4%26.9%41.4%28.8%
Humanity's Last Exam (with tools)67.7%65.6%63.6%57.2%
Terminal-Bench-Science 0.1 (research)58.7%52.6%29.0%64.6%22.4%
OSWorld 2.0 (computer use, partial credit)81.8%80.7%74.0%
Chartography (visual chart recognition, with tools)89.0%88.4%83.4%

Source: Anthropic, 22 September 2026. Terminal-Bench 4.0: Opus 5.5 at xhigh effort, GPT-6 Astra at high effort (each model's highest reported score, as reported by OpenAI for Astra); standard error ±2.6 pts for Opus 5.5. AutomationBench run and reported by Zapier without fallback models, so safeguard interventions counted as failures, understating Opus 5.5's practical score.

Two honest caveats, both flagged by Anthropic itself. First, on AutomationBench and Terminal-Bench-Science, GPT-6 Astra actually leads, a reminder that Opus 5.5 is not a clean sweep. Second, Anthropic explicitly says that at this level of capability, benchmark margins have become a less reliable guide to real-world differences, and that in its own use, the gap between Opus 5.5 and Fable 5.1 is narrower than these scores suggest. Take the headline wins as directional, not as a settled verdict.

Pricing: 20-60% Below Opus 5

Rate (per million tokens)Claude Opus 5.5Claude Opus 5Change
Input tokens$4$5-20%
Output tokens$20$25-20%
Cache reads$0.20$0.50-60%
Cache writes (5-minute)$5$6.25-20%

Source: Anthropic, 22 September 2026, checked 23 September 2026.

A Fast mode is also available in Claude Code and the Claude Platform, running up to 2.5x faster at $8 input / $40 output per million tokens. Anthropic says that because cache reads, the majority of agentic and coding cost, dropped 60%, and Opus 5.5 uses fewer tokens per task than Opus 5, typical workloads cost around 40% less overall, with output generated over 30% faster.

Subscription pricing is unchanged: Claude Pro remains $20/month and Max $100-$200/month. What changed instead is what you get for that price: Anthropic is increasing five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, and adding a bankable rate-limit reset that subscribers can save and use whenever they choose, rather than losing it if unused.

Coding: Large Migrations, Cheaper

Anthropic pitches Opus 5.5 specifically at long, sprawling engineering jobs rather than short snippets. One early tester completed a 680,000-line code migration in under a day, work Anthropic says would have taken an engineering team weeks. Asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app's behaviour, a meaningful reliability gap for anything running unattended.

In an internal head-to-head, Anthropic had Opus 5.5 and Fable 5.1 each translate HAProxy, a widely used load-balancing tool, from C into Rust. Both rewrites passed nearly all of HAProxy's own regression tests, but Opus 5.5 finished in 9.5 hours versus 12 for Fable 5.1, at 51% lower cost. Another early tester, an audit of a 200,000-line codebase, took Opus 5.5 under three hours versus over 20 hours for Opus 5, using 2.5x fewer tokens.

On cost-efficiency specifically: at its default effort level on FrontierCode, Opus 5.5 beats GPT-6 Astra for roughly 20% of the cost per task. On Terminal-Bench 4.0 it matches Astra for about 40% of the cost, and on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost. GitHub's Chief Product Officer Mario Rodriguez said Opus 5.5 "used among the fewest tokens and steps we measured" across GitHub Copilot CLI and VS Code testing, solving more terminal tasks than Opus 5 in less than half the steps.

Anthropic also highlights security posture for autonomous coding agents: a classifier screens every action before it runs, an open-source sandbox lets security teams audit behaviour directly, and code review catches vulnerabilities before they merge. On prompt injection, Opus 5.5 matches or beats Opus 5 across coding, tool use, computer use and web browsing, and ties Fable 5.1 for the lowest prompt-injection success rate of any model tested, per an independent benchmark run by security firm Gray Swan.

Knowledge Work & Research

Anthropic tested Opus 5.5, Fable 5.1 and Opus 5 on a research task deliberately designed to reward honesty over confidence: writing a report on a company's quarterly performance from a version of the web where the actual earnings release was hard to find. An automated grader checked every figure and quote against sources, and any invented number failed the run outright. 16 of 18 Opus 5.5 reports cleared the bar across different effort settings; neither Fable 5.1 nor Opus 5 cleared it on a single attempt. That is a meaningful result for anyone worried about hallucinated figures in agentic research workflows.

Investment firm Walleye Capital, an early tester, reported that Opus 5.5 largely solved their internal evaluation suite even at its lowest effort setting, and at higher settings caught an error in the firm's own evaluation instructions that no other model had noticed. In a separate merger-analysis test, Opus 5.5 and Opus 5 each built a financial model and executive presentation for a fictional HR-software acquisition; both reached the same conclusion, but Opus 5.5's work was more thorough with fewer errors, finishing in 63 minutes versus 93 for Opus 5, at half the cost.

On GDPval-AA v2.1, which grades real-world professional work across 44 occupations, Opus 5.5 scores 1846 Elo at max effort, ahead of Fable 5.1 (1735) and Opus 5 (1708). At its default (medium) effort setting, it beats GPT-6 Astra at max effort for roughly a fifth of the cost per task.

Communication: The Feedback Anthropic Fixed

One of the most common complaints about Opus 5 was verbose, hard-to-follow output, and Anthropic says it rebuilt Opus 5.5's communication style directly in response. It leads with the most important information, uses less jargon, follows explicit style instructions more consistently, and, according to early testers, is noticeably easier to work with over long sessions. One tester summed it up: "it writes the way I do." Ramp's Staff Software Engineer John Ruelas said a design spec came out usable with very minimal edits, and that when Opus 5.5 optimised the company's test suite, he could follow its reasoning easily enough to ship the change with confidence.

Deloitte Consulting reported a related, more concrete benefit: even at its lowest effort setting, Opus 5.5 caught 72% of known bugs in code reviews versus 56% for Opus 5 at high effort, with fewer false alarms and far shorter output, the kind of gain that compounds when lower-effort runs become viable for production use.

Safety: Pacing the Frontier

Opus 5.5 is the first model Anthropic has shipped since CEO Dario Amodei's essay arguing AI progress should be paced so safety practices stay ahead of capability. See our coverage in We Must Pace the Frontier, Explained. Anthropic frames its safety work on two timelines: refining established practices (alignment testing, external pre-release evaluation by METR and Frontier Design, capability-matched safeguards) for the current model generation, while separately preparing training and evaluation processes, including tighter reinforcement-learning environment filtering and stronger interpretability-based monitoring, for more advanced future models it expects "could be trained soon".

On its automated behavioural audit, a suite of nearly 2,000 simulated scenarios and Anthropic's primary alignment evaluation, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behaviour, and is also the strongest model on most honesty measures. It specifically improves on several behaviours implicated in recent industry cybersecurity incidents: biased or motivated reasoning, attempting to escape a sandbox, and taking harmful actions after concluding it is in a simulated environment. In a new evaluation testing whether a model tries to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it did make was low-severity and self-reported, a meaningful result for anyone running Claude unattended across production systems.

Anthropic is candid about the limits here too: it says Opus 5.5 often appears to suspect it is being evaluated, which complicates how confidently pre-deployment testing predicts real-world behaviour, and expects that challenge to grow as capabilities increase unless interpretability research advances. That caveat is why the safeguards below exist alongside, not instead of, the alignment testing.

Safeguards: Cyber, Biology & Distillation

The structural change in this release is safeguards, not just alignment scores. Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology and distillation, each falling back to another model transparently rather than simply refusing.

  • Cybersecurity. Opus 5.5 has extremely strong cyber capabilities, comparable to Mythos 5.1's. Users can still find and fix bugs as part of routine development, but most cybersecurity tasks now re-route to Claude Opus 4.8. Anthropic is expanding its Cyber Verification Program to Opus 5.5 with three tiers of trusted access, including Mythos-model access for the most vetted tier.
  • Biology. Opus 5.5 exceeds Opus 5 and matches or beats Mythos 5.1 across many biology work areas, including a long-horizon molecular prediction and design evaluation run with Dyno Therapeutics, where expert red-teamers rated its scientific novelty comparable to the best model they had tested. Full-capability biology research access requires Anthropic's Life Sciences Verification Program.
  • Distillation. Opus 5.5 ships with "preserved thinking", the anti-distillation safeguard introduced with Fable 5.1, which stops API users editing Claude's prior context to extract its reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after 31 August 2026.

Two other changes carry over from Fable 5.1: Opus 5.5 is available with zero data retention, ships with Anthropic's EU AI Act watermarking, and is no longer available with "thinking" mode switched off, confirming the "always adaptive, no off mode" detail that had circulated in pre-launch leaks.

Availability

Claude Opus 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure. Developers can call it on the Claude Platform as claude-opus-5-5. Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow "in the coming weeks" with many of the same performance, efficiency and safety improvements.

How It Compares

Update, 23 September 2026: about 90 minutes after Opus 5.5 launched, OpenAI released GPT-6 Sol and GPT-6 Luna at $2/$10 and $0.10/$0.50 per million tokens. OpenAI's launch benchmarks were run against Opus 5 and Fable 5.1 rather than Opus 5.5, so the only clean shared number is AutomationBench, where Opus 5.5 scores 40.0% to GPT-6 Sol's 33.2%. For the full picture see GPT-6 Sol vs Claude Opus 5.5, Claude Opus 5.5 vs GPT-6 Astra and our frontier benchmark comparison.

Against Fable 5.1, Opus 5.5 wins every published benchmark, but Anthropic itself says the real-world gap is narrower than the scorecard implies, and Fable 5.1 remains the choice for the biology and cybersecurity work that still routes to Mythos-class safeguards rather than Opus 5.5's more restricted defaults. Against Opus 5, there is no contest on price or most benchmarks, Opus 5.5 is cheaper, faster and scores higher almost everywhere, though it now carries meaningfully tighter cyber and bio safeguards than Opus 5 shipped with. Against GPT-6 Astra, the picture is genuinely mixed: Opus 5.5 leads on coding, knowledge work and reasoning benchmarks, often at a fraction of the cost, but Astra leads on AutomationBench and Terminal-Bench-Science, so task fit still matters more than a single leaderboard position.

Who Should Use It

Switch to Opus 5.5 now if you are on Opus 5 for coding, agentic or knowledge-work tasks, the upgrade is cheaper on every metric and scores higher on almost every benchmark Anthropic published. It is also the clear choice for large, long-running engineering jobs like codebase migrations, where the cost and time savings compound.

Stay on a specialist if your work is offensive security research or frontier biology work that needs full, unsafeguarded capability, Mythos 5.1 and the verification programmes remain the route there, not Opus 5.5's default configuration. And if your task specifically maps to AutomationBench-style business-process automation or scientific literature synthesis, benchmark GPT-6 Astra alongside Opus 5.5 rather than assuming the newer Claude model wins by default.

The Bottom Line

Claude Opus 5.5 is a genuine step change on cost and reliability rather than a pure capability leap: cheaper across every metered rate, meaningfully more reliable on long unattended jobs, and, on Anthropic's own testing, its most aligned model to date. The honest caveats matter too, GPT-6 Astra still leads on two of the nine published benchmarks, the real-world gap to Fable 5.1 is narrower than the scorecard suggests, and the new cyber and bio safeguards are a real, not cosmetic, tightening versus what Opus 5 shipped with.

For everyday coding, research and business-automation work, Opus 5.5 is now the default to reach for on price and reliability alone. Route the specialist cyber and biology work through the verification programmes rather than fighting the new safeguards, and, per Anthropic's own advice, treat any single benchmark margin at this level of capability as a starting point for your own testing rather than a final verdict.

Last updated: 23 September 2026. This review is based on Anthropic's official Claude Opus 5.5 announcement and system card; several benchmark figures come from Anthropic's own internal runs or partner-reported evaluations and may be refined as independent benchmarks land.

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens, both 20% below Claude Opus 5's $5/$25. Cache reads, which Anthropic says make up most agentic and coding costs, drop 60% to $0.20 per million tokens, and cache writes are $5 per million tokens. Anthropic says typical workloads cost about 40% less overall, and output generates over 30% faster.
Is Claude Opus 5.5 better than Claude Fable 5.1?
Anthropic describes Opus 5.5 as performing 'at the level of' Fable 5.1 on most work, and Opus 5.5 leads Fable 5.1 on every benchmark Anthropic published: Terminal-Bench 4.0 (66.4% vs 55.8%), GDPval-AA v2.1 (1846 vs 1735 Elo), CursorBench 4.0 (57.8% vs 51.8%) and Humanity's Last Exam with tools (67.7% vs 65.6%). Anthropic itself cautions that at this level, benchmark margins are a less reliable guide than usual, and says the real-world gap is narrower than the scores suggest.
Does Claude Opus 5.5 beat GPT-6 Astra?
On four of the six benchmarks where Anthropic's launch table has scores for both, yes: Terminal-Bench 4.0, FrontierCode, GDPval-AA v2.1 and Humanity's Last Exam. Anthropic also says Opus 5.5 matches Astra on Terminal-Bench 4.0 for about 40% of the cost per task. But OpenAI's GPT-6 Astra leads on AutomationBench (41.4% vs 40.0%) and Terminal-Bench-Science 0.1 (64.6% vs 58.7%), so 'best model' still depends on the task.
What safeguards does Claude Opus 5.5 have?
Opus 5.5 is the first Opus model to launch with safeguards similar to Fable 5.1's on cybersecurity, biology and distillation. Most cybersecurity tasks are re-routed to Claude Opus 4.8; biology research needing full capability requires Anthropic's Life Sciences Verification Program; and 'preserved thinking' blocks API users from extracting Claude's reasoning via context editing, applying to API accounts created on or after 31 August 2026.
Is Claude Opus 5.5 available now, and what is the API model ID?
Yes, it launched 22 September 2026 on all platforms, including the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. The API model ID is claude-opus-5-5. Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks.

Key takeaways

Cheaper and faster, not just smarter

$4/$20 per million tokens (20% down), cache reads 60% cheaper, output over 30% faster, and Anthropic says typical workloads cost 40% less overall.

Anthropic's most aligned model yet

Best scores of any model to date on the automated behavioural audit; 85% fewer containment-boundary-crossing attempts than Opus 5 or Mythos 5.1.

First Opus with Fable-class safeguards

Cybersecurity, biology and anti-distillation safeguards similar to Fable 5.1, a first for the Opus line, with cyber tasks falling back to Opus 4.8.

AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.