AI Tools Review
Claude Opus 5: Benchmarks, System Card & Review

Claude Opus 5: Benchmarks, System Card & Review

24 July 2026

Quick Answer:

Claude Opus 5 is Anthropic's 24 July 2026 flagship, succeeding Opus 4.8. Anthropic calls it a "thoughtful and proactive" model that comes close to the frontier intelligence of Claude Fable 5 at about half the price. It is the new state of the art on coding and knowledge-work suites - 43.3% on Frontier-Bench v0.1 (more than double Opus 4.8), 30.2% on ARC-AGI-3 (roughly 3x the next-best model) and 1861 on GDPval-AA v2 - and it is Anthropic's most aligned model to date, scoring 2.3 on the automated behavioural audit. Pricing is unchanged from Opus 4.8 at $5/$25 per million tokens. The one clear ceiling: it stays deliberately behind Mythos 5 on offensive cyber and autonomous biology.

Anthropic's pitch for Opus 5 is unusually blunt: near-frontier intelligence, half the price, and the safest model it has ever shipped. For once the benchmark card mostly backs the claim - and the places where it does not are the most revealing part of the launch.

This is a full system-card-style review: what changed, what the evaluations actually show, where Opus 5 wins, where it is intentionally held back, and how to think about it against the rest of the July 2026 frontier.

Executive Summary

Claude Opus 5 is the rare flagship that is both the most capable and the most aligned model in its maker's line-up at the same time. On Anthropic's own comparison card it leads Fable 5, Opus 4.8 and GPT-5.6 Sol on the majority of rows - agentic terminal coding, knowledge work, novel problem-solving, computer use, business workflows - while its automated behavioural audit score falls to the lowest Anthropic has recorded. The headline story is not a single benchmark; it is that the reliability and safety curve moved in the right direction without a price rise.

The model is designed to be used every day. It is now the default model on Claude Max and the strongest option on Claude Pro, and Anthropic emphasises token efficiency - it tends to reach an answer with less wasted effort than its predecessors. A tunable effort setting (low, medium, high, xhigh, max) lets you trade intelligence against speed and cost, and the cost-performance curves show Opus 5 sitting above every rival at a given price on the tasks that matter.

  • Best for: agentic software engineering, knowledge work, computer use, business-process automation and everyday high-volume use where cost per task matters.
  • Headline numbers: 43.3% Frontier-Bench v0.1, 30.2% ARC-AGI-3, 1861 GDPval-AA v2, 70.6% OSWorld 2.0, 26.0% AutomationBench - leading the field on most.
  • Defining trait: it verifies its own work and iterates until it succeeds, rather than declaring victory early.
  • Main caveat: deliberately behind Mythos 5 on offensive cyber exploitation and autonomous biology research.

Lineage: From Opus 4.8 to Opus 5

Opus 5 arrives only weeks after Opus 4.8, and the jump in the version number - from a 4.x point release to a whole new major - is earned by the size of the capability step rather than any single new feature. Where 4.8 was a consolidation release that hardened reliability, Opus 5 is a genuine capability leap: on Frontier-Bench it more than doubles 4.8's score, and on the novel-reasoning ARC-AGI-3 test it moves from a near-zero 1.5% to 30.2%.

The Opus line sits alongside Anthropic's two research-frontier siblings, Fable 5 and Mythos 5. Fable 5 is the raw-intelligence frontier model; Mythos 5 is the specialist that leads on cybersecurity and long-horizon biology. Opus 5's role is different again: it is the model built to be run constantly, closing most of the gap to Fable 5 while costing roughly half as much per task. That positioning - "close to the frontier, at everyday prices" - is the through-line of the entire launch.

The Effort Ladder and Training

Anthropic keeps the precise architecture proprietary while disclosing capabilities and safety properties in depth. Two things about how Opus 5 is built matter for anyone deploying it. The first is the effort setting, a control that spans low, medium, high, xhigh and max. It lets you optimise for intelligence or conserve tokens for faster, cheaper results, and it is the single biggest lever on both quality and spend. The published benchmark scores come from the higher rungs of that ladder; at lower effort Opus 5 behaves like a fast, capable workhorse. Treating the two as one undifferentiated "model" is the most common way teams over- or under-spend.

The second is a training emphasis on verification and iteration. Anthropic reports that Opus 5 is much stronger at checking its own work and iterating carefully until it succeeds, rather than stopping at a plausible-looking first answer. The early-access anecdotes make the point vividly: given a machine-part drawing it could not directly view, Opus 5 wrote its own computer-vision pipeline to recover the geometry from raw pixels and rebuilt the part in FreeCAD - succeeding repeatedly where no competing model in the same setup could after five attempts. Given a real bug in a popular open-source package manager, it found the root cause and fixed an edge case the community's own patch had missed, while a rival model fixed only the surface symptom and declared the bug resolved.

That behaviour - building a test harness when none exists, refusing to trust an unverified result - is what Anthropic means by "thoughtful and proactive". It is also the trait most likely to translate into real-world reliability, because it attacks the failure mode that hurts agentic systems most: confidently reporting success on work that was never actually completed.

Capabilities Deep Dive

Agentic software engineering

This is where Opus 5 is strongest relative to the field. On Frontier-Bench v0.1, an agentic terminal-coding benchmark, it scores 43.3% - ahead of GPT-5.6 Sol (34.4%), Fable 5 (33.7%) and more than double Opus 4.8 (21.1%). On FrontierCode v1.1 it edges the field at 53.4%, and it holds coherent plans across long chains of tool calls with far better error recovery than 4.8. The one honest exception is DeepSWE v1.1, where GPT-5.6 Sol (72.7%) and Fable 5 (69.7%) both nose ahead of Opus 5 (68.8%) - a reminder that "best coding model" is harness- and task-dependent.

Novel problem-solving

ARC-AGI-3 is the eye-catching result. It measures whether a model can solve genuinely novel problems it has not seen patterns for, and Opus 5's 30.2% is about three times GPT-5.6 Sol's 7.8% and roughly twenty times Opus 4.8's 1.5%. Fable 5 does not post a score. This is the clearest evidence that the gains are in reasoning rather than memorised benchmark structure.

Knowledge work and business automation

On GDPval-AA v2, which grades economically valuable knowledge work, Opus 5 scores 1861, clear of Fable 5 (1747), GPT-5.6 Sol (1736) and Opus 4.8 (1593). On Zapier's AutomationBench, which tests whether a model can complete business tasks end to end, it reaches 26.0% - around 1.5x the next-best model for the same cost per task, and even at its lowest effort setting it passes more tasks than any other model. For anyone building automated back-office workflows, that end-to-end completion rate is the number that matters.

Computer use and agentic search

On OSWorld 2.0, a computer-use benchmark, Opus 5 posts 70.6%, ahead of Fable 5 (66.1%), GPT-5.6 Sol (62.6%) and Opus 4.8 (55.7%) - and Anthropic says it surpasses Fable 5's best result at just over a third of the cost. On the BrowseComp agentic-search test it scores 90.8%, narrowly ahead of GPT-5.6 Sol (90.4%). These are the capabilities that underpin desktop and browser agents, and Opus 5 leads on both.

Multidisciplinary reasoning and visual output

On Humanity's Last Exam, Opus 5 scores 56.3% without tools and 64.7% with tools, effectively level with Fable 5 (56.5% / 63.9%) and well ahead of Opus 4.8. Anthropic also highlights markedly stronger visual outputs - it built an interactive wind-tunnel simulation visualising airflow over aerodynamic and non-aerodynamic objects - which matters for anyone using the model to generate diagrams, simulations or front-end interfaces.

The Benchmark Card: Real Numbers

Anthropic's headline comparison card puts Opus 5 against Fable 5, Opus 4.8 and OpenAI's GPT-5.6 Sol across twelve evaluations. Opus 5 is highlighted as the leader on the majority of them, and the table is the single most useful artefact in the whole announcement.

Benchmark comparison table: Claude Opus 5 versus Fable 5, Opus 4.8 and GPT-5.6 Sol across agentic coding, knowledge work, novel problem-solving, agentic search, reasoning, computer use, business workflows, legal, health and biology.
Opus 5 versus Fable 5, Opus 4.8 and GPT-5.6 Sol. Opus 5 leads most rows; a few cells substitute Mythos 5 (Health 66.0%, BioMysteryBench human-solved 89.0%). Source: Anthropic.
  • Agentic terminal coding (Frontier-Bench v0.1): 43.3% - ahead of GPT-5.6 Sol (34.4%), Fable 5 (33.7%) and Opus 4.8 (21.1%).
  • Knowledge work (GDPval-AA v2): 1861 - clear of Fable 5 (1747), GPT-5.6 Sol (1736) and Opus 4.8 (1593).
  • Novel problem-solving (ARC-AGI-3): 30.2% - versus 7.8% for GPT-5.6 Sol and 1.5% for Opus 4.8.
  • Agentic search (BrowseComp): 90.8% - narrowly leading GPT-5.6 Sol (90.4%).
  • Computer use (OSWorld 2.0): 70.6% - ahead of Fable 5 (66.1%) and GPT-5.6 Sol (62.6%).
  • Business workflows (AutomationBench): 26.0% - well clear of the ~17-18% cluster from the other three.
  • Agentic coding (DeepSWE v1.1): 68.8% - behind GPT-5.6 Sol (72.7%) and Fable 5 (69.7%).
  • Health (HealthBench Professional): 59.8% - here the specialist Mythos 5 (66.0%) leads, with GPT-5.6 Sol (60.5%) also ahead.

Two honest caveats. First, a handful of cells swap in Mythos 5 rather than Fable 5 - notably Health and the "human-solved" tier of BioMysteryBench - so read the column labels carefully; those are the tasks where Anthropic's specialist model, not Opus 5, is the benchmark to beat. Second, agentic scores are harness-sensitive: the same model scores differently under different scaffolds, so treat the cross-model gaps as directional rather than exact. For the wider competitive picture, see our model wars roundup.

Cost-Effectiveness and Effort

The benchmark card tells you Opus 5 is fast; the effort-versus-cost plot tells you it is efficient. Because the effort setting spans five rungs, Anthropic can chart score against dollars spent per attempt - and on Frontier-Bench, Opus 5's curve sits above every rival at any given price point on the high, xhigh and max rungs.

Line chart of agentic coding score versus cost per attempt on a log scale for Frontier-Bench v0.1: Opus 5 reaches about 44% and sits above every rival's cost-performance curve, Fable 5 tops out near 33.7% at higher cost, GPT-5.6 Sol near 37.5%, Opus 4.8 near 18.8%.
Agentic coding score by effort level on Frontier-Bench v0.1, plotted against cost per attempt (log scale). Opus 5 reaches roughly 44% and leads the cost-performance frontier. Source: Anthropic (internal run, mini-SWE-agent harness).

The practical reading: Opus 5 peaks around 44% on Frontier-Bench at roughly $14-15 per attempt, while Fable 5 needs close to $28 per attempt to reach its ceiling near 33.7%. GPT-5.6 Sol tops out around 37.5%, and Opus 4.8 sits far below at about 18.8%. Anthropic notes Opus 5 more than doubles Opus 4.8's performance at a lower cost per task - the exact combination of "more capable" and "cheaper" that is usually a marketing fiction and here is a plotted line. On CursorBench 3.2 at max effort, Anthropic adds that Opus 5 lands within 0.5% of Fable 5's peak score at half the cost per task.

The footnote is worth reading too: these are internal runs on the mini-SWE-agent harness with a GKE backend, mean reward over five attempts per task, and Opus 4.8 served as the fallback on safety-classifier refusals for both Opus 5 and Fable 5. That fallback detail matters when you compare exploitation-heavy cyber scores later.

System Card: Safety and Alignment

The most striking claim in the launch is that Opus 5 is Anthropic's most aligned model to date. During pre-deployment testing, its automated behavioural audit found Opus 5 adheres to Claude's Constitution better than Opus 4.8, Sonnet 5 or Fable 5; exhibits the lowest rates of deceptive behaviour; and is the least susceptible to being tricked into misuse. It is also Anthropic's safest model yet at avoiding reckless, hard-to-reverse actions - the property that matters most for agents with real system access.

Bar chart of overall misaligned-behaviour score (1-10, lower is better) from Anthropic's automated behavioural audit: Opus 4.8 2.85, Mythos 5 2.81, Sonnet 5 3.35, Opus 5 2.30.
Overall misaligned-behaviour score (lower is better): Opus 5 scores 2.30, below Opus 4.8 (2.85), Mythos 5 (2.81) and Sonnet 5 (3.35). Source: Anthropic.

On the automated behavioural audit, Opus 5 scores 2.3 on overall misaligned behaviour - the lowest of Anthropic's recent models, and comfortably under Sonnet 5's 3.35. Calibration and honesty are notoriously hard to move in the right direction while raising raw capability, because a more capable model is also a more capable deceiver. That Opus 5 pushes both dials the right way at once is the launch's quietest but most important result. For agentic deployments the standard discipline still applies: scoped permissions, human-in-the-loop checkpoints for irreversible actions, and logging you actually review. Better alignment reduces the risk; it does not remove the need for containment.

Cybersecurity: OSS-Fuzz and Safeguards

Opus 5 does not advance the frontier in risky dual-use capability. In evaluations run with private-sector and government partners, Anthropic found it remains behind Mythos 5 in both biology research and offensive cybersecurity. As with Opus 4.8, Anthropic deliberately avoided training Opus 5 on cyber tasks - yet it improved substantially anyway as a by-product of becoming more generally capable, and it now comes close to Mythos 5 at finding vulnerabilities. The crucial gap is in exploitation: turning a vulnerability into a working exploit.

Two bar charts from the OSS-Fuzz evaluation. Left, vulnerability identification pass rate: Opus 4.8 61.5%, Mythos 5 80.0%, Opus 5 79.4%. Right, exploitation success count: Opus 4.8 0, Mythos 5 13, Opus 5 4.
OSS-Fuzz: Opus 5 nearly matches Mythos 5 at identifying vulnerabilities (79.4% vs 80.0%) but lags far behind on developing exploits (4 challenges vs 13). Source: Anthropic.

The two panels tell the whole story. On vulnerability identification, Opus 5 (79.4%) is almost level with Mythos 5 (80.0%) and well above Opus 4.8 (61.5%). On exploitation success - the step that turns a finding into a material cyber threat - Opus 5 fully solves just 4 challenges against Mythos 5's 13, while Opus 4.8 solves none. Anthropic's read is that Opus 5 has become genuinely useful for defensive vulnerability discovery without closing the gap on the offensive capability that carries the most risk.

The safeguards reflect that. Opus 5's cyber classifiers are proportionally less restrictive than Fable 5's: they allow finding vulnerabilities in source code but block binary-based vulnerability scanning (a method more associated with malicious actors), penetration testing and exploit generation. Anthropic expects them to intervene around 85% less often than Fable 5's, and in Claude.ai, Claude Code and Claude Cowork any flagged request falls back to Opus 4.8 by default. Enterprises and researchers in Anthropic's Cyber Verification Program get access to a less-restricted version for legitimate security work.

Life Sciences and Scientific Research

For scientific research, Opus 5 is a meaningful step up from Opus 4.8, beating it on every one of Anthropic's life-sciences evaluations - spanning structural biology, organic chemistry and bioinformatics. The gains are largest in organic chemistry: on tasks like inferring molecular structures from spectroscopy data it scores 10.2 percentage points higher than Opus 4.8. On protein work - for example predicting how variations in a protein's sequence affect its function - it scores 7.7 percentage points higher. Because its safeguards are similar to Opus 4.8's, Opus 5 is now Anthropic's most capable generally available model for scientific research.

There is a deliberate ceiling. On long-running, autonomous research tasks - where Anthropic expects AI to pose the most substantial biology-related risk - Opus 5 still shows important limitations, and Mythos 5 remains the stronger model for that work. As part of the launch, biology-related requests that were blocked on Fable 5 now route to Opus 5 rather than Opus 4.8, making it the practical default for legitimate life-sciences use.

Pricing, Access and Deployment

Opus 5 is available today on all platforms at the same price as Opus 4.8: $5 (around £3.95) per million input tokens and $25 (around £19.75) per million output tokens. Developers can call claude-opus-5 on the Claude API. It is the new default model on Claude Max and the strongest model on Claude Pro, which is a notable shift - Anthropic is putting its newest flagship in front of everyday subscribers rather than reserving it for the top tier.

A Fast mode runs Opus 5 at roughly 2.5x the default speed, priced at twice the base rate on the Claude Platform and through usage credits in Claude Code. Two beta features ship alongside: mid-conversation tool changes on the Claude Platform, which let developers change which tools Claude can use without invalidating the prompt cache; and automatic fallbacks on the API, so requests flagged by the safety classifiers on Opus 5 (or Fable 5) can route to another model rather than being blocked. As with prior Opus models, Opus 5 has no data-retention requirements for general access.

As always, treat any per-token figure as a starting point and benchmark it against your own token profile - effort level, context length and tool-call volume dominate real spend far more than the sticker price. The good news is that Opus 5's cost-performance curve gives you a genuinely cheaper path to a given quality bar than Opus 4.8 did.

Limitations and Known Issues

  • Not the cyber or bio frontier: Mythos 5 remains ahead on offensive cybersecurity exploitation and on long-running autonomous biology research - by design.
  • Loses a few coding rows: GPT-5.6 Sol leads DeepSWE v1.1 (72.7% vs 68.8%), and Fable 5 edges Opus 5 on one or two suites. "Best coding model" is task-dependent.
  • Effort-level cost: the headline scores come from the higher effort rungs, which are the expensive ones; low effort is cheaper but noticeably weaker.
  • Classifier fallbacks: flagged cyber and some bio requests silently fall back to Opus 4.8 unless you are in the Cyber Verification Program, which can surprise security teams.
  • Benchmarks are Anthropic's own: several figures come from internal runs; independent third-party numbers will refine the picture over the coming weeks.

How It Compares

Against Fable 5, Opus 5 is the value play: it closes most of the intelligence gap and beats Fable 5 outright on several agentic suites, at roughly half the cost per task - though Fable 5 keeps a narrow edge on a couple of coding benchmarks. Against Mythos 5, the division is cleaner: Mythos 5 owns cybersecurity exploitation and autonomous biology, while Opus 5 leads on general agentic and knowledge work. Our Fable 5 vs Mythos 5 comparison unpacks that split in detail.

Against OpenAI's GPT-5.6 Sol, the two trade blows: Opus 5 leads on Frontier-Bench, ARC-AGI-3, GDPval, OSWorld and AutomationBench, while GPT-5.6 Sol takes DeepSWE and stays competitive on agentic search. And against its own predecessor Opus 4.8, there is no contest - Opus 5 wins every row on the card, usually by a wide margin, at the same price. The honest summary is that capability is converging at the top and the differentiators are increasingly reliability, safety posture and cost per task rather than a single leaderboard position - and on cost per task, Opus 5 is currently the one to beat.

Who Should Use It

Switch to Opus 5 now if you run agentic coding, knowledge work, computer-use or business-automation workloads, if you are on Opus 4.8 (the upgrade is close to free and strictly better), or if you want near-Fable-5 quality at roughly half the cost per task. For most everyday production use, it is now the obvious default - which is precisely why Anthropic made it the default on Max.

Stay on a specialist if your work is offensive security or long-horizon autonomous biology research, where Mythos 5 remains ahead, or if your pipeline depends on a specific coding benchmark - like DeepSWE-style tasks - where GPT-5.6 Sol or Fable 5 currently edge it. And if you are a security researcher who needs the fuller cyber capability, look at the Cyber Verification Program rather than fighting the default classifiers.

The Bottom Line

Claude Opus 5 is the most complete model Anthropic has shipped: state of the art on the coding and knowledge-work benchmarks that pay the bills, three times the next-best model on novel reasoning, and - unusually - the most aligned and safest model in the line-up at the same time, all at Opus 4.8's price. The benchmark card is genuinely impressive precisely because it is not a clean sweep: the rows Opus 5 loses, and the cyber and bio ceilings Anthropic built in on purpose, are what make the wins credible.

For everyday, high-volume, agentic work, Opus 5 is the new default to reach for, and the cost-performance curve is the clearest reason why. Respect the safety guidance, understand which effort rung you are paying for, and route the specialist tasks to the specialist models - do that, and Opus 5's "near-frontier intelligence at half the price" stops being a slogan and starts being a line on your invoice.

Last updated: July 2026. This review is based on Anthropic's official Claude Opus 5 announcement and benchmark disclosures; several figures come from Anthropic's internal runs and may be refined as independent benchmarks land.

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

Is Claude Opus 5 better than Opus 4.8?
Yes, substantially. On Frontier-Bench v0.1 agentic terminal coding, Opus 5 scores 43.3% against Opus 4.8's 21.1% - more than double - and at a lower cost per task. It also leads Opus 4.8 on every life-sciences evaluation Anthropic runs, and is Anthropic's most aligned model to date at 2.3 on the automated behavioural audit versus 2.85 for Opus 4.8.
How much does Claude Opus 5 cost?
Opus 5 costs the same as Opus 4.8: $5 (around £3.95) per million input tokens and $25 (around £19.75) per million output tokens. A Fast mode runs roughly 2.5x the default speed at twice the base price. Anthropic positions it as coming close to Fable 5's intelligence at about half the price per task.
Is Opus 5 better than Fable 5 or Mythos 5?
It depends on the task. Opus 5 leads Fable 5 on Frontier-Bench, GDPval-AA, ARC-AGI-3, OSWorld 2.0 and AutomationBench, often at a fraction of the cost. But it stays behind Mythos 5 on cybersecurity exploitation and on long-running autonomous biology research, and Fable 5 still edges it on a few coding suites such as DeepSWE.
Why is Opus 5 restricted on cybersecurity tasks?
Opus 5 has become strong at finding software vulnerabilities as a by-product of general capability, even though Anthropic did not train it on cyber tasks. Its classifiers allow source-code vulnerability discovery but block binary-based scanning, penetration testing and exploit generation. Anthropic expects them to intervene about 85% less often than Fable 5's, with flagged requests falling back to Opus 4.8.
What is the ARC-AGI-3 result for Opus 5?
On ARC-AGI-3, a test of novel problem-solving, Opus 5 scores 30.2% - roughly three times the next-best model (GPT-5.6 Sol at 7.8%) and about twenty times Opus 4.8's 1.5%. Anthropic cites it as evidence that Opus 5's gains are in genuine reasoning, not just benchmark familiarity.
AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert Verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.