AI Tools Review

Claude Opus 4.8

By Anthropic

Released: 2026-05-28

LLM
Agents
Coding
Reasoning
Anthropic
Paid
New

Claude Opus 4.8 is Anthropic's June 2026 flagship model, succeeding Opus 4.7. It posts a headline score of 81 on the hardest agentic coding and reasoning suites, holds long-horizon tool-use plans together across far more steps, and is notably more candid about its own uncertainty - refusing to fabricate rather than confidently pressing on. It is the default choice for serious agentic and software-engineering workloads.

Visit Claude Opus 4.8

The SWE-bench Pro leader

Opus 4.8 holds the top score on the hardest widely-run software-engineering benchmark at 69.2% — clear of Claude Sonnet 5 (63.2%), GPT-5.6 Sol (64.6%) and GLM-5.2 (62.1%). Only Fable 5 sits above it.

Honest under uncertainty

The defining behavioural change of the 4.8 release: the model is markedly more candid about what it does not know, refusing to fabricate rather than confidently pressing on. In production agents, that trait is worth benchmark points.

The long-horizon workhorse

Opus 4.8 holds multi-step tool-use plans together across far more steps than its predecessors — the quality that decides whether an autonomous agent finishes a real task or wanders off halfway through.

Claude Opus 4.8 is Anthropic's workhorse flagship: the model that leads SWE-bench Pro, anchors serious Claude Code deployments, and — since Fable 5 arrived above it — has settled into the role it was arguably always built for. It is no longer the most intelligent Claude, and at $1.80 per measured task it is far from the cheapest option anywhere. What it remains is the most dependable heavy-duty engineering model in Anthropic's range, and this review examines whether that is still worth paying for.

Where Opus 4.8 now sits

Claude Opus 4.8 summary card: positioning and best-for list in the campaign index-card style.
At a glance — where this model fits.

Opus 4.8 shipped on 28 May 2026 as the successor to Opus 4.7, and for a few weeks it was the most capable model Anthropic sold. The arrival of Claude Fable 5 — the first occupant of the new Mythos-class tier — changed its job description without changing the model. Opus is no longer the ceiling of the Claude range; it is the flagship of the everyday, the model Anthropic expects serious engineering teams to run when Fable-class pricing is unjustifiable and Sonnet-class capability is not quite enough.

That repositioning suits it. On the Artificial Analysis Intelligence Index v4.1, Opus 4.8 scores 56 — third in the field overall, behind Fable 5's 60 and GPT-5.6 Sol's 59, and comfortably ahead of Sonnet 5's 53. But the aggregate index has never been where Opus makes its case. Its case is a specific one: on the hardest, most economically consequential category of model work — real software engineering on real codebases — it holds the top score of any model priced for production use.

The competitive context has also sharpened around it. OpenAI's GPT-5.6 family landed on 9 July with an outright Coding Agent Index lead for Sol; GLM-5.2 brought open weights within striking distance at a fraction of the price; and Anthropic's own Sonnet 5 closed to within three index points at roughly half the per-task cost. Opus 4.8 in mid-2026 is therefore a model under pressure from every direction — above, below and beside — which is exactly why a clear-eyed look at what it uniquely does well matters.

Capabilities: reliability as the headline feature

Anthropic's own framing of the 4.8 release centred on two qualities that do not reduce to a single benchmark number: longer reliable horizons and better self-assessment. The first is about how many steps of an agentic plan the model can execute before drifting — tool calls, file edits, verification passes, course corrections — and Opus 4.8 holds those long chains together across far more steps than its predecessors. For anyone running autonomous coding agents, this is the property that matters most: an agent that is brilliant for twenty steps and lost by step forty is worse than one that is merely very good for two hundred.

The second quality is candour. Opus 4.8 is notably more honest about its own uncertainty than earlier Claude flagships — it refuses to fabricate rather than confidently pressing on, flags the parts of an answer it is unsure of, and declines tasks it cannot do rather than delivering plausible failure. Anthropic's system card discussion of this trait produced the memorable complaint from some early users that the model is 'too honest' — it will tell you your plan is flawed rather than politely executing it. In production, where a confident wrong answer costs more than an admission of uncertainty, this is a feature, and arguably the release's most underrated one.

The third capability pillar is surgical code editing. Opus 4.8 is unusually precise about minimal diffs — changing what needs to change and leaving the rest of the file alone — which compounds across long agentic sessions where sprawling edits create merge conflicts and review burden. Combined with the 1M-token context window, which holds substantial codebases whole, the model is engineered around a single scenario: a long-running agent doing careful work in a large repository. That scenario is exactly where its benchmark lead lives.

  • Long-horizon reliability: multi-step agentic plans hold together across far more steps
  • Improved self-assessment: refuses to fabricate; flags its own uncertainty
  • Surgical, minimal code edits that reduce review and merge burden
  • 1M-token context window for whole-codebase work

The benchmark picture

The number that defines this model is 69.2% on SWE-bench Pro — the hardest widely-run software-engineering benchmark, and the one that best predicts performance on genuinely difficult, multi-file engineering tasks. That score leads every model priced for production use: GPT-5.6 Sol posts 64.6%, Sonnet 5 63.2%, GLM-5.2 62.1%. Only Claude Fable 5, at 80.3% and $2.75 per task, sits above it. If your workload concentrates at the difficult end of software engineering, this single row of the table is most of the purchasing decision.

The wider picture is competitive rather than dominant. On the Coding Agent Index, measured in the Claude Code harness, Opus 4.8 scores 73 — behind GPT-5.6 Sol's 80 in Codex, Terra's 77 and Fable 5's 77, and ahead of GLM-5.2's 58. On Terminal-Bench 2.1 it posts 74.6%, well behind Sol's 88.8% record — terminal-driven agentic work is now clearly OpenAI's ground. On Humanity's Last Exam with tools it manages 57.9%, a hair ahead of Sonnet 5's 57.4%, and on GDPval-AA v2 knowledge work its 1,615 Elo is fractionally behind Sonnet 5's 1,618 — a symbolic inversion that says the mid-tier model has caught the flagship on the broad middle of professional tasks.

Independent aggregate measurement tells the same story from a different angle. Artificial Analysis's Intelligence Index v4.1 places Opus 4.8 at 56, third overall, at a measured cost of $1.80 per Intelligence Index task. On our AITR Value For Money Index — intelligence divided by cost per task — it scores 31, near the bottom of the field alongside the other premium closed models. The honest summary: one decisive lead on the benchmark that matters most for hard engineering, mid-pack results in the categories rivals have specialised in, and economics that only make sense when that one lead is what you are buying.

Where it sits on the Intelligence Index

Artificial Analysis Intelligence Index v4.1 across every scored frontier model — this model highlighted.

Source: Artificial Analysis (9 July 2026). Interactive — hover any bar. Explore the full benchmarks →

Pricing and the value question

Claude Opus 4.8 specification card: intelligence, coding index, cost per task, API pricing, context window and value score.
The numbers in one card — data from our benchmarks tracker.

Opus 4.8 costs $5 per million input tokens and $25 per million output tokens — unchanged flagship-tier pricing, and straightforward to model. The more informative figure is Artificial Analysis's measured $1.80 per Intelligence Index task, which captures what the model actually costs to run through a standard unit of benchmarked work once its reasoning verbosity is accounted for. That places it below Fable 5's $2.75 but well above GPT-5.6 Sol's $1.04, Sonnet 5's $1.53 and an order of magnitude above the open-weights value tier.

The comparison that will make most budget-holders pause is internal. Claude Sonnet 5 delivers 53 of Opus 4.8's 56 index points at $1.53 per task — and until 31 August 2026 carries introductory per-token pricing of $2/$10 against Opus's $5/$25. On knowledge work, Sonnet actually edges ahead. The gap that survives this comparison is precisely the SWE-bench Pro gap: 69.2% against 63.2%. Whether Opus 4.8 is worth its premium therefore reduces to a single question — does your workload live in the six points between those numbers?

The external comparison is no more comfortable. GPT-5.6 Sol scores higher on the Coding Agent Index at roughly 10% lower per-task coding cost in its own harness, and Grok 4.5 delivers 54 index points — two behind Opus — at $0.31 per task, roughly a sixth of the price. Opus 4.8's value case in July 2026 is real but narrow: it is the strongest production-priced model on the hardest engineering benchmark, deployed inside the most mature agentic harness ecosystem, for teams whose failure costs dwarf their token bills. Outside that description, cheaper options now cover most of the ground.

Cost per Intelligence Index task, in context

What a unit of benchmarked work actually costs across the field. Lower is better.

Source: Artificial Analysis (9 July 2026). Interactive — hover any bar. Explore the full benchmarks →

In practice: the Claude Code anchor

Opus 4.8's natural habitat is Claude Code, and most of its benchmark record was measured there. The pairing matters because agentic coding performance is a property of model-plus-harness, not model alone: the 73 on the Coding Agent Index and the SWE-bench Pro lead both describe Opus operating with Claude Code's tool scaffolding, permission model and long-session ergonomics. Teams already running Claude Code get the model's paper performance more or less as advertised; teams on other stacks should trial before assuming the numbers transfer.

Its behavioural traits compound in long sessions. The honesty-under-uncertainty profile means an Opus-driven agent tends to stop and ask rather than invent an API that does not exist; the surgical-edit discipline keeps diffs reviewable even after hours of autonomous work; and the long-horizon reliability means multi-stage refactors — the migrations, dependency upgrades and cross-service changes that define hard enterprise engineering — complete rather than stall. None of these qualities shows up as a single benchmark row, and all of them show up in whether you trust the agent enough to stop watching it.

The sensible deployment pattern in mid-2026 is tiered. Route routine engineering traffic to Sonnet 5 — the benchmarks say it will handle most of it — and reserve Opus 4.8 for the work that earns its premium: the gnarly multi-service refactor, the unfamiliar codebase, the long dependency chain where a cheaper model's failure costs a day of human debugging. Above it, Fable 5 exists for the problems where even Opus is not enough, at a price that enforces its own discipline.

Limitations and honest caveats

The clearest limitation is that Opus 4.8's leads have narrowed to one. Terminal-Bench 2.1 belongs to GPT-5.6 Sol by fourteen points. The Coding Agent Index belongs to Sol by seven. Knowledge work has been caught — by Anthropic's own cheaper model. What remains is SWE-bench Pro, and while it is arguably the most important single benchmark in the field, a value proposition resting on one number is inherently more fragile than it looks on a pricing page.

The economics warrant equal candour. At $1.80 per measured task and a Value For Money Index of 31, Opus 4.8 is among the most expensive ways to buy intelligence we track. GLM-5.2 delivers 51 index points, open weights and a 58 coding score at $0.32 per task; Grok 4.5 delivers 54 points at $0.31. Neither matches Opus on hard engineering — but both redefine what 'close enough' costs, and any honest assessment must acknowledge that the gap between Opus and the value tier has never bought less margin than it does now.

Two smaller notes. The 'too honest' trait, whatever its production virtues, genuinely frustrates users who want the model to attempt a task despite its own reservations — teams should calibrate prompts and expectations accordingly. And the model's positioning is now explicitly transitional: it is the second-best Claude by design, and Anthropic's release cadence suggests its successor will inherit the workhorse role within months rather than years. Committing to Opus 4.8 is committing to a point on a fast-moving line.

Verdict: who should use it

Claude Opus 4.8 verdict card with our one-line assessment.
The verdict, briefly.

Claude Opus 4.8 is the best production-priced model in the world at the specific thing serious engineering organisations most need: completing hard, real-repository software-engineering work reliably. The 69.2% SWE-bench Pro score is the cleanest expression of that, and the behavioural profile around it — honest uncertainty, surgical edits, long-horizon stability — is what makes the score deliverable in practice rather than just on the leaderboard.

It is the right choice for teams running autonomous or semi-autonomous coding agents against complex production codebases inside Claude Code, for whom the six-point SWE-bench Pro margin over Sonnet 5 translates directly into fewer failed runs and less human rescue work. It is the wrong choice as a general-purpose default in 2026: Sonnet 5 covers the broad middle at half the cost, Sol wins the terminal and beats it on coding-index economics, and the open-weights tier has made 'good enough' nearly free.

Use it, in short, the way its own maker now positions it — as the heavy-duty tier of a routed system, not the answer to every prompt. Deployed that way, Opus 4.8 remains one of the easiest models in the field to justify. Deployed as a default, it is a premium habit that Sonnet 5's launch quietly made optional.

  • Choose it for: hard agentic engineering in Claude Code, complex refactors, high failure-cost work
  • Skip it for: routine coding and knowledge work — Sonnet 5 covers these at roughly half the cost
  • Watch for: the narrowing premium as GPT-5.6, GLM-5.2 and Grok 4.5 compress the value gap

Claude Opus 4.8: benchmark results

SWE-bench Pro69.2%

Leads all production-priced models; Fable 5 (80.3%) sits above at a premium tier

Artificial Analysis Intelligence Index v4.156

Third overall — Fable 5: 60, GPT-5.6 Sol: 59, Sonnet 5: 53

Coding Agent Index (Claude Code)73

GPT-5.6 Sol leads at 80 in Codex

Terminal-Bench 2.174.6%

GPT-5.6 Sol holds the record at 88.8%

Humanity's Last Exam (with tools)57.9%

Sonnet 5: 57.4%

Figures from Anthropic launch tables and Artificial Analysis (Intelligence Index v4.1 and Coding Agent Index, 9 July 2026). Coding Agent Index measured in the Claude Code harness.

Where Claude Opus 4.8 fits

Hard agentic software engineering

The SWE-bench Pro leader among production-priced models, built for multi-file, multi-service engineering work in Claude Code where failed runs cost real engineering hours.

Long-horizon autonomous agents

Multi-stage refactors, migrations and dependency upgrades that unfold across hundreds of steps — the scenario Opus 4.8's longer reliable horizon was specifically built for.

High-stakes review and analysis

Work where the model's honesty under uncertainty pays: security-sensitive changes, architectural reviews and any task where a confident fabrication is costlier than a flagged unknown.

Whole-codebase reasoning

The 1M-token context window holds substantial repositories in a single session, supporting cross-cutting analysis without retrieval pipelines.

The heavy tier of a routed stack

Paired with Sonnet 5 as the default and Fable 5 as the escalation ceiling, Opus 4.8 is the natural middle-heavy tier: reserved for the tasks that earn its $1.80-per-task cost.

Sources & further reading

Anthropic Model Timeline

Claude Fable 5

1M tokens context

Claude Sonnet 5

1M tokens context

Claude Sonnet 5

1M tokens context

Claude Mythos 5
Claude Fable 5

1M tokens context

Claude Opus 4.8Current

Long-context context

Claude Cowork
Anthropic: Claude Opus 4.5

200k tokens context

Anthropic: Claude Haiku 4.5

200k tokens context

Claude 4.5 Haiku

200k tokens context

Anthropic: Claude Sonnet 4.5

1,000k tokens context

Anthropic: Claude Opus 4.1

200k tokens context

Anthropic: Claude Opus 4

200k tokens context

Anthropic: Claude Sonnet 4

1,000k tokens context

Anthropic: Claude 3.7 Sonnet (thinking)

200k tokens context

Anthropic: Claude 3.7 Sonnet

200k tokens context

Anthropic: Claude 3.5 Haiku

200k tokens context

Anthropic: Claude 3.5 Sonnet

200k tokens context

Anthropic: Claude 3 Haiku

200k tokens context

Frequently Asked Questions

Is Claude Opus 4.8 still Anthropic's best model?

No — Claude Fable 5, the first Mythos-class model, sits above it with an Intelligence Index of 60 to Opus 4.8's 56 and an 80.3% SWE-bench Pro score to Opus's 69.2%. But Fable 5 costs $2.75 per measured task against Opus's $1.80 and carries no published per-token pricing, so Opus 4.8 remains the top of Anthropic's conventionally priced range — the workhorse flagship.

How does Opus 4.8 compare with Claude Sonnet 5?

Sonnet 5 has closed most of the gap: 53 vs 56 on the Intelligence Index, 57.4% vs 57.9% on Humanity's Last Exam with tools, and it actually edges ahead on GDPval-AA v2 knowledge work (1,618 vs 1,615 Elo) at roughly half the per-token cost. The gap that remains is hard engineering: Opus leads SWE-bench Pro 69.2% to 63.2%. If your workload lives in those six points, Opus earns its premium; otherwise Sonnet 5 is the rational default.

How much does Claude Opus 4.8 cost?

API pricing is $5 per million input tokens and $25 per million output tokens. Artificial Analysis measures its real-world cost at $1.80 per Intelligence Index task — below Fable 5's $2.75, above GPT-5.6 Sol's $1.04 and Sonnet 5's $1.53. On our AITR Value For Money Index it scores 31, reflecting a premium model priced for capability rather than volume.

Is Opus 4.8 better than GPT-5.6 for coding?

It depends which coding. GPT-5.6 Sol leads the Coding Agent Index (80 in Codex vs Opus's 73 in Claude Code) and holds the Terminal-Bench 2.1 record at 88.8% against Opus's 74.6%. Opus 4.8 leads on SWE-bench Pro, 69.2% to Sol's 64.6% — the benchmark that best represents hard, real-repository engineering. Terminal-heavy agentic work favours Sol; the hardest repository work favours Opus.

What does 'too honest' mean about Opus 4.8?

The 4.8 release made the model markedly more candid about its own uncertainty: it refuses to fabricate, flags what it does not know, and will push back on flawed plans rather than executing them politely. Most production teams consider this a major asset — confident fabrication is the expensive failure mode in autonomous agents — but users who want an attempt regardless of the model's reservations can find the behaviour frustrating.

Specifications

pricingFlagship Opus tier
context WindowLong-context

AI Evaluation

4.9
Expert Rating
Text4.9/5
Coding4.9/5

The strongest model Anthropic has shipped. The longer reliable horizon and improved self-assessment matter more in production than the benchmark - though you only feel the gain if your workflow is built to exploit it.

Pros

  • Best-in-class agentic reliability
  • More honest about uncertainty
  • Surgical code edits

Cons

  • Premium pricing
  • 'Too honest' for users who want an answer no matter what