AI Tools Review

Claude Sonnet 5

By Anthropic

Released: 2026-06-30

LLM
Agents
Coding
Anthropic
Value
Paid
New

Claude Sonnet 5 is Anthropic's mid-size default, released 30 June 2026 with a 1M-token context window as standard. It scores 53 on the Intelligence Index - three points off Opus 4.8 - and actually edges the flagship on GDPval knowledge work, at roughly half the per-token cost, with introductory $2/$10 pricing until 31 August 2026.

Visit Claude Sonnet 5

Intelligence Index 53

Scores 53 on Artificial Analysis's Intelligence Index v4.1 at max effort, above GPT-5.6 Luna and GLM-5.2 (both 51), and within three points of Opus 4.8's 56.

Beats Opus on knowledge work

On GDPval-AA v2, Sonnet 5 posts an Elo of 1,618, narrowly ahead of Opus 4.8's 1,615. The first time a Sonnet has edged past the flagship on a headline economic-work benchmark.

Strong value case

At $1.53 per Intelligence Index task and an introductory $2/$10 per million tokens until 31 August 2026, it scores 35 on our Value For Money Index, roughly 40-60% cheaper per token than Opus 4.8.

Claude Sonnet 5 is Anthropic's mid-size model and, for most people, the sensible default. Released on 30 June 2026 with a full system card, it brings a 1M-token context window, the most agentic behaviour of any Sonnet to date, and a benchmark profile that presses uncomfortably close to Opus 4.8, and on GDPval knowledge work, actually passes it. With introductory pricing of $2/$10 per million tokens until the end of August, the value argument is hard to ignore. Here is our full assessment.

What Claude Sonnet 5 is

Claude Sonnet 5 summary card: positioning and best-for list in the campaign index-card style.
At a glance: where this model fits.

Claude Sonnet 5 landed on 30 June 2026, published alongside its system card in Anthropic's usual fashion. The date carried some incidental drama: it was the same day the export order on Fable 5 lifted, which meant Anthropic effectively shared its own news cycle with the return of its larger sibling to unrestricted availability. That timing shaped early coverage, but it should not obscure what Sonnet 5 actually is: the new centre of gravity in Anthropic's line-up, the model most teams will reach for first and, on this showing, the one fewest will need to reach beyond.

In Anthropic's tiering, Sonnet is the mid-size model: bigger and more capable than the economy tier, cheaper and faster to run than the Opus flagship. Sonnet 5 leans into that role with generous plumbing. The context window is 1M tokens: notably, that figure is both the default and the maximum, so there is no reduced-context standard mode with a premium long-context variant hiding behind a different price. Output is capped at 128k tokens in normal operation, which is ample for long reports, large diffs and multi-file code generation, and it can be extended to 300k tokens via a beta header on the Message Batches API for genuinely bulk workloads.

The practical upshot is that Sonnet 5 is built for the way models are actually used in mid-2026: long-running agentic sessions, whole-repository code work and document sets that would have required retrieval gymnastics a couple of years ago. A million tokens of context as standard removes a whole category of engineering workarounds. If the previous Sonnet generation was the model you used when Opus was too expensive, Sonnet 5 makes a stronger claim: it is the model you use because, for most tasks, the flagship no longer offers enough of a margin to justify the premium.

Capabilities: the most agentic Sonnet yet

Anthropic's framing at launch was that this is the most agentic Sonnet it has shipped, and the benchmark evidence supports the claim. On Terminal-Bench 2.1, which measures a model's ability to operate autonomously in a command-line environment, Sonnet 5 scores 80.4%. On OSWorld-Verified, the computer-use benchmark that has models drive real desktop applications, it posts 81.2%. Both are the sort of numbers that were flagship territory not long ago, and both matter far more to real deployments in 2026 than static question-answering scores, because they describe how the model behaves when it is given tools, a goal and room to work.

The more interesting story is how far the gap to Opus 4.8 has narrowed, and where it has inverted. On the Artificial Analysis Intelligence Index, Sonnet 5's 53 sits three points behind Opus 4.8's 56, a visible but no longer decisive margin. On Humanity's Last Exam with tools, the two are practically level: 57.4% against Opus 4.8's 57.9%. And on GDPval-AA v2, the benchmark that grades models on economically valuable knowledge work across real professional tasks, Sonnet 5's Elo of 1,618 actually edges past Opus 4.8's 1,615. It is a narrow lead, but a symbolically significant one.

That GDPval result deserves emphasis because of what the benchmark represents. It is not a puzzle set; it is an attempt to measure the sort of analysis, drafting and structured professional output that businesses actually pay for. A mid-size model beating its own flagship there suggests Anthropic has tuned Sonnet 5 specifically for the workhorse role: the reports, briefs, spreadsheet reasoning and long agentic errands that make up the bulk of enterprise usage. Opus 4.8 retains a clear edge on the hardest coding problems, as the next section shows, but for general knowledge work the honest reading is that Sonnet 5 is no longer the compromise option.

  • Terminal-Bench 2.1: 80.4%, strong autonomous command-line operation
  • OSWorld-Verified: 81.2%, capable real-desktop computer use
  • GDPval-AA v2: 1,618 Elo, narrowly ahead of Opus 4.8 (1,615) on knowledge work
  • Humanity's Last Exam with tools: 57.4%, within half a point of Opus 4.8 (57.9%)

The benchmark picture

Start with the aggregate view. Artificial Analysis's Intelligence Index v4.1 places Sonnet 5 at 53 when run at max effort. That puts it above two notable rivals, GPT-5.6 Luna and GLM-5.2, both on 51, and below the two models above it in Anthropic's own stable: Opus 4.8 on 56 and Fable 5 on 60. In other words, on the day it launched, the main competition for Sonnet 5 came from inside the building. Against the external field at its price point, it leads.

Coding tells a more nuanced story. On SWE-bench Verified, the established measure of resolving real GitHub issues, Sonnet 5 scores 85.2%, a figure that reflects how saturated that benchmark has become across frontier models. The harder SWE-bench Pro is where separation appears: Sonnet 5 manages 63.2% against Opus 4.8's 69.2%. Six percentage points on the toughest coding evaluation is the clearest remaining argument for paying Opus prices. If your workload lives at the difficult end of software engineering (gnarly multi-service refactors, unfamiliar codebases, long dependency chains) the flagship still earns its premium.

Everywhere else, the deltas are small or reversed. Terminal-Bench 2.1 at 80.4% and OSWorld-Verified at 81.2% establish agentic competence; Humanity's Last Exam with tools at 57.4% sits within half a point of Opus 4.8; and GDPval-AA v2 at 1,618 Elo tips over it. The pattern is consistent with a deliberate design choice: concede a measured gap on maximum-difficulty reasoning and coding, and close it, or better, on the broad middle of real-world work. Full comparative tables live on our benchmarks page; the figures here come from Anthropic's launch table and Artificial Analysis, as of July 2026.

Where it sits on the Intelligence Index

Artificial Analysis Intelligence Index v4.1 across every scored frontier model: this model highlighted.

Source: Artificial Analysis (9 July 2026). Interactive, hover any bar. Explore the full benchmarks →

Safety and the system card

Sonnet 5 shipped with its system card on day one, and the document is worth reading rather than skimming. The headline is a broad improvement over Sonnet 4.6 on agentic safety: the dimension that matters most now that models routinely operate tools, browse, and execute multi-step plans with limited supervision. Anthropic reports better refusal of malicious requests, greater resistance to prompt-injection hijacks, and lower rates of both hallucination and sycophancy. For a model explicitly pitched as the most agentic Sonnet yet, those prompt-injection numbers are arguably the single most important line in the card: an agent that can be steered by hostile content in a webpage or document is a liability regardless of how capable it is.

Deployment posture is conservative by current industry standards. Sonnet 5 ships under ASL-3-equivalent protections, with cyber safeguards enabled by default rather than offered as an opt-in. Anthropic also states that the model does not cross its automated AI R&D capability threshold, the internal red line concerning models that could meaningfully accelerate the development of more capable AI systems. For enterprise buyers with governance teams to satisfy, that combination of a published system card, default-on safeguards and an explicit capability-threshold statement is genuinely useful, and remains rarer across the industry than it should be.

One caveat deserves honest treatment: the system card notes that Sonnet 5 shows somewhat higher rates of misaligned behaviour than Opus 4.8. Anthropic's consistent pattern has been that its largest models are also its best-aligned, and Sonnet 5 does not break that pattern. This is not a red flag so much as a parameter for deployment decisions. For most workloads the improvements over Sonnet 4.6 are the operative fact; for the most sensitive autonomous deployments, long-horizon agents with broad permissions and minimal human review, the residual gap to Opus 4.8 is a legitimate input into which model to run.

Pricing and value

Claude Sonnet 5 specification card: intelligence, coding index, cost per task, API pricing, context window and value score.
The numbers in one card: data from our benchmarks tracker.

The launch pricing is aggressive. Until 31 August 2026, Sonnet 5 costs $2 per million input tokens and $10 per million output tokens, an introductory window clearly designed to pull workloads across before the price settles at $3/$15. Even at the standard rate, the comparison with Opus 4.8 at $5/$25 is stark: depending on your input-output mix, Sonnet 5 works out roughly 40-60% cheaper per token than the flagship. During the introductory window the gap is wider still, and teams with migration work to do have a two-month incentive to do it now.

Raw token prices only matter relative to what you get for them, which is why we lean on efficiency metrics. Artificial Analysis puts Sonnet 5's cost per Intelligence Index task at $1.53, the blended cost of completing one unit of its benchmark workload. Feeding that into our own Value For Money Index, which divides measured intelligence by cost per task, Sonnet 5 scores 35. That is the kind of figure that reframes the Opus question entirely: when a model delivers roughly 95% of the flagship's aggregate intelligence score, matches or beats it on knowledge work, and costs a fraction as much to run, the burden of proof shifts to whoever is arguing for the bigger model.

The pricing structure also interacts sensibly with the specifications. Because the 1M-token context is the default rather than a surcharged tier, long-context work is priced on the same simple schedule, and the Message Batches API route to 300k-token outputs gives high-volume users a path to bulk generation. The one thing to plan for is the step change on 1 September 2026, when input rises from $2 to $3 and output from $10 to $15. Budget against the standard rate, treat the introductory window as a bonus, and the economics still hold up comfortably.

Cost per Intelligence Index task, in context

What a unit of benchmarked work actually costs across the field. Lower is better.

Source: Artificial Analysis (9 July 2026). Interactive, hover any bar. Explore the full benchmarks →

Limitations

The clearest limitation is at the hard end of software engineering. The 63.2% on SWE-bench Pro, six points behind Opus 4.8's 69.2%, is not a rounding error: it represents a real class of difficult, long-horizon coding problems where the flagship succeeds and Sonnet 5 does not. Teams doing frontier-difficulty engineering work, or running autonomous coding agents against complex production codebases, should benchmark both models on their own tasks before standardising on the cheaper one. The 85.2% on SWE-bench Verified is reassuring but, like every frontier model's score on that benchmark, tells you less than it used to.

The alignment picture carries its own asterisk. Improvements over Sonnet 4.6 across refusals, prompt-injection resistance, hallucination and sycophancy are all welcome, but the system card is candid that misaligned behaviour rates remain somewhat higher than Opus 4.8's. For a chat deployment this is unlikely to matter much. For a highly autonomous agent with broad tool access, it is a reason either to pay for Opus or to invest in tighter scaffolding (narrower permissions, human checkpoints, output review) around Sonnet 5.

Two smaller practicalities. First, the headline pricing is temporary: the $2/$10 rate expires on 31 August 2026, and anyone modelling costs on the introductory figure will see a 50% rise at the start of September. Second, while the 128k-token output cap is generous, workloads that genuinely need the 300k extension must route through the Message Batches API with a beta header, an asynchronous path that will not suit interactive applications. And within Anthropic's own range, Sonnet 5 is emphatically not the ceiling: Fable 5 sits well above it at 60 on the Intelligence Index, with Opus 4.8 at 56, so users chasing maximum capability regardless of cost still have somewhere else to go.

Verdict: who should use it

Claude Sonnet 5 verdict card with our one-line assessment.
The verdict, briefly.

Sonnet 5 is the easiest recommendation Anthropic has given us in some time. It is the correct default for the broad majority of production workloads: agentic assistants, document and knowledge work, mainstream software engineering, and anything that benefits from a 1M-token context at a mid-tier price. The Intelligence Index gap to Opus 4.8 has shrunk to three points, the knowledge-work benchmark has actually flipped in Sonnet's favour, and the per-token cost is roughly half. For most buyers, the question is no longer whether Sonnet is good enough: it is whether Opus is different enough.

There remain two audiences who should look elsewhere. Teams whose work concentrates at the hardest end of coding, where the six-point SWE-bench Pro gap lives, will still get their money's worth from Opus 4.8, as will anyone running maximally autonomous agents who wants the flagship's lower misaligned-behaviour rates. And organisations chasing peak capability full stop now have Fable 5 back on the table as of the same launch day, sitting at 60 on the Intelligence Index. Sonnet 5 does not compete with either of those on their own terms; it competes on the ratio of capability to cost, and there it currently has no serious rival in Anthropic's range.

Our advice is straightforward. If you are on Sonnet 4.6, upgrade, the capability, safety and context improvements are unambiguous. If you are on Opus 4.8 for general workloads rather than frontier-difficulty ones, run a two-week evaluation on your own tasks during the introductory pricing window; there is a fair chance you halve your bill for no measurable quality loss. And if you are choosing an Anthropic model for the first time, start here. A Value For Money Index of 35 is the number that summarises this release: Sonnet 5 is not the smartest model you can buy in July 2026, but it may well be the smartest way to spend the money.

Claude Sonnet 5: benchmark scores

Artificial Analysis Intelligence Index v4.153

Max effort. Opus 4.8: 56; Fable 5: 60; GPT-5.6 Luna and GLM-5.2: 51.

SWE-bench Pro63.2%

Opus 4.8 leads with 69.2%.

Terminal-Bench 2.180.4%

Autonomous command-line operation.

OSWorld-Verified81.2%

Real-desktop computer use.

GDPval-AA v21,618 Elo

Elo shown against 100-point bar for display. Edges Opus 4.8's 1,615.

Figures from Anthropic launch table and Artificial Analysis, July 2026.

Where Claude Sonnet 5 fits

Agentic assistants and tool-driven workflows

With 80.4% on Terminal-Bench 2.1 and 81.2% on OSWorld-Verified, Sonnet 5 is well suited to agents that operate terminals, browsers and desktop applications, and its improved prompt-injection resistance makes it a safer choice for agents exposed to untrusted content.

Professional knowledge work

Its GDPval-AA v2 Elo of 1,618, narrowly ahead of Opus 4.8, makes it a strong pick for report drafting, analysis, structured business writing and the broad middle of economically valuable white-collar tasks.

Mainstream software engineering

An 85.2% score on SWE-bench Verified covers the bulk of day-to-day engineering: issue resolution, refactors, test writing and code review, at less than half the per-token cost of Opus 4.8.

Long-context document processing

The 1M-token context window (standard, not a premium tier) handles whole codebases, contract sets, discovery bundles and research corpora without retrieval workarounds, with 128k-token outputs for long deliverables.

High-volume batch generation

Via the Message Batches API with a beta header, outputs extend to 300k tokens, making Sonnet 5 practical for bulk report generation, dataset transformation and large-scale content pipelines at introductory rates of $2/$10 per million tokens.

Sources & further reading

Anthropic Model Timeline

Claude Code
Claude Fable 5

1M tokens context

Claude Sonnet 5Current

1M tokens context

Claude Sonnet 5Current

1M tokens context

Claude Mythos 5
Claude Fable 5

1M tokens context

Claude Opus 4.8

Long-context context

Claude Cowork
Anthropic: Claude Opus 4.5

200k tokens context

Anthropic: Claude Haiku 4.5

200k tokens context

Claude 4.5 Haiku

200k tokens context

Anthropic: Claude Sonnet 4.5

1,000k tokens context

Anthropic: Claude Opus 4.1

200k tokens context

Anthropic: Claude Opus 4

200k tokens context

Anthropic: Claude Sonnet 4

1,000k tokens context

Anthropic: Claude 3.7 Sonnet (thinking)

200k tokens context

Anthropic: Claude 3.7 Sonnet

200k tokens context

Anthropic: Claude 3.5 Haiku

200k tokens context

Anthropic: Claude 3.5 Sonnet

200k tokens context

Anthropic: Claude 3 Haiku

200k tokens context

Frequently Asked Questions

How much does Claude Sonnet 5 cost?

Until 31 August 2026, introductory pricing is $2 per million input tokens and $10 per million output tokens. From 1 September 2026 the standard rate is $3/$15. That compares with $5/$25 for Opus 4.8, making Sonnet 5 roughly 40-60% cheaper per token depending on your usage mix.

Is Claude Sonnet 5 better than Opus 4.8?

Mostly no, but the gap is narrow, and on knowledge work it inverts. Opus 4.8 leads on the Intelligence Index (56 vs 53) and clearly on SWE-bench Pro (69.2% vs 63.2%), but Sonnet 5 edges ahead on GDPval-AA v2 (1,618 Elo vs 1,615) and sits within half a point on Humanity's Last Exam with tools (57.4% vs 57.9%), at less than half the price.

How large is Sonnet 5's context window?

1M tokens, which is both the default and the maximum: there is no separate long-context tier. Output is capped at 128k tokens, extendable to 300k tokens through the Message Batches API using a beta header.

Is Claude Sonnet 5 safe for agentic use?

The system card reports clear improvements over Sonnet 4.6: better refusal of malicious requests, more resistance to prompt-injection hijacks, and lower hallucination and sycophancy rates. It ships under ASL-3-equivalent protections with cyber safeguards on by default. One caveat: it shows somewhat higher rates of misaligned behaviour than Opus 4.8, which matters most for highly autonomous deployments.

When was Claude Sonnet 5 released?

On 30 June 2026, accompanied by its full system card, coincidentally the same day the export order on Fable 5 lifted. Introductory pricing runs from launch until 31 August 2026.

Specifications

pricing$2/$10 intro to 31 Aug, then $3/$15 per 1M
context Window1M tokens

AI Evaluation

4.8
Expert Rating
Text4.8/5
Coding4.7/5

The sensible default: most of the Opus experience at roughly half the cost, a 1M context as standard, and the best value in Anthropic's range. Only frontier-difficulty coding still argues for paying more.

Pros

  • Edges Opus 4.8 on GDPval knowledge work
  • 1M context standard, not a premium tier
  • Strong agentic safety profile (ASL-3, default-on safeguards)

Cons

  • Six points behind Opus on SWE-bench Pro
  • Intro pricing expires 31 August 2026
  • Somewhat higher misaligned-behaviour rates than Opus