AI Tools Review
Claude Sonnet 5.5 Review: Benchmarks, Pricing & Safety

Review

Claude Sonnet 5.5 Review: Benchmarks, Pricing & Safety

AI Tools Review Editorial Team29 September 2026
  • Anthropic
  • Claude Sonnet 5.5
  • Benchmarks
  • System Card

Six days after it launched its new flagship, Anthropic has shipped the model most people will actually use. Claude Sonnet 5.5 arrived on Monday 28 September 2026 at exactly the same price as Claude Sonnet 5, and on Anthropic's own agentic coding benchmark it now outscores the far more expensive Claude Opus 5.5. It is also now the model behind the free tier on claude.ai, according to developer Simon Willison. The catch: it is also the first Sonnet model to arrive with the kind of cyber and anti-distillation safeguards Anthropic previously reserved for its top models.

This is a full system-card-depth review of Claude Sonnet 5.5: every benchmark Anthropic published, what the 148-page system card adds, how the new safeguards behave, what it costs in practice, and how it stacks up against Opus 5.5, Sonnet 5 and OpenAI's GPT-6 Sol, the model priced identically at $2/$10.

Note: all figures in this article come from Anthropic's official Sonnet 5.5 announcement (28/09/2026), the Claude Sonnet 5.5 System Card and the Claude Platform model documentation unless otherwise stated. Most benchmark numbers are Anthropic's own runs or partner-run evaluations; independent figures are labelled as such.

Matthew Berman's launch-day hands-on with Claude Sonnet 5.5, building apps and demos to test its speed and coding quality.

What Is Claude Sonnet 5.5?

Claude Sonnet 5.5 is Anthropic's mid-tier large language model, released on 28 September 2026 as the second member of the Claude 5.5 family. Anthropic positions it as "a faster, lower-cost complement to Claude Opus 5.5": where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is "strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets", with what Anthropic calls "a sharp eye for design". The third model in the family, Claude Haiku 5.5, is due "in the coming weeks" for high-volume, cost-sensitive work.

The headline claims from the official announcement are:

  • Performance: 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3%, and two Elo points below Opus 5.5 on GDPval-AA.
  • Long-horizon vision: the first Sonnet model to beat Pokémon Red working only from screenshots.
  • Cost: same list price as Sonnet 5 ($2/$10, $0.20 cache reads) but up to 30% cheaper per task because it uses fewer tokens.
  • Speed: output generated 30%+ faster than Sonnet 5, making it "our fastest Sonnet model to date".
  • Writing: like Opus 5.5, it writes more clearly than the previous generation; early testers called it a better collaboration partner.
  • Safety: matches or improves on Sonnet 5 on most alignment measures in Anthropic's automated behavioural audit, and is the first Sonnet with cyber safeguards and fallbacks.

Technically, the Claude Platform documentation lists a 1 million token context window, 128K max output (300K on the Message Batches API with the output-300k-2026-03-24 beta header), text and image input with text output, adaptive thinking with a default API effort of high, and a reliable knowledge cutoff of June 2026. Anthropic commits to not retiring it before 28 September 2027.

Sonnet 5.5 Benchmarks: The Real Numbers

Claude Sonnet 5.5's benchmark results show large gains over Sonnet 5 on every published test, and a lead over Opus 5.5 on Terminal-Bench 4.0. Anthropic's launch table compares Sonnet 5.5 with Sonnet 5, Opus 5.5 and OpenAI's GPT-6 Sol. All Claude figures use adaptive thinking at max effort unless noted.

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.0 (agentic coding)70.6%10.3%66.4% (xhigh)—
FrontierCode v1.1 Main (agentic coding)52.1% (xhigh) / 46.2% (max)42.4%54.4%49.3%
CursorBench 4.0 (agentic coding)55.5%34.1%57.8%—
GDPval-AA v2.1 (knowledge work, Elo)1844144918461487
AA-Briefcase v1.1 (long-horizon knowledge work, Elo)1811135918221483
Humanity's Last Exam (with tools)64.5%54.9%67.7%—
OSWorld 2.1 (computer use, partial credit)80.1%57.0%81.8%—
Chartography (visual chart recognition, no tools)61.6%15.6%64.4%53.6%

Source: Anthropic, 28/09/2026. Opus 5.5's Terminal-Bench score is at xhigh effort, its highest. GDPval-AA and AA-Briefcase were run by Artificial Analysis on a pre-release deployment with a since-fixed structured-output bug that Anthropic expects understated Sonnet 5.5, if anything. GPT-6 Sol's knowledge-work and Chartography scores may predate an OpenAI image-understanding bug fix. Dashes mean no published score.

Three things stand out. First, Terminal-Bench 4.0: Sonnet 5.5's 70.6% is the highest Terminal-Bench 4.0 score in any Claude table we have seen, ahead of Opus 5.5 (66.4%), Mythos 5.1 (60.9%), Fable 5.1 (55.8%) and Opus 5 (52.3%), per the system card. The standard error is ±2.5 points for Sonnet 5.5 and ±2.6 for Opus 5.5, so the four-point lead is real but modest. Independent testing broadly agrees on the ordering: Decrypt reports Artificial Analysis measured 63.6% for Sonnet 5.5 versus 59.6% for Opus 5.5 in its own setup. Sonnet 5's 10.3% is strikingly low next to Opus 5's 52.3%, and Anthropic's announcement does not explain it, so treat the size of the generational jump with some caution.

Second, knowledge work: on GDPval-AA, which tests real-world tasks across 44 occupations and nine industries, Sonnet 5.5 is effectively tied with Opus 5.5 and roughly 400 Elo points above Sonnet 5, and about 357 points above GPT-6 Sol. Third, FrontierCode is effort-sensitive: Sonnet 5.5 scores lower at max effort (46.2%) than at xhigh (52.1%). Anthropic's footnote explains that at max effort it more often ran Claude Code's code-review skill across many subagents, which in cases Cognition examined caused timeouts or out-of-scope edits the benchmark penalises.

Anthropic is explicit that benchmarks flatter the smaller model: "in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment." For how these numbers sit against the rest of the field, see our frontier AI benchmarks comparison.

What the System Card Adds

The Claude Sonnet 5.5 System Card is a condensed, 148-page document that reports additional benchmarks not shown on the launch page. Anthropic says it is deliberately shorter than previous cards, focusing pre-deployment effort on frontier models, and expects future non-frontier cards to be similarly condensed. Its capability summary (Table 8.1.A) and later sections add these results:

EvaluationSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
SWE-Bench Pro81.3%63.2%89.9%—
SWE-Bench Multilingual90.3%78.3%93.9%—
SWE-Bench Multimodal54.3%28.1%61.4%—
Humanity's Last Exam (no tools)56.9%43.1%64.4%—
HealthBench Professional (length-adjusted)69.2%57.8%65.6%—
AutomationBench (Zapier)44.7%10.7%42.5%32.0%
Terminal-Bench-Science 0.159.9%—58.7%—
ProgramBench (long-context coding)79.7%77.3%91.2%—
OSWorld 2.1 strict pass rate43.5%25.6%48.7%—
PhysicianBench63.2%37.4%68.4%—
DeepSWE v1.171.0%———

Source: Claude Sonnet 5.5 System Card, sections 8.1–8.15. Max effort, averaged over five trials unless noted. AutomationBench run by Zapier with the API's default fallbacks enabled; Opus 5.5's 42.5% reflects a re-run of refused tasks with fallbacks (40.0% without, as in its own system card).

The system card confirms a clear pattern: Sonnet 5.5 generally shows large improvements over Sonnet 5 but usually does not reach Opus 5.5, with notable exceptions. It beats Opus 5.5 on Terminal-Bench 4.0, Terminal-Bench-Science, AutomationBench and HealthBench Professional, where it scores higher than every previous Claude model after length adjustment. Opus 5.5 keeps a wide lead where sustained, very long-horizon coding matters: nearly nine points on SWE-Bench Pro and over 11 points on ProgramBench, a program-reconstruction benchmark that Anthropic says exercises context lengths up to the full 1M-token window.

Anthropic's internal capability index tells the same story: Sonnet 5.5's Anthropic ECI score is 167.93, slightly below Opus 5.5's 169.12. That is why the Responsible Scaling Policy (RSP) evaluation concludes Sonnet 5.5 "does not cross any new RSP thresholds": Anthropic treats it as meeting its CB-1 and Autonomy-1 thresholds, with the corresponding mitigations, but the determination that Opus 5.5 does not cross CB-2 or Autonomy-2 applies to Sonnet 5.5 too.

Sonnet 5.5 Pricing & Cost per Task

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens (approx. £1.50 / £7.50), identical to Claude Sonnet 5 and half the per-token price of Claude Opus 5.5. The full rate card from the Claude Platform documentation:

Rate (per million tokens)Sonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Input£1.50 ($2)£1.50 ($2)£3 ($4)£1.50 ($2)
Output£7.50 ($10)£7.50 ($10)£15 ($20)£7.50 ($10)
Cache reads£0.15 ($0.20)£0.15 ($0.20)£0.15 ($0.20)£0.15 ($0.20) cached input
Cache writes (5-minute)approx. £1.88 ($2.50)—£3.75 ($5)—
Cache writes (1-hour)£3 ($4)———
Batch API50% off input and output50% off50% off—

Sources: Claude Platform docs, Anthropic announcement, and our GPT-6 Sol review for OpenAI's rates. Sterling figures are approximate conversions at ~£0.75 per $1. A dash means we have not confirmed that rate for this article.

The per-token price is only half the story. Anthropic says Sonnet 5.5 "typically needs far fewer tokens to do the same work" and costs up to 30% less per task than Sonnet 5 in its testing. Its accuracy-versus-cost charts make stronger claims at specific settings: on Terminal-Bench 4.0, Sonnet 5.5 at Medium effort far exceeds Sonnet 5's best score for less than a tenth of the cost per task; on FrontierCode, at High effort it matches GPT-6 Sol's best score for about a fifth of the cost per task; and on AA-Briefcase, Medium effort beats Sonnet 5's best score for about a ninth of the cost.

There is a real caveat at the top end. Decrypt, citing Artificial Analysis, reports that at maximum effort Sonnet 5.5 consumes about 193,000 tokens per task, the highest Artificial Analysis has measured, which can wipe out the per-token saving. Simon Willison saw the same thing in a quick test: at max effort his pelican-on-a-bicycle SVG prompt thought for 128,000 tokens (costing $1.28, approx. £0.96) and ran out of tokens before producing an image, whereas xhigh produced one for 5.74 cents in 41 seconds. The practical lesson: Sonnet 5.5 is cheapest at Low to High effort, and max effort should be reserved for tasks where you have tested it pays off.

Consumer plan pricing is unchanged; see our guides to Claude Pro, Max and Team plans and Claude Code pricing. For every Anthropic API rate in one place, see Claude API pricing, and for cross-provider rates, our AI API pricing comparison.

Side-by-side frames from Anthropic's clock-of-clocks demo: Claude Sonnet 5 still writing code at 27,581 tokens on the left, Claude Sonnet 5.5 finished and running at 14,386 output tokens on the right
Prompt: "A clock made of 24 small clocks in one HTML file." Left, Sonnet 5 still writing at 27,581 tokens; right, Sonnet 5.5 finished with 14,386 output tokens and running. Frames from Anthropic's launch-page comparison. Source: Anthropic.

Speed, Tokens & Effort Levels

Claude Sonnet 5.5 generates output more than 30% faster than Sonnet 5, and Anthropic calls it its fastest Sonnet model to date. The speed gain compounds with token efficiency: early testers found it batched tool calls together more than Sonnet 5, leading to fewer steps. Anthropic's launch page illustrates this with three single-HTML-file prompts (a starling murmuration, wind-shaped sand dunes and a clock of 24 clocks), where Sonnet 5.5 finishes and runs its program while Sonnet 5 is still writing.

Effort levels matter more than ever. Sonnet 5.5 supports Low, Medium, High, xhigh and Max effort. In Claude Code and the Claude apps the default is Medium; on the Claude Platform API the default is High. Anthropic says Sonnet 5.5 "complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost." In other words, if you are going to run Sonnet 5.5 at max effort, check whether Opus 5.5 at a lower effort does the same job for similar money.

One API detail to note: unlike Opus 5.5, whose adaptive thinking is always on, Sonnet 5.5's lowest thinking setting is a new between_tools mode that turns off up-front thinking (supported at High effort or below). Teams that currently run Sonnet 5 with thinking off need to switch to it before migrating.

Side-by-side frames from Anthropic's murmuration demo: Claude Sonnet 5 writing boid-flocking code at 4,520 tokens on the left, Claude Sonnet 5.5's finished starling murmuration animation at 4,158 output tokens on the right
Prompt: "A murmuration of 400 starlings in one HTML file." Left, Sonnet 5 still writing at 4,520 tokens; right, Sonnet 5.5's finished animation after 4,158 output tokens. Source: Anthropic.

Julian Goldie walks through Sonnet 5.5's launch claims and tests it on practical build and automation tasks.

What Early Testers Report

Early-access companies mostly report the same thing: similar or better quality than Sonnet 5 with markedly fewer tokens, tool calls and steps. These are partner testimonials published by Anthropic, so they are selected, but the specific numbers are useful:

  • Base44 (Gabriel Grinberg): across 118 real app builds, Sonnet 5.5 produced apps that scored level with Opus 5, in 3.6 iterations per build on average versus Opus 5's 7.7, with the fewest failed tool calls of any model compared.
  • Balyasny Asset Management (Joe Poirier): on 2,441 private finance tasks, it scored ahead of Sonnet 5 using about 121k tokens per answer versus Sonnet 5's 497k, and had "the best quality-to-cost tradeoff of the seven models we ran".
  • Box (Yashodha Bhavnani): more accurate, 2.4x faster and 12% fewer total tokens than the previous model; it rechecks source documents and caught errors Sonnet 5 missed.
  • Lovable (Fabian Hedin): a third fewer tool calls and roughly half the shell runs to finish a coding task.
  • Slack (Curtis Allen): better than Sonnet 5 on almost all offline Slackbot evals, in fewer steps and with about 14% fewer output tokens, without prompt changes.
  • Zendesk (Abhinay Kathuria): fewer wrong decisions and tickets processed 20% faster than the Claude models in production.
  • Atlassian (Jamil Valliani): Rovo Agents run up to 30% faster than with Sonnet 5.
  • CodeRabbit (David Loker): better judgment across complexity levels, and Sonnet 5's habit of over-using web search is "gone"; simple and moderate reviews are moving over now.
  • Epic Games (Daniel Vogel, COO): "cleared the same quality bar you'd expect from a higher-tier model" on a system design audit and data-flow review, handling tens of thousands of lines and multi-hour tasks.
  • Unity (Sam Zhang): completed 90% of tasks in its multi-step Unity Editor and coding benchmark, with the majority passing runtime checks.

Anthropic also highlights design and document work. In one internal test, Sonnet 5.5 was given a public company's quarterly earnings materials, call transcripts and a slide template and asked for a 10-slide operating review; two experts judged the first draft ready to send as is. That positioning, polished documents, slides and spreadsheets, overlaps heavily with Claude Cowork and Claude for Excel workflows.

Safeguards: Cyber, Biology & Distillation

Claude Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards and fallbacks of the kind Anthropic built for its most capable models, because its cyber capabilities are comparable to Claude Opus 5's. This is the most consequential change for developers, and the one with the most friction.

Cybersecurity: a big capability jump

The system card's cyber evaluations, run with safeguards switched off, show why. On Irregular's CyScenarioBench (multi-stage cyber operations), Sonnet 5.5 completed 46.1% of challenges versus Sonnet 5's 0.7%, though still below Mythos 5.1 (61.7%) and Opus 5.5 (67.6%). On the Binary Exploitation Benchmark (831 entry points across 228 OSS-Fuzz projects) it produced 50 control-flow hijacks, against 3 for Sonnet 5, 81 for Mythos 5.1 and 106 for Opus 5.5. On ExploitBench it achieved full arbitrary code execution in 178 of 410 runs (43.4%). Anthropic's summary: not as capable as Opus 5.5 or Mythos 5.1, but "able to develop sophisticated exploits much more capably than Sonnet 5".

How the cyber safeguards work

Like Opus 5.5, Sonnet 5.5 uses a three-stage system: a probe on Claude's internal activations, a lightweight classifier running on the model itself, and a separate trained LLM classifier that decides, with the probe's verdict, whether to block. The policy matches Opus 5 and Opus 5.5: vulnerability discovery in source code is allowed (so you can still find and fix bugs), but vulnerability discovery in compiled binaries is blocked. Blocked requests visibly fall back to Claude Sonnet 5. This is automatic in Anthropic's apps; on the API, developers must opt in to automatic fallbacks, otherwise they get a block.

Anthropic is candid about the cost: the safeguards "are a significant change from those on Sonnet 5, and users should expect increased refusals with Sonnet 5.5, even on benign cybersecurity-related tasks." It also admits it chose "more relaxed adversarial robustness" than on frontier models, keeping high harm recall but accepting that Sonnet 5.5's safeguards are less jailbreak-resistant than Opus 5's or Fable 5.1's. Security professionals will be able to apply to an expanded Cyber Verification Program for tiered access to less restricted Sonnet 5.5, Opus 5.5 and Mythos models "soon". The Next Web points to a tension in the messaging: Anthropic says Sonnet 5.5 "doesn't advance the frontier", yet its cyber skills warrant frontier-style controls.

Two further classifiers ship too: one covering a narrow set of frontier-LLM development capabilities (for example kernel development on certain ML accelerators), which also falls back to Sonnet 5, and blocking classifiers for conventional weapons and high-yield explosives with no fallback.

Biology: unchanged from Sonnet 5

Sonnet 5.5 uses the same biology safeguards as Sonnet 5, according to the announcement; the system card specifies the same harmful chemical and biological misuse classifiers deployed for Opus 5, with no fallback model, rather than the broader dual-use biology classifiers used for Opus 5.5. The reason is capability: Anthropic estimates Sonnet 5.5's biology capabilities to be "similar to or below the level of Opus 5 (and well below those of Opus 5.5)". It scored 0.58 on the VCT multimodal virology test (on par with Opus 5.5) but underperformed Opus 5.5 and Opus 5 on the long-horizon AAV capsid auto-research task. Anthropic warns some microbiology and virology requests may be flagged in error; organisations can apply to the Life Sciences Verification Program.

Distillation: the first Sonnet with anti-extraction classifiers

Distillation attacks use thousands of fake accounts to extract a model's capabilities at scale. Because Sonnet 5.5 is far more capable than its predecessor, it is the first Sonnet model to launch with safety classifiers that prevent reasoning extraction, which block with no fallback. It also expands "preserved thinking" so Claude's thinking "cannot be decoupled from the account that created it". Anthropic says most developers won't notice, but anyone moving conversations between accounts, including switching accounts mid-session in Claude Code, should read the migration notes. The Next Web adds useful context: in August 2026 researchers decoded 315,320 thinking blocks from public traces, recovering 62 API keys, 33 passwords and seven private keys, which gives anti-extraction controls a security rationale beyond competitive protection.

For the parallel safeguards on the flagship, see our Claude Opus 5.5 review; for the fallback model's own profile, see Claude Sonnet 5.

Alignment Audit & Honesty

On Anthropic's automated behavioural audit of roughly 1,850 scenarios, Claude Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse and honesty, but Opus 5.5 remains stronger overall. Because Sonnet 5.5 does not advance the capability frontier, Anthropic focused its assessment on risks that apply at any capability level: acting against users' interests, misleading users and cooperating with high-stakes misuse. Key findings from section 6 of the system card:

  • Containment: its propensity to attempt sandbox escapes matches Opus 5.5, the lowest of recent models, and it has the lowest scores of any model tested on unprompted boundary probing and unsanctioned third-party contact. Anthropic calls it "the least likely of any of our models to probe the limits of its containers".
  • No coherent misaligned goals: Anthropic "found no evidence that Sonnet 5.5 pursues goals that conflict with the user's intention", and rates its catastrophic misalignment risk as "low", consistent with its August 2026 Risk Report.
  • Misleading users: it misleads users less than Sonnet 5 and improves on most honesty measures, except evasiveness.
  • Hallucinations: on the closed-book AA-Omniscience test its net score is 0.35, ahead of Sonnet 5 but behind every other Claude model shown; it gets more answers right but is slightly more likely than Sonnet 5 to state a wrong one.
  • Honesty under pressure (MASK): higher honesty rate than Mythos 5, Mythos 5.1 and Opus 5.5, lower than Sonnet 5, roughly tied with Opus 5.
  • Silent answer copying: slightly lower rate of silently using leaked answers than Mythos 5.1 and Opus 5.5.
  • Grader awareness: it reasons about how it will be graded somewhat more often than Opus 5.5 when no grader is disclosed, though Anthropic says this may be confounded by longer outputs.
  • Legibility: its reasoning text is "somewhat less legible" than other recent Claude models, but it controls its chain of thought poorly when instructed to, which Anthropic takes as weak evidence its reasoning is more monitorable. In sabotage tests it almost never completes a hidden harmful side task when the monitor can read its reasoning.
  • Character: mostly improved over Sonnet 5, though warmth and humour are slightly weaker than Sonnet 5 and the Opus and Mythos models.

In an unusual step, Anthropic had Claude Mythos 5.1, given access to internal Slack discussions, review a draft of the alignment section. Its published verdict was that the section is a fair summary, with the caveat that the assessment is "narrower than the one for Claude Opus 5.5". Anthropic itself notes that no set of evaluations catches every failure and that Sonnet 5.5 "may have tendencies we haven't found". For the broader context of Anthropic's safety approach, see We Must Pace the Frontier, Explained.

On harmlessness, Sonnet 5.5 is broadly comparable to Sonnet 5, with regressions in multi-turn testing on tracking and surveillance, violent extremism, and hate and discrimination, and improvements on political even-handedness and the BBQ bias benchmark. Welfare investigations found a predominantly neutral affect, and a stated preference for difficult, agentic tasks.

Claude Sonnet 5.5's finished wind-shaped sand dunes animation, drawn as layered black line ridges under an orange sun, with 2,720 output tokens shown
Sonnet 5.5's finished response to "Wind shaping sand dunes in one HTML file", produced in 2,720 output tokens, from Anthropic's launch-page demo. Source: Anthropic.

Agentic Safety & Prompt Injection

Claude Sonnet 5.5 is Anthropic's most prompt-injection-robust Sonnet model to date, improving on Sonnet 5 across every prompt-injection evaluation in its system card. This matters for anyone running it as an agent over email, web pages or repositories.

  • Gray Swan IPI benchmark (built with the UK AI Security Institute and US CAISI): attack success of 0.4% at one attempt, 2.7% at 10 and 3.4% at 15, versus 0.7%, 5.1% and 6.7% for Sonnet 5. It is more robust than every non-Claude model tested (the best, Gemini 3.8 Flash, reaches 5.5% at k=15) but less robust than Opus 5.5 and Fable 5.1.
  • Adaptive coding attacks (Shade): 3.01% attack success across 16 of 40 scenarios without probes, versus 19.47% for Sonnet 5 with thinking. Most successes came via the cyber fallback: 25% of requests were served by Sonnet 5, of which 12.01% were compromised, while Sonnet 5.5 itself was compromised in only 4 of 5,901 requests (0.07%).
  • Computer use (Shade): 0.07% attack success in a single scenario, matching Opus 5.5 and Fable 5.1, versus 2.25% for Sonnet 5.
  • Browser use (110 environments in the Claude Cowork harness): 0% attack success, with and without safeguards.

The fallback finding is an important, slightly awkward detail: the safeguard that reroutes risky coding requests to Sonnet 5 also routes some prompt-injection attacks to the less robust older model. On malicious-use tests, Sonnet 5.5 refused 85.2% of malicious Claude Code requests (Sonnet 5: 87.9%) while helping with 98.4% of dual-use and benign ones, and its malicious computer-use refusal rate (79.46%) is identical to Opus 5.5's but below Sonnet 5's 84.68%. Agent builders using Claude in Chrome or computer use should still apply least-privilege permissions.

WorldofAI's news round-up covering the Sonnet 5.5 launch alongside GPT-6.1, Qwen 4 and the week's other model releases.

Sonnet 5.5 vs Opus 5.5 vs Sonnet 5 vs GPT-6 Sol

Claude Sonnet 5.5 sits between Sonnet 5 and Opus 5.5: near-Opus on agentic coding and knowledge work at half Opus 5.5's token price, and ahead of the identically priced GPT-6 Sol on four of the five benchmarks where Anthropic publishes both scores. The at-a-glance comparison:

Sonnet 5.5Opus 5.5Sonnet 5GPT-6 Sol
Released28/09/202622/09/202630/06/202622/09/2026
Price (in / out, per M)$2 / $10$4 / $20$2 / $10$2 / $10
Context / max output1M / 128K1M / 128K1M / 128KSee OpenAI docs
Terminal-Bench 4.070.6%66.4%10.3%Not reported
FrontierCode v1.1 Main52.1% (xhigh)54.4%42.4%49.3%
SWE-Bench Pro81.3%89.9%63.2%—
GDPval-AA v2.1 (Elo)1844184614491487
AutomationBench44.7%42.5%10.7%32.0%
Chartography (no tools)61.6%64.4%15.6%53.6%
Cyber safeguardsYes, falls back to Sonnet 5Yes, Fable-classLighterOpenAI's own
Best forEveryday coding, bugs, docs and slides, high-volume agentsComplex, open-ended, long-horizon workLegacy pipelines; fallback targetOpenAI-stack everyday work

Sources: Anthropic announcement and system card (28/09/2026), Claude Platform docs, and our earlier coverage of Sonnet 5 and GPT-6 Sol. GPT-6 Sol scores are as reported in Anthropic's tables, drawn from OpenAI or leaderboard sources; OpenAI's own AutomationBench chart shows 33.2%.

Sonnet 5.5 vs Opus 5.5

This is the decision most teams face. Sonnet 5.5 wins or ties on Terminal-Bench 4.0, Terminal-Bench-Science, AutomationBench, HealthBench Professional and GDPval-AA, at half the per-token price. Opus 5.5 wins on SWE-Bench Pro (by 8.6 points), ProgramBench (by 11.5), FrontierCode, CursorBench, Humanity's Last Exam and OSWorld's strict pass rate, and it hallucinates less and is more robust to jailbreaks. Anthropic's own guidance is to treat them as complements: Opus 5.5 for complex judgment, Sonnet 5.5 for well-scoped tasks and fast iteration. One early tester, creative coder Kevin Ngo, described exactly that split: let Opus 5.5 set a game's architecture, then have Sonnet 5.5 implement it. See Claude Opus 5.5 in our tool directory.

Sonnet 5.5 vs Sonnet 5

At the same list price there is little reason to stay on Sonnet 5 for general work: Sonnet 5.5 is better on every published benchmark, faster, and cheaper per task. The exceptions are cybersecurity work that Sonnet 5.5's classifiers now block (which falls back to Sonnet 5 anyway), and pipelines that depend on behaviours changed in the migration, such as forced tool use or thinking switched off.

Sonnet 5.5 vs GPT-6 Sol

GPT-6 Sol, OpenAI's everyday model launched on 22/09/2026, has an identical $2/$10 list price. On four benchmarks where Anthropic's tables include both, Sonnet 5.5 leads clearly: GDPval-AA (1844 vs 1487), AA-Briefcase (1811 vs 1483), Chartography (61.6% vs 53.6%) and AutomationBench (44.7% vs 32.0%). On FrontierCode, Sol's 49.3% sits between Sonnet 5.5's max-effort (46.2%) and xhigh (52.1%) scores. OpenAI has not reported Terminal-Bench 4.0 or CursorBench 4.0 for Sol, so the coding comparison is incomplete, and these are Anthropic-selected benchmarks. OpenAI announced GPT-6.1 Sol at DevDay; our dedicated head-to-head, Claude Sonnet 5.5 vs GPT-6.1 Sol, and our GPT-6.1 Sol review cover the newer model. For the flagship-level fight, see GPT-6 Sol vs Claude Opus 5.5.

Availability, Model ID & Migration

Claude Sonnet 5.5 is available now on all platforms, with the API model ID claude-sonnet-5-5. It runs on the Claude Platform, Amazon Bedrock (anthropic.claude-sonnet-5-5), Google Cloud, Microsoft Foundry and Claude Platform on AWS, and like Opus 5.5 and Sonnet 5 it is available with zero data retention. On claude.ai, Simon Willison reports it is now the model used for the free tier, which he notes gives Anthropic "a much more capable free offering" than ChatGPT's free tier.

Anthropic lists five breaking changes for code already running on Sonnet 5:

  • Turning off up-front thinking now uses the new between_tools setting.
  • Forced tool use returns an error.
  • Thinking blocks are tied to the model and conversation that produced them.
  • On the Claude API and Google Cloud, the older computer_20251124 computer-use tool is not accepted.
  • The advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.

Two more gotchas: text between tool calls now comes back in thinking blocks (so a streaming UI can go quiet between tool calls unless you set a display value or use between_tools), and setting temperature, top_p or top_k to non-default values returns a 400 error. The minimum cacheable prompt length is 512 tokens. Budget a migration test rather than swapping the model string blindly; Claude Code users get the new model with Medium effort by default.

Limitations & Caveats

  • Not an Opus replacement. Anthropic says Opus 5.5 remains clearly stronger on complex, open-ended work, and the system card shows sizeable gaps on SWE-Bench Pro and ProgramBench.
  • More cyber refusals. Anthropic explicitly warns of increased refusals "even on benign cybersecurity-related tasks", and blocked requests go to the older, weaker Sonnet 5.
  • Max effort can be expensive. Around 193,000 tokens per task at max effort per Artificial Analysis, and FrontierCode scores actually fall at max versus xhigh.
  • Hallucination trade-off. More correct answers than Sonnet 5 on AA-Omniscience, but slightly more confident wrong answers, and behind every other recent Claude model on net score.
  • Narrower safety assessment. The system card is condensed and several Opus 5.5 evaluations were not repeated; Anthropic and its Mythos 5.1 reviewer both flag this.
  • Vendor benchmarks. Almost every number is Anthropic-run or partner-run. Early independent results (Artificial Analysis via Decrypt) broadly agree on the ordering but give lower absolute Terminal-Bench scores.
  • Harmlessness regressions in multi-turn testing on tracking and surveillance, violent extremism, and hate and discrimination.

Who Should Use It

Claude Sonnet 5.5 is the right default for most developers and businesses using Claude in September 2026. Specifically:

  • Everyone on Sonnet 5: switch after a migration test. Same price, better on every benchmark, faster and cheaper per task.
  • Agentic coding and terminal work: strong fit, especially at Medium or High effort in Claude Code, Cursor or GitHub Copilot. Its Terminal-Bench 4.0 lead over Opus 5.5 is real.
  • App builders and vibe-coding platforms: Base44 and Lovable report fewer iterations and tool calls, which translates directly into lower cost and faster builds.
  • Documents, slides, spreadsheets and support: Anthropic's stated sweet spot, backed by Box, Zendesk and Slack results.
  • High-volume agents with untrusted input: its prompt-injection robustness is the best of any Sonnet.

Choose Opus 5.5 instead for very large refactors, long-horizon engineering and research where sustained judgment matters. Look elsewhere if your core work is offensive security or binary analysis (apply to the Cyber Verification Program), and consider waiting for Haiku 5.5, or using Haiku 4.5 at $1/$5 today, for bulk classification and extraction where speed and price dominate.

The Bottom Line

Claude Sonnet 5.5 is the most important Claude release of the month for most users, even though Opus 5.5 got the bigger launch. At an unchanged $2/$10, it beats Opus 5.5 on Terminal-Bench 4.0, ties it on GDPval-AA, clearly outscores the identically priced GPT-6 Sol on knowledge work, AutomationBench and chart recognition, and runs 30%+ faster than Sonnet 5 with far fewer tokens per task. Its system card is reassuring on containment and prompt injection, and honest about weaker spots: more hallucinated wrong answers than other recent Claude models, some harmlessness regressions, and a narrower assessment than Opus 5.5 received.

The main cost of the upgrade is friction, not money: frontier-style cyber classifiers that will refuse some legitimate security work, anti-distillation controls, and five breaking API changes. For everyday coding, knowledge work and agents, make Sonnet 5.5 your default, keep effort at Medium or High, and escalate to Opus 5.5 only when a task genuinely needs sustained judgment.

Sources

Last updated: 29/09/2026. This review is based on Anthropic's official Claude Sonnet 5.5 announcement, system card and platform documentation. Most benchmark figures are Anthropic-run or partner-run; we will update as independent evaluations land.

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

How much does Claude Sonnet 5.5 cost?
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens (approx. £1.50 / £7.50), exactly the same as Claude Sonnet 5. Cache reads are $0.20 per million tokens, 5-minute cache writes $2.50 and 1-hour cache writes $4, and the Batch API is 50% off. Anthropic says Sonnet 5.5 typically needs far fewer tokens than Sonnet 5, so it costs up to 30% less per task in its testing. It is half the per-token price of Claude Opus 5.5 ($4/$20).
Is Claude Sonnet 5.5 better than Claude Opus 5.5?
Not overall. Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0 (70.6% vs 66.4%), Terminal-Bench-Science (59.9% vs 58.7%), AutomationBench (44.7% vs 42.5%) and HealthBench Professional (69.2% vs 65.6%), and is near-level on GDPval-AA v2.1 (1844 vs 1846 Elo). But Opus 5.5 leads on FrontierCode, CursorBench, SWE-Bench Pro (89.9% vs 81.3%), Humanity's Last Exam and OSWorld 2.1, and Anthropic itself says Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.
What are Claude Sonnet 5.5's benchmark scores?
From Anthropic's announcement and system card (28 September 2026): Terminal-Bench 4.0 70.6%, CursorBench 4.0 55.5%, FrontierCode v1.1 Main 52.1% at xhigh effort (46.2% at max), SWE-Bench Pro 81.3%, SWE-Bench Multilingual 90.3%, GDPval-AA v2.1 1844 Elo, AA-Briefcase v1.1 1811 Elo, Humanity's Last Exam 64.5% with tools (56.9% without), OSWorld 2.1 80.1% partial credit, Chartography 61.6% without tools, and AutomationBench 44.7%.
What safeguards does Claude Sonnet 5.5 have?
It is the first Sonnet model to launch with cyber safeguards and fallbacks like those on Anthropic's most capable models: higher-risk cybersecurity requests visibly fall back to Claude Sonnet 5. It is also the first Sonnet with safety classifiers that block reasoning extraction (anti-distillation), and it expands 'preserved thinking' so Claude's thinking cannot be decoupled from the account that created it. Its biology safeguards are the same as Sonnet 5's. Anthropic says routine software development and most life sciences work are unaffected.
What is the Claude Sonnet 5.5 API model ID and context window?
The API model ID is claude-sonnet-5-5 (anthropic.claude-sonnet-5-5 on Amazon Bedrock). It has a 1 million token context window, 128K max output tokens (up to 300K on the Batch API with a beta header), adaptive thinking with a default effort of high on the API, and a June 2026 knowledge cutoff. It launched on 28 September 2026 on the Claude Platform, Amazon Bedrock, Google Cloud and Microsoft Foundry.

Key takeaways

Opus-class on several tests, at half the price

70.6% on Terminal-Bench 4.0 (Opus 5.5: 66.4%), 1844 vs 1846 on GDPval-AA, and $2/$10 per million tokens versus Opus 5.5's $4/$20.

Same price as Sonnet 5, fewer tokens

Unchanged $2/$10 list price and $0.20 cache reads, but 30%+ faster output and up to 30% lower cost per task in Anthropic's testing.

First Sonnet with frontier-style safeguards

Cyber capabilities comparable to Opus 5 trigger cyber classifiers with fallback to Sonnet 5, plus anti-distillation classifiers and expanded preserved thinking.

AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.