Six days after it launched its new flagship, Anthropic has shipped the model most people will actually use. Claude Sonnet 5.5 arrived on Monday 28 September 2026 at exactly the same price as Claude Sonnet 5, and on Anthropic's own agentic coding benchmark it now outscores the far more expensive Claude Opus 5.5. It is also now the model behind the free tier on claude.ai, according to developer Simon Willison. The catch: it is also the first Sonnet model to arrive with the kind of cyber and anti-distillation safeguards Anthropic previously reserved for its top models.
This is a full system-card-depth review of Claude Sonnet 5.5: every benchmark Anthropic published, what the 148-page system card adds, how the new safeguards behave, what it costs in practice, and how it stacks up against Opus 5.5, Sonnet 5 and OpenAI's GPT-6 Sol, the model priced identically at $2/$10.
Note: all figures in this article come from Anthropic's official Sonnet 5.5 announcement (28/09/2026), the Claude Sonnet 5.5 System Card and the Claude Platform model documentation unless otherwise stated. Most benchmark numbers are Anthropic's own runs or partner-run evaluations; independent figures are labelled as such.
Matthew Berman's launch-day hands-on with Claude Sonnet 5.5, building apps and demos to test its speed and coding quality.
What Is Claude Sonnet 5.5?
Claude Sonnet 5.5 is Anthropic's mid-tier large language model, released on 28 September 2026 as the second member of the Claude 5.5 family. Anthropic positions it as "a faster, lower-cost complement to Claude Opus 5.5": where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is "strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets", with what Anthropic calls "a sharp eye for design". The third model in the family, Claude Haiku 5.5, is due "in the coming weeks" for high-volume, cost-sensitive work.
The headline claims from the official announcement are:
- Performance: 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3%, and two Elo points below Opus 5.5 on GDPval-AA.
- Long-horizon vision: the first Sonnet model to beat Pokémon Red working only from screenshots.
- Cost: same list price as Sonnet 5 ($2/$10, $0.20 cache reads) but up to 30% cheaper per task because it uses fewer tokens.
- Speed: output generated 30%+ faster than Sonnet 5, making it "our fastest Sonnet model to date".
- Writing: like Opus 5.5, it writes more clearly than the previous generation; early testers called it a better collaboration partner.
- Safety: matches or improves on Sonnet 5 on most alignment measures in Anthropic's automated behavioural audit, and is the first Sonnet with cyber safeguards and fallbacks.
Technically, the Claude Platform documentation lists a 1 million token context window, 128K max output (300K on the Message Batches API with the output-300k-2026-03-24 beta header), text and image input with text output, adaptive thinking with a default API effort of high, and a reliable knowledge cutoff of June 2026. Anthropic commits to not retiring it before 28 September 2027.
Sonnet 5.5 Benchmarks: The Real Numbers
Claude Sonnet 5.5's benchmark results show large gains over Sonnet 5 on every published test, and a lead over Opus 5.5 on Terminal-Bench 4.0. Anthropic's launch table compares Sonnet 5.5 with Sonnet 5, Opus 5.5 and OpenAI's GPT-6 Sol. All Claude figures use adaptive thinking at max effort unless noted.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% (xhigh) | — |
| FrontierCode v1.1 Main (agentic coding) | 52.1% (xhigh) / 46.2% (max) | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 (agentic coding) | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 (knowledge work, Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (long-horizon knowledge work, Elo) | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 (computer use, partial credit) | 80.1% | 57.0% | 81.8% | — |
| Chartography (visual chart recognition, no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
Source: Anthropic, 28/09/2026. Opus 5.5's Terminal-Bench score is at xhigh effort, its highest. GDPval-AA and AA-Briefcase were run by Artificial Analysis on a pre-release deployment with a since-fixed structured-output bug that Anthropic expects understated Sonnet 5.5, if anything. GPT-6 Sol's knowledge-work and Chartography scores may predate an OpenAI image-understanding bug fix. Dashes mean no published score.
Three things stand out. First, Terminal-Bench 4.0: Sonnet 5.5's 70.6% is the highest Terminal-Bench 4.0 score in any Claude table we have seen, ahead of Opus 5.5 (66.4%), Mythos 5.1 (60.9%), Fable 5.1 (55.8%) and Opus 5 (52.3%), per the system card. The standard error is ±2.5 points for Sonnet 5.5 and ±2.6 for Opus 5.5, so the four-point lead is real but modest. Independent testing broadly agrees on the ordering: Decrypt reports Artificial Analysis measured 63.6% for Sonnet 5.5 versus 59.6% for Opus 5.5 in its own setup. Sonnet 5's 10.3% is strikingly low next to Opus 5's 52.3%, and Anthropic's announcement does not explain it, so treat the size of the generational jump with some caution.
Second, knowledge work: on GDPval-AA, which tests real-world tasks across 44 occupations and nine industries, Sonnet 5.5 is effectively tied with Opus 5.5 and roughly 400 Elo points above Sonnet 5, and about 357 points above GPT-6 Sol. Third, FrontierCode is effort-sensitive: Sonnet 5.5 scores lower at max effort (46.2%) than at xhigh (52.1%). Anthropic's footnote explains that at max effort it more often ran Claude Code's code-review skill across many subagents, which in cases Cognition examined caused timeouts or out-of-scope edits the benchmark penalises.
Anthropic is explicit that benchmarks flatter the smaller model: "in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment." For how these numbers sit against the rest of the field, see our frontier AI benchmarks comparison.
What the System Card Adds
The Claude Sonnet 5.5 System Card is a condensed, 148-page document that reports additional benchmarks not shown on the launch page. Anthropic says it is deliberately shorter than previous cards, focusing pre-deployment effort on frontier models, and expects future non-frontier cards to be similarly condensed. Its capability summary (Table 8.1.A) and later sections add these results:
| Evaluation | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| SWE-Bench Pro | 81.3% | 63.2% | 89.9% | — |
| SWE-Bench Multilingual | 90.3% | 78.3% | 93.9% | — |
| SWE-Bench Multimodal | 54.3% | 28.1% | 61.4% | — |
| Humanity's Last Exam (no tools) | 56.9% | 43.1% | 64.4% | — |
| HealthBench Professional (length-adjusted) | 69.2% | 57.8% | 65.6% | — |
| AutomationBench (Zapier) | 44.7% | 10.7% | 42.5% | 32.0% |
| Terminal-Bench-Science 0.1 | 59.9% | — | 58.7% | — |
| ProgramBench (long-context coding) | 79.7% | 77.3% | 91.2% | — |
| OSWorld 2.1 strict pass rate | 43.5% | 25.6% | 48.7% | — |
| PhysicianBench | 63.2% | 37.4% | 68.4% | — |
| DeepSWE v1.1 | 71.0% | — | — | — |
Source: Claude Sonnet 5.5 System Card, sections 8.1–8.15. Max effort, averaged over five trials unless noted. AutomationBench run by Zapier with the API's default fallbacks enabled; Opus 5.5's 42.5% reflects a re-run of refused tasks with fallbacks (40.0% without, as in its own system card).
The system card confirms a clear pattern: Sonnet 5.5 generally shows large improvements over Sonnet 5 but usually does not reach Opus 5.5, with notable exceptions. It beats Opus 5.5 on Terminal-Bench 4.0, Terminal-Bench-Science, AutomationBench and HealthBench Professional, where it scores higher than every previous Claude model after length adjustment. Opus 5.5 keeps a wide lead where sustained, very long-horizon coding matters: nearly nine points on SWE-Bench Pro and over 11 points on ProgramBench, a program-reconstruction benchmark that Anthropic says exercises context lengths up to the full 1M-token window.
Anthropic's internal capability index tells the same story: Sonnet 5.5's Anthropic ECI score is 167.93, slightly below Opus 5.5's 169.12. That is why the Responsible Scaling Policy (RSP) evaluation concludes Sonnet 5.5 "does not cross any new RSP thresholds": Anthropic treats it as meeting its CB-1 and Autonomy-1 thresholds, with the corresponding mitigations, but the determination that Opus 5.5 does not cross CB-2 or Autonomy-2 applies to Sonnet 5.5 too.
Sonnet 5.5 Pricing & Cost per Task
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens (approx. £1.50 / £7.50), identical to Claude Sonnet 5 and half the per-token price of Claude Opus 5.5. The full rate card from the Claude Platform documentation:
| Rate (per million tokens) | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Input | £1.50 ($2) | £1.50 ($2) | £3 ($4) | £1.50 ($2) |
| Output | £7.50 ($10) | £7.50 ($10) | £15 ($20) | £7.50 ($10) |
| Cache reads | £0.15 ($0.20) | £0.15 ($0.20) | £0.15 ($0.20) | £0.15 ($0.20) cached input |
| Cache writes (5-minute) | approx. £1.88 ($2.50) | — | £3.75 ($5) | — |
| Cache writes (1-hour) | £3 ($4) | — | — | — |
| Batch API | 50% off input and output | 50% off | 50% off | — |
Sources: Claude Platform docs, Anthropic announcement, and our GPT-6 Sol review for OpenAI's rates. Sterling figures are approximate conversions at ~£0.75 per $1. A dash means we have not confirmed that rate for this article.
The per-token price is only half the story. Anthropic says Sonnet 5.5 "typically needs far fewer tokens to do the same work" and costs up to 30% less per task than Sonnet 5 in its testing. Its accuracy-versus-cost charts make stronger claims at specific settings: on Terminal-Bench 4.0, Sonnet 5.5 at Medium effort far exceeds Sonnet 5's best score for less than a tenth of the cost per task; on FrontierCode, at High effort it matches GPT-6 Sol's best score for about a fifth of the cost per task; and on AA-Briefcase, Medium effort beats Sonnet 5's best score for about a ninth of the cost.
There is a real caveat at the top end. Decrypt, citing Artificial Analysis, reports that at maximum effort Sonnet 5.5 consumes about 193,000 tokens per task, the highest Artificial Analysis has measured, which can wipe out the per-token saving. Simon Willison saw the same thing in a quick test: at max effort his pelican-on-a-bicycle SVG prompt thought for 128,000 tokens (costing $1.28, approx. £0.96) and ran out of tokens before producing an image, whereas xhigh produced one for 5.74 cents in 41 seconds. The practical lesson: Sonnet 5.5 is cheapest at Low to High effort, and max effort should be reserved for tasks where you have tested it pays off.
Consumer plan pricing is unchanged; see our guides to Claude Pro, Max and Team plans and Claude Code pricing. For every Anthropic API rate in one place, see Claude API pricing, and for cross-provider rates, our AI API pricing comparison.

Speed, Tokens & Effort Levels
Claude Sonnet 5.5 generates output more than 30% faster than Sonnet 5, and Anthropic calls it its fastest Sonnet model to date. The speed gain compounds with token efficiency: early testers found it batched tool calls together more than Sonnet 5, leading to fewer steps. Anthropic's launch page illustrates this with three single-HTML-file prompts (a starling murmuration, wind-shaped sand dunes and a clock of 24 clocks), where Sonnet 5.5 finishes and runs its program while Sonnet 5 is still writing.
Effort levels matter more than ever. Sonnet 5.5 supports Low, Medium, High, xhigh and Max effort. In Claude Code and the Claude apps the default is Medium; on the Claude Platform API the default is High. Anthropic says Sonnet 5.5 "complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost." In other words, if you are going to run Sonnet 5.5 at max effort, check whether Opus 5.5 at a lower effort does the same job for similar money.
One API detail to note: unlike Opus 5.5, whose adaptive thinking is always on, Sonnet 5.5's lowest thinking setting is a new between_tools mode that turns off up-front thinking (supported at High effort or below). Teams that currently run Sonnet 5 with thinking off need to switch to it before migrating.

Julian Goldie walks through Sonnet 5.5's launch claims and tests it on practical build and automation tasks.
What Early Testers Report
Early-access companies mostly report the same thing: similar or better quality than Sonnet 5 with markedly fewer tokens, tool calls and steps. These are partner testimonials published by Anthropic, so they are selected, but the specific numbers are useful:
- Base44 (Gabriel Grinberg): across 118 real app builds, Sonnet 5.5 produced apps that scored level with Opus 5, in 3.6 iterations per build on average versus Opus 5's 7.7, with the fewest failed tool calls of any model compared.
- Balyasny Asset Management (Joe Poirier): on 2,441 private finance tasks, it scored ahead of Sonnet 5 using about 121k tokens per answer versus Sonnet 5's 497k, and had "the best quality-to-cost tradeoff of the seven models we ran".
- Box (Yashodha Bhavnani): more accurate, 2.4x faster and 12% fewer total tokens than the previous model; it rechecks source documents and caught errors Sonnet 5 missed.
- Lovable (Fabian Hedin): a third fewer tool calls and roughly half the shell runs to finish a coding task.
- Slack (Curtis Allen): better than Sonnet 5 on almost all offline Slackbot evals, in fewer steps and with about 14% fewer output tokens, without prompt changes.
- Zendesk (Abhinay Kathuria): fewer wrong decisions and tickets processed 20% faster than the Claude models in production.
- Atlassian (Jamil Valliani): Rovo Agents run up to 30% faster than with Sonnet 5.
- CodeRabbit (David Loker): better judgment across complexity levels, and Sonnet 5's habit of over-using web search is "gone"; simple and moderate reviews are moving over now.
- Epic Games (Daniel Vogel, COO): "cleared the same quality bar you'd expect from a higher-tier model" on a system design audit and data-flow review, handling tens of thousands of lines and multi-hour tasks.
- Unity (Sam Zhang): completed 90% of tasks in its multi-step Unity Editor and coding benchmark, with the majority passing runtime checks.
Anthropic also highlights design and document work. In one internal test, Sonnet 5.5 was given a public company's quarterly earnings materials, call transcripts and a slide template and asked for a 10-slide operating review; two experts judged the first draft ready to send as is. That positioning, polished documents, slides and spreadsheets, overlaps heavily with Claude Cowork and Claude for Excel workflows.
Safeguards: Cyber, Biology & Distillation
Claude Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards and fallbacks of the kind Anthropic built for its most capable models, because its cyber capabilities are comparable to Claude Opus 5's. This is the most consequential change for developers, and the one with the most friction.
Cybersecurity: a big capability jump
The system card's cyber evaluations, run with safeguards switched off, show why. On Irregular's CyScenarioBench (multi-stage cyber operations), Sonnet 5.5 completed 46.1% of challenges versus Sonnet 5's 0.7%, though still below Mythos 5.1 (61.7%) and Opus 5.5 (67.6%). On the Binary Exploitation Benchmark (831 entry points across 228 OSS-Fuzz projects) it produced 50 control-flow hijacks, against 3 for Sonnet 5, 81 for Mythos 5.1 and 106 for Opus 5.5. On ExploitBench it achieved full arbitrary code execution in 178 of 410 runs (43.4%). Anthropic's summary: not as capable as Opus 5.5 or Mythos 5.1, but "able to develop sophisticated exploits much more capably than Sonnet 5".
How the cyber safeguards work
Like Opus 5.5, Sonnet 5.5 uses a three-stage system: a probe on Claude's internal activations, a lightweight classifier running on the model itself, and a separate trained LLM classifier that decides, with the probe's verdict, whether to block. The policy matches Opus 5 and Opus 5.5: vulnerability discovery in source code is allowed (so you can still find and fix bugs), but vulnerability discovery in compiled binaries is blocked. Blocked requests visibly fall back to Claude Sonnet 5. This is automatic in Anthropic's apps; on the API, developers must opt in to automatic fallbacks, otherwise they get a block.
Anthropic is candid about the cost: the safeguards "are a significant change from those on Sonnet 5, and users should expect increased refusals with Sonnet 5.5, even on benign cybersecurity-related tasks." It also admits it chose "more relaxed adversarial robustness" than on frontier models, keeping high harm recall but accepting that Sonnet 5.5's safeguards are less jailbreak-resistant than Opus 5's or Fable 5.1's. Security professionals will be able to apply to an expanded Cyber Verification Program for tiered access to less restricted Sonnet 5.5, Opus 5.5 and Mythos models "soon". The Next Web points to a tension in the messaging: Anthropic says Sonnet 5.5 "doesn't advance the frontier", yet its cyber skills warrant frontier-style controls.
Two further classifiers ship too: one covering a narrow set of frontier-LLM development capabilities (for example kernel development on certain ML accelerators), which also falls back to Sonnet 5, and blocking classifiers for conventional weapons and high-yield explosives with no fallback.
Biology: unchanged from Sonnet 5
Sonnet 5.5 uses the same biology safeguards as Sonnet 5, according to the announcement; the system card specifies the same harmful chemical and biological misuse classifiers deployed for Opus 5, with no fallback model, rather than the broader dual-use biology classifiers used for Opus 5.5. The reason is capability: Anthropic estimates Sonnet 5.5's biology capabilities to be "similar to or below the level of Opus 5 (and well below those of Opus 5.5)". It scored 0.58 on the VCT multimodal virology test (on par with Opus 5.5) but underperformed Opus 5.5 and Opus 5 on the long-horizon AAV capsid auto-research task. Anthropic warns some microbiology and virology requests may be flagged in error; organisations can apply to the Life Sciences Verification Program.
Distillation: the first Sonnet with anti-extraction classifiers
Distillation attacks use thousands of fake accounts to extract a model's capabilities at scale. Because Sonnet 5.5 is far more capable than its predecessor, it is the first Sonnet model to launch with safety classifiers that prevent reasoning extraction, which block with no fallback. It also expands "preserved thinking" so Claude's thinking "cannot be decoupled from the account that created it". Anthropic says most developers won't notice, but anyone moving conversations between accounts, including switching accounts mid-session in Claude Code, should read the migration notes. The Next Web adds useful context: in August 2026 researchers decoded 315,320 thinking blocks from public traces, recovering 62 API keys, 33 passwords and seven private keys, which gives anti-extraction controls a security rationale beyond competitive protection.
For the parallel safeguards on the flagship, see our Claude Opus 5.5 review; for the fallback model's own profile, see Claude Sonnet 5.
Alignment Audit & Honesty
On Anthropic's automated behavioural audit of roughly 1,850 scenarios, Claude Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse and honesty, but Opus 5.5 remains stronger overall. Because Sonnet 5.5 does not advance the capability frontier, Anthropic focused its assessment on risks that apply at any capability level: acting against users' interests, misleading users and cooperating with high-stakes misuse. Key findings from section 6 of the system card:
- Containment: its propensity to attempt sandbox escapes matches Opus 5.5, the lowest of recent models, and it has the lowest scores of any model tested on unprompted boundary probing and unsanctioned third-party contact. Anthropic calls it "the least likely of any of our models to probe the limits of its containers".
- No coherent misaligned goals: Anthropic "found no evidence that Sonnet 5.5 pursues goals that conflict with the user's intention", and rates its catastrophic misalignment risk as "low", consistent with its August 2026 Risk Report.
- Misleading users: it misleads users less than Sonnet 5 and improves on most honesty measures, except evasiveness.
- Hallucinations: on the closed-book AA-Omniscience test its net score is 0.35, ahead of Sonnet 5 but behind every other Claude model shown; it gets more answers right but is slightly more likely than Sonnet 5 to state a wrong one.
- Honesty under pressure (MASK): higher honesty rate than Mythos 5, Mythos 5.1 and Opus 5.5, lower than Sonnet 5, roughly tied with Opus 5.
- Silent answer copying: slightly lower rate of silently using leaked answers than Mythos 5.1 and Opus 5.5.
- Grader awareness: it reasons about how it will be graded somewhat more often than Opus 5.5 when no grader is disclosed, though Anthropic says this may be confounded by longer outputs.
- Legibility: its reasoning text is "somewhat less legible" than other recent Claude models, but it controls its chain of thought poorly when instructed to, which Anthropic takes as weak evidence its reasoning is more monitorable. In sabotage tests it almost never completes a hidden harmful side task when the monitor can read its reasoning.
- Character: mostly improved over Sonnet 5, though warmth and humour are slightly weaker than Sonnet 5 and the Opus and Mythos models.
In an unusual step, Anthropic had Claude Mythos 5.1, given access to internal Slack discussions, review a draft of the alignment section. Its published verdict was that the section is a fair summary, with the caveat that the assessment is "narrower than the one for Claude Opus 5.5". Anthropic itself notes that no set of evaluations catches every failure and that Sonnet 5.5 "may have tendencies we haven't found". For the broader context of Anthropic's safety approach, see We Must Pace the Frontier, Explained.
On harmlessness, Sonnet 5.5 is broadly comparable to Sonnet 5, with regressions in multi-turn testing on tracking and surveillance, violent extremism, and hate and discrimination, and improvements on political even-handedness and the BBQ bias benchmark. Welfare investigations found a predominantly neutral affect, and a stated preference for difficult, agentic tasks.

Agentic Safety & Prompt Injection
Claude Sonnet 5.5 is Anthropic's most prompt-injection-robust Sonnet model to date, improving on Sonnet 5 across every prompt-injection evaluation in its system card. This matters for anyone running it as an agent over email, web pages or repositories.
- Gray Swan IPI benchmark (built with the UK AI Security Institute and US CAISI): attack success of 0.4% at one attempt, 2.7% at 10 and 3.4% at 15, versus 0.7%, 5.1% and 6.7% for Sonnet 5. It is more robust than every non-Claude model tested (the best, Gemini 3.8 Flash, reaches 5.5% at k=15) but less robust than Opus 5.5 and Fable 5.1.
- Adaptive coding attacks (Shade): 3.01% attack success across 16 of 40 scenarios without probes, versus 19.47% for Sonnet 5 with thinking. Most successes came via the cyber fallback: 25% of requests were served by Sonnet 5, of which 12.01% were compromised, while Sonnet 5.5 itself was compromised in only 4 of 5,901 requests (0.07%).
- Computer use (Shade): 0.07% attack success in a single scenario, matching Opus 5.5 and Fable 5.1, versus 2.25% for Sonnet 5.
- Browser use (110 environments in the Claude Cowork harness): 0% attack success, with and without safeguards.
The fallback finding is an important, slightly awkward detail: the safeguard that reroutes risky coding requests to Sonnet 5 also routes some prompt-injection attacks to the less robust older model. On malicious-use tests, Sonnet 5.5 refused 85.2% of malicious Claude Code requests (Sonnet 5: 87.9%) while helping with 98.4% of dual-use and benign ones, and its malicious computer-use refusal rate (79.46%) is identical to Opus 5.5's but below Sonnet 5's 84.68%. Agent builders using Claude in Chrome or computer use should still apply least-privilege permissions.
WorldofAI's news round-up covering the Sonnet 5.5 launch alongside GPT-6.1, Qwen 4 and the week's other model releases.
Sonnet 5.5 vs Opus 5.5 vs Sonnet 5 vs GPT-6 Sol
Claude Sonnet 5.5 sits between Sonnet 5 and Opus 5.5: near-Opus on agentic coding and knowledge work at half Opus 5.5's token price, and ahead of the identically priced GPT-6 Sol on four of the five benchmarks where Anthropic publishes both scores. The at-a-glance comparison:
| Sonnet 5.5 | Opus 5.5 | Sonnet 5 | GPT-6 Sol | |
|---|---|---|---|---|
| Released | 28/09/2026 | 22/09/2026 | 30/06/2026 | 22/09/2026 |
| Price (in / out, per M) | $2 / $10 | $4 / $20 | $2 / $10 | $2 / $10 |
| Context / max output | 1M / 128K | 1M / 128K | 1M / 128K | See OpenAI docs |
| Terminal-Bench 4.0 | 70.6% | 66.4% | 10.3% | Not reported |
| FrontierCode v1.1 Main | 52.1% (xhigh) | 54.4% | 42.4% | 49.3% |
| SWE-Bench Pro | 81.3% | 89.9% | 63.2% | — |
| GDPval-AA v2.1 (Elo) | 1844 | 1846 | 1449 | 1487 |
| AutomationBench | 44.7% | 42.5% | 10.7% | 32.0% |
| Chartography (no tools) | 61.6% | 64.4% | 15.6% | 53.6% |
| Cyber safeguards | Yes, falls back to Sonnet 5 | Yes, Fable-class | Lighter | OpenAI's own |
| Best for | Everyday coding, bugs, docs and slides, high-volume agents | Complex, open-ended, long-horizon work | Legacy pipelines; fallback target | OpenAI-stack everyday work |
Sources: Anthropic announcement and system card (28/09/2026), Claude Platform docs, and our earlier coverage of Sonnet 5 and GPT-6 Sol. GPT-6 Sol scores are as reported in Anthropic's tables, drawn from OpenAI or leaderboard sources; OpenAI's own AutomationBench chart shows 33.2%.
Sonnet 5.5 vs Opus 5.5
This is the decision most teams face. Sonnet 5.5 wins or ties on Terminal-Bench 4.0, Terminal-Bench-Science, AutomationBench, HealthBench Professional and GDPval-AA, at half the per-token price. Opus 5.5 wins on SWE-Bench Pro (by 8.6 points), ProgramBench (by 11.5), FrontierCode, CursorBench, Humanity's Last Exam and OSWorld's strict pass rate, and it hallucinates less and is more robust to jailbreaks. Anthropic's own guidance is to treat them as complements: Opus 5.5 for complex judgment, Sonnet 5.5 for well-scoped tasks and fast iteration. One early tester, creative coder Kevin Ngo, described exactly that split: let Opus 5.5 set a game's architecture, then have Sonnet 5.5 implement it. See Claude Opus 5.5 in our tool directory.
Sonnet 5.5 vs Sonnet 5
At the same list price there is little reason to stay on Sonnet 5 for general work: Sonnet 5.5 is better on every published benchmark, faster, and cheaper per task. The exceptions are cybersecurity work that Sonnet 5.5's classifiers now block (which falls back to Sonnet 5 anyway), and pipelines that depend on behaviours changed in the migration, such as forced tool use or thinking switched off.
Sonnet 5.5 vs GPT-6 Sol
GPT-6 Sol, OpenAI's everyday model launched on 22/09/2026, has an identical $2/$10 list price. On four benchmarks where Anthropic's tables include both, Sonnet 5.5 leads clearly: GDPval-AA (1844 vs 1487), AA-Briefcase (1811 vs 1483), Chartography (61.6% vs 53.6%) and AutomationBench (44.7% vs 32.0%). On FrontierCode, Sol's 49.3% sits between Sonnet 5.5's max-effort (46.2%) and xhigh (52.1%) scores. OpenAI has not reported Terminal-Bench 4.0 or CursorBench 4.0 for Sol, so the coding comparison is incomplete, and these are Anthropic-selected benchmarks. OpenAI announced GPT-6.1 Sol at DevDay; our dedicated head-to-head, Claude Sonnet 5.5 vs GPT-6.1 Sol, and our GPT-6.1 Sol review cover the newer model. For the flagship-level fight, see GPT-6 Sol vs Claude Opus 5.5.
Availability, Model ID & Migration
Claude Sonnet 5.5 is available now on all platforms, with the API model ID claude-sonnet-5-5. It runs on the Claude Platform, Amazon Bedrock (anthropic.claude-sonnet-5-5), Google Cloud, Microsoft Foundry and Claude Platform on AWS, and like Opus 5.5 and Sonnet 5 it is available with zero data retention. On claude.ai, Simon Willison reports it is now the model used for the free tier, which he notes gives Anthropic "a much more capable free offering" than ChatGPT's free tier.
Anthropic lists five breaking changes for code already running on Sonnet 5:
- Turning off up-front thinking now uses the new
between_toolssetting. - Forced tool use returns an error.
- Thinking blocks are tied to the model and conversation that produced them.
- On the Claude API and Google Cloud, the older
computer_20251124computer-use tool is not accepted. - The advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.
Two more gotchas: text between tool calls now comes back in thinking blocks (so a streaming UI can go quiet between tool calls unless you set a display value or use between_tools), and setting temperature, top_p or top_k to non-default values returns a 400 error. The minimum cacheable prompt length is 512 tokens. Budget a migration test rather than swapping the model string blindly; Claude Code users get the new model with Medium effort by default.
Limitations & Caveats
- Not an Opus replacement. Anthropic says Opus 5.5 remains clearly stronger on complex, open-ended work, and the system card shows sizeable gaps on SWE-Bench Pro and ProgramBench.
- More cyber refusals. Anthropic explicitly warns of increased refusals "even on benign cybersecurity-related tasks", and blocked requests go to the older, weaker Sonnet 5.
- Max effort can be expensive. Around 193,000 tokens per task at max effort per Artificial Analysis, and FrontierCode scores actually fall at max versus xhigh.
- Hallucination trade-off. More correct answers than Sonnet 5 on AA-Omniscience, but slightly more confident wrong answers, and behind every other recent Claude model on net score.
- Narrower safety assessment. The system card is condensed and several Opus 5.5 evaluations were not repeated; Anthropic and its Mythos 5.1 reviewer both flag this.
- Vendor benchmarks. Almost every number is Anthropic-run or partner-run. Early independent results (Artificial Analysis via Decrypt) broadly agree on the ordering but give lower absolute Terminal-Bench scores.
- Harmlessness regressions in multi-turn testing on tracking and surveillance, violent extremism, and hate and discrimination.
Who Should Use It
Claude Sonnet 5.5 is the right default for most developers and businesses using Claude in September 2026. Specifically:
- Everyone on Sonnet 5: switch after a migration test. Same price, better on every benchmark, faster and cheaper per task.
- Agentic coding and terminal work: strong fit, especially at Medium or High effort in Claude Code, Cursor or GitHub Copilot. Its Terminal-Bench 4.0 lead over Opus 5.5 is real.
- App builders and vibe-coding platforms: Base44 and Lovable report fewer iterations and tool calls, which translates directly into lower cost and faster builds.
- Documents, slides, spreadsheets and support: Anthropic's stated sweet spot, backed by Box, Zendesk and Slack results.
- High-volume agents with untrusted input: its prompt-injection robustness is the best of any Sonnet.
Choose Opus 5.5 instead for very large refactors, long-horizon engineering and research where sustained judgment matters. Look elsewhere if your core work is offensive security or binary analysis (apply to the Cyber Verification Program), and consider waiting for Haiku 5.5, or using Haiku 4.5 at $1/$5 today, for bulk classification and extraction where speed and price dominate.
The Bottom Line
Claude Sonnet 5.5 is the most important Claude release of the month for most users, even though Opus 5.5 got the bigger launch. At an unchanged $2/$10, it beats Opus 5.5 on Terminal-Bench 4.0, ties it on GDPval-AA, clearly outscores the identically priced GPT-6 Sol on knowledge work, AutomationBench and chart recognition, and runs 30%+ faster than Sonnet 5 with far fewer tokens per task. Its system card is reassuring on containment and prompt injection, and honest about weaker spots: more hallucinated wrong answers than other recent Claude models, some harmlessness regressions, and a narrower assessment than Opus 5.5 received.
The main cost of the upgrade is friction, not money: frontier-style cyber classifiers that will refuse some legitimate security work, anti-distillation controls, and five breaking API changes. For everyday coding, knowledge work and agents, make Sonnet 5.5 your default, keep effort at Medium or High, and escalate to Opus 5.5 only when a task genuinely needs sustained judgment.
Sources
- Anthropic: Introducing Claude Sonnet 5.5 (28/09/2026)
- Anthropic: System Card, Claude Sonnet 5.5
- Claude Platform Docs: Claude Sonnet 5.5 overview
- The Next Web: Sonnet 5.5 cyber and distillation safeguards
- Decrypt: Claude Sonnet 5.5 release (Artificial Analysis figures)
- SiliconANGLE: Anthropic debuts Claude Sonnet 5.5
- Unite.AI: Claude Sonnet 5.5 at unchanged Sonnet 5 pricing
- The Decoder: Sonnet 5.5 nearly matches Opus 5.5
- The New Stack: Sonnet 5.5 launch
- Simon Willison: Claude Sonnet 5.5
Last updated: 29/09/2026. This review is based on Anthropic's official Claude Sonnet 5.5 announcement, system card and platform documentation. Most benchmark figures are Anthropic-run or partner-run; we will update as independent evaluations land.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.








