AI Tools Review
Claude Sonnet 5.5 vs GPT-6.1 Sol: Benchmarks & Cost

Comparison

Claude Sonnet 5.5 vs GPT-6.1 Sol: Benchmarks & Cost

AI Tools Review Editorial Team29 September 2026
  • Claude Sonnet 5.5
  • GPT-6.1 Sol
  • Anthropic
  • OpenAI

For the second time in a week, Anthropic and OpenAI have shipped rival models within a day of each other at an identical price. Claude Sonnet 5.5 landed on Monday 28 September 2026; GPT-6.1 Sol arrived on stage at OpenAI DevDay on Tuesday 29 September, billed as "near-Astra intelligence" at a fifth of GPT-6 Astra's token price. Both cost $2 in and $10 out per million tokens. This is the head-to-head: what each vendor actually published, what independent evaluators have measured so far, and what the price parity means once caching, effort settings and token counts are factored in.

If you want the background on each model first, read our Claude Sonnet 5.5 review and GPT-6.1 Sol review, plus the OpenAI DevDay 2026 recap.

Method note: we only put two numbers side by side when both come from the same benchmark version and, ideally, the same evaluator. Anthropic's launch table compares Sonnet 5.5 against the older GPT-6 Sol, not GPT-6.1 Sol, and OpenAI's launch charts compare GPT-6.1 Sol against Claude Opus 5.5, not Sonnet 5.5. Where one vendor has not published a score, we say "not published" rather than guessing. Sterling conversions use approx. £0.75 per $1.

Matthew Berman tests Claude Sonnet 5.5 on launch day by building full playable projects, including the Age of Empires-style Crownfall.

Claude Sonnet 5.5 or GPT-6.1 Sol: What Is the Verdict?

Claude Sonnet 5.5 is the stronger and faster model; GPT-6.1 Sol is the cheaper model to actually run. That is the short answer from the first independent data, and it holds across most of the evaluations published so far.

  • Capability: Artificial Analysis ranks Sonnet 5.5 at 56 on its Intelligence Index (v4.3.2), second only to Claude Opus 5.5 (58), with GPT-6.1 Sol at 52, one point behind GPT-6 Astra (53).
  • Where Sonnet 5.5 wins: agentic terminal coding, scientific coding, long-horizon knowledge work (GDPval-AA, AA-Briefcase) and multi-tool business workflows (AutomationBench).
  • Where GPT-6.1 Sol wins: question answering over complex PDFs (GDP.pdf), factual reliability (AA-Omniscience), and raw token efficiency.
  • Speed: Sonnet 5.5 generates roughly twice as many tokens per second at standard speed; OpenAI counters with a paid Ultrafast tier (up to 300 tokens per second) that was "coming in days" for Sol at launch.
  • Cost: identical list prices, but Sol's cached input is half the price, and on Artificial Analysis's max-effort runs Sol cost $0.72 per task versus $7.60 for Sonnet 5.5, because Sonnet thought for about five times as many tokens.

The practical upshot: if you are paying per task at high volume, GPT-6.1 Sol is hard to beat on price-performance. If your bottleneck is the quality of agentic coding or long professional deliverables, and you can control Sonnet 5.5's effort level, Sonnet 5.5 is the better tool. Neither vendor's numbers alone settle it, so treat this as a shortlist for your own testing.

How Do the Specs and Pricing Compare?

Claude Sonnet 5.5 and GPT-6.1 Sol share the same headline API price and roughly the same context window, but differ on cache pricing, long-context charges, speed tiers and knowledge cutoff.

SpecClaude Sonnet 5.5GPT-6.1 Sol
DeveloperAnthropicOpenAI
Release28/09/202629/09/2026 (DevDay)
API model IDclaude-sonnet-5-5gpt-6.1-sol
Input (per 1M tokens)$2 (approx. £1.50)$2 (approx. £1.50)
Output (per 1M tokens)$10 (approx. £7.50)$10 (approx. £7.50)
Cached input / cache reads$0.20$0.10
Cache writes$2.50 (5-minute), $4 (1-hour)$2.50
Batch discount50% off input and output50% off (Batch/Flex)
Long-context surchargeNone across the 1M windowAbove 272K tokens: 2x input and cache, 1.5x output
Faster tiersNo fast mode listed for Sonnet 5.5Fast mode at 2x price; Ultrafast (6x price) announced, "coming" for Sol
Context window1M tokens1.05M tokens
Max output128K (300K on Batch, beta)128K
Knowledge cutoffJune 202630/04/2026
Reasoning controlAdaptive thinking, effort levels; default High on API, Medium in Claude apps and Claude Codelow, medium (default), high, xhigh, max; no none/minimal
InputsText and imagesText and images
Where to use itClaude apps (incl. free tier), Claude Code, Claude API, Amazon Bedrock, Google Cloud, Microsoft FoundryOpenAI API, ChatGPT Work and Codex (Plus, Pro, Business, Enterprise, Edu); not yet in standard ChatGPT chat

Sources: Claude Platform docs, Anthropic, OpenAI developer docs, VentureBeat. Checked 29/09/2026.

Two pricing details matter more than the matching headline. First, GPT-6.1 Sol halves the cached-input price from GPT-6 Sol's $0.20 to $0.10, which is 5% of its standard input rate; Sonnet 5.5 keeps Anthropic's usual 10% ratio at $0.20. For agents that re-read the same large context on every turn, cache reads are often the biggest line item. Second, OpenAI surcharges prompts above 272,000 tokens, whereas Anthropic prices Sonnet 5.5's whole 1M-token window at the same per-token rate. For whole-codebase or large document-pack prompts, that flips the cost advantage. For the wider market, see our AI API pricing comparison and Claude API pricing guide.

How Do Claude Sonnet 5.5 and GPT-6.1 Sol Compare on Benchmarks?

The only benchmark set where both models have been tested by the same evaluator, on the same versions, is Artificial Analysis's Intelligence Index suite. Both vendors' own launch tables compare against other rivals, so they cannot be combined directly. Here are the independent numbers first.

Artificial Analysis evaluationClaude Sonnet 5.5 (max)GPT-6.1 Sol (max)Leader
Intelligence Index v4.3.25652Sonnet 5.5
Terminal-Bench 4.0 (agentic coding)64%56%Sonnet 5.5
SciCode (scientific coding)61%54%Sonnet 5.5
GDPval-AA v2.1 (knowledge work, Elo)18441575Sonnet 5.5
AA-Briefcase v1.1 (long-horizon work, Elo)18111564Sonnet 5.5
AutomationBench-AA (SaaS workflows)71%65%Sonnet 5.5
Humanity's Last Exam55%53%Sonnet 5.5 (narrow)
GDP.pdf (PDF question answering)26%31%GPT-6.1 Sol
AA-Omniscience (knowledge and hallucination, index)3242GPT-6.1 Sol
CritPt (frontier physics)31%32%GPT-6.1 Sol (narrow)
AA-LCR v1.1 (long-context reasoning)83%83%Tie
Output tokens per index task193K38KGPT-6.1 Sol (fewer)
Cost per index task$7.60$0.72GPT-6.1 Sol
Cost to run whole index$8,977$1,082GPT-6.1 Sol

Source: Artificial Analysis model comparison, accessed 29/09/2026. Sonnet 5.5 run with adaptive reasoning at max effort and default fallbacks. Anthropic notes that AA's GDPval-AA and AA-Briefcase runs used a pre-release deployment with a since-fixed structured-output bug, which it expects to understate Sonnet 5.5 slightly, if at all.

Next, the vendor-published numbers. Each row shows what each company actually reported; "not published" means the vendor has not released that score for its model.

Vendor-reported benchmarkClaude Sonnet 5.5 (Anthropic)GPT-6.1 Sol (OpenAI)
Terminal-Bench 4.070.6%Not published
CursorBench 4.055.5%Not published
FrontierCode 1.1 (Main)46.2% (max)Not published
DeepSWE v1.1Not published75.2% (high)
AutomationBench 1.0.6 (Zapier)44.75% (max, Zapier leaderboard)About 36% (max, OpenAI-reported)
OSWorld (computer use)80.1% on OSWorld 2.1, partial credit71.4% on OSWorld 2.0 (max)
GDP.pdfNot published32.0% (high)
Terminal-Bench-Science 0.1Not published57.0% (max)
Humanity's Last Exam (with tools)64.5%Not published
Chartography (no tools)61.6%Not published
LMArena (text / WebDev)Not yet rankedNot yet ranked

Sources: Anthropic; OpenAI's launch figures as reported by Vellum, Handy AI and OrcaRouter; Zapier AutomationBench. OSWorld versions differ, so those two numbers are not directly comparable. Some third-party write-ups cite a DeepSWE score for Sonnet 5.5; we could not trace it to an Anthropic source, so we leave it out.

Three honest caveats. First, Anthropic's own Terminal-Bench 4.0 figure for Sonnet 5.5 (70.6%) is higher than Artificial Analysis's independent run (64%), a reminder that harness and effort settings move scores by several points. Second, Sonnet 5.5's lead is bought partly with tokens: at max effort it produced about five times as much output per task as GPT-6.1 Sol. Third, both vendors have a history of shipping quick fixes after launch (Anthropic notes OpenAI recently fixed an image-understanding bug in GPT-6 Sol), so expect these numbers to move over the next few weeks. For the wider leaderboard, see our frontier AI benchmarks round-up.

Which Is Better for Coding, Sonnet 5.5 or GPT-6.1 Sol?

Claude Sonnet 5.5 has the stronger independent coding results, whilst GPT-6.1 Sol has the stronger vendor claim on repository-scale software engineering. The two companies simply chose different headline tests, so the independent Terminal-Bench 4.0 and SciCode runs are the fairest direct comparison, and Sonnet 5.5 leads both by seven to eight points.

Anthropic's pitch is that Sonnet 5.5 is a huge jump over Claude Sonnet 5: 70.6% on Terminal-Bench 4.0 against Sonnet 5's 10.3%, and 55.5% on CursorBench 4.0, within about two points of Claude Opus 5.5. It positions the model for "well-scoped everyday tasks" and bug fixing rather than the open-ended architectural work it still reserves for Opus 5.5. Customer quotes back this up: Lovable reported a third fewer tool calls and roughly half the shell runs per task, and Base44 said that across 118 real app builds, Sonnet 5.5 reached Opus 5-level quality in 3.6 iterations versus 7.7.

OpenAI's headline for GPT-6.1 Sol is 75.2% on DeepSWE v1.1 at high effort, which it says matches GPT-6 Astra at roughly a fifth of the cost and beats GPT-6 Sol's best score by 6.4 points. It is a meaningful claim, but Anthropic has not published a DeepSWE number for Sonnet 5.5, so there is no like-for-like vendor comparison.

One Sonnet-specific quirk worth knowing: Anthropic notes Sonnet 5.5 scores lower at Max effort than at Xhigh on FrontierCode, because at Max it more often ran Claude Code's multi-subagent code-review skill, leading to timeouts or out-of-scope edits. Simon Willison hit a related issue, with Max effort burning 128,000 tokens ($1.28) without producing output on a simple SVG test, whilst Xhigh finished in 41 seconds for under 6 cents. The lesson for developers: do not assume Max is best for Sonnet 5.5; Xhigh or High is often the sweet spot.

Our read: for agentic terminal work and IDE-style multi-file edits, Sonnet 5.5 has the better evidence today. For large, well-defined repository maintenance where cost per task matters, GPT-6.1 Sol at high effort is a very credible alternative. Test both on a slice of your own backlog. Both models also run in third-party tools such as Cursor and GitHub Copilot, where model availability can vary by plan.

Side-by-side demo: Claude Sonnet 5 still writing code at 27,581 tokens whilst Claude Sonnet 5.5 has finished its clock-of-clocks animation using 14,386 output tokens and has been running for 20.9 seconds
Anthropic's "clock of clocks" demo: Sonnet 5 (left) is still writing at 27,581 tokens whilst Sonnet 5.5 (right) has already finished with 14,386 output tokens and been running for 20.9 seconds. Source: Anthropic.

Which Is Better for Agents and Tool Use?

Claude Sonnet 5.5 currently leads on multi-step, multi-tool business automation, the category closest to what most companies mean by "agents". On Zapier's AutomationBench 1.0.6, which tests end-to-end workflows across 47 real tools in sales, marketing, operations, support, finance and HR, Sonnet 5.5 at max effort sits at the top of the public leaderboard with 44.75% at $1.14 per task, ahead of Opus 5.5 (42.47%) and GPT-6 Astra (41.4%). At Xhigh it scores 36.83% for $0.48 per task.

OpenAI reports GPT-6.1 Sol at about 36% at max effort, 2.2 points above Claude Opus 5.5 at medium effort and 4.8 points up on GPT-6 Sol, at roughly a third of Opus 5.5's cost. GPT-6.1 Sol was not yet on Zapier's public board when we checked, so this is a vendor figure. Artificial Analysis's own variant, AutomationBench-AA, independently puts Sonnet 5.5 ahead, 71% to 65%.

On computer use, both are strong but not comparable: Anthropic reports 80.1% for Sonnet 5.5 on OSWorld 2.1 with partial credit, while OpenAI reports 71.4% for GPT-6.1 Sol on OSWorld 2.0 at max effort, within 2.1 points of GPT-6 Astra at about one-seventh of its cost per task. Different versions and scoring mean neither number can be declared the winner.

On agent reliability, the signals point in different directions. OpenAI says GPT-6.1 Sol fails to flag a broken search tool 2.1% of the time (versus 4.9% for GPT-6 Sol) and attempts to bypass explicit restrictions far less often than its predecessor (23.5% versus 64.4%), though still more often than GPT-6 Astra (17.4%). Anthropic says Sonnet 5.5 has the lowest rate of attempted sandbox escapes and unprompted boundary probing of any Claude model it tested, and is its most prompt-injection-robust Sonnet yet. Zendesk reported Sonnet 5.5 made fewer wrong decisions and resolved support tickets 20% faster than its production models.

For fully managed agents, OpenAI also launched persistent "Dots" agents at DevDay; see OpenAI Dots agents explained. For building your own, the older OpenAI Agents API and Anthropic's Claude Code hooks remain the main building blocks.

Liam Ottley and team react live to the DevDay 2026 keynote where OpenAI announced GPT-6.1 Sol, Dots agents and the Ultrafast speed tier.

Which Is Better for Writing and Knowledge Work?

Claude Sonnet 5.5 is clearly ahead on producing professional deliverables; GPT-6.1 Sol is ahead on answering precise questions from documents and on factual reliability. That split is the most useful finding in the independent data.

On GDPval-AA v2.1, which grades real work products across 44 occupations, Sonnet 5.5 scores 1844 Elo against GPT-6.1 Sol's 1575, a gap of over 250 Elo, and it is just two points behind Opus 5.5. AA-Briefcase, a newer long-horizon knowledge-work test, tells the same story (1811 vs 1564). Anthropic says Sonnet 5.5 shares Opus 5.5's clearer writing style and is strong at "polished documents, slides, and spreadsheets". Box reported it rechecks source documents and catches errors Sonnet 5 missed, whilst running 2.4x faster; Balyasny Asset Management said it scored ahead on 2,441 finance tasks using 121K tokens per answer versus Sonnet 5's 497K.

GPT-6.1 Sol hits back on precision. On GDP.pdf, which tests answers to professional questions over complex PDFs, OpenAI reports 32.0%, level with GPT-6 Astra and above Opus 5.5 with fallbacks; Artificial Analysis independently has Sol ahead of Sonnet 5.5, 31% to 26%. On AA-Omniscience, which rewards correct answers and penalises confident wrong ones, Sol scores 42 versus Sonnet's 32. OpenAI also says Sol cut factual errors from 11.4% to 7.7% at low effort versus GPT-6 Sol.

Practical takeaway: for drafting reports, decks, analyses and long written work, start with Sonnet 5.5. For extraction, document Q&A and research where a confident wrong answer is costly, GPT-6.1 Sol has the better independent evidence. Our earlier piece on Claude's writing strengths gives the longer history.

Which Is Faster, Sonnet 5.5 or GPT-6.1 Sol?

At standard speed, Claude Sonnet 5.5 is roughly twice as fast. Artificial Analysis measured 138 output tokens per second for Sonnet 5.5 (max effort, with fallbacks) against 67 for GPT-6.1 Sol (max effort). Anthropic says Sonnet 5.5 generates output over 30% faster than Sonnet 5, making it its fastest Sonnet to date, and Atlassian reported Rovo agents running up to 30% faster.

Artificial Analysis output speed bar chart in tokens per second, showing Claude Sonnet 5.5 (max with fallback) at 138, GPT-6.1 Sol (max) at 67, GPT-6 Sol (max) at 76 and GPT-6 Astra (max) at 57
Output speed in tokens per second: Claude Sonnet 5.5 at 138 versus GPT-6.1 Sol at 67 at max effort on standard tiers. Source: Artificial Analysis (via VentureBeat).

Tokens per second is only half of speed. Because Sonnet 5.5 at max effort generated about five times as many tokens per task, its end-to-end time on Artificial Analysis's 500-token response test was slower (about 378 seconds versus 275 for Sol, both dominated by thinking time at max effort). At the lower effort settings most people actually use, both respond far faster; Anthropic sets Medium as the default in Claude Code and the Claude apps for that reason.

OpenAI's speed answer is Ultrafast, a paid tier announced at DevDay that it says reaches up to 300 generated tokens per second: up to 8x faster in Codex and up to 6x faster through the API, with API usage billed at 6x standard pricing. At launch, Ultrafast was available for GPT-6 Astra (for Pro 500 and Enterprise) and described as coming within days for GPT-6.1 Sol. OpenAI has not published a separate Sol Ultrafast price; if the 6x rule applies, that would be about $12 in / $60 out per million tokens (approx. £9 / £45), which is our calculation, not an OpenAI figure. There is also a Fast mode at 2x standard price. Anthropic lists no fast mode for Sonnet 5.5; its Fast mode is offered on Opus 5.5.

What Does Each Model Cost per Task?

Identical list prices do not mean identical bills. The worked examples below use each vendor's published rates and assume the same token counts for both models, so they isolate the pricing structure. In reality, token counts differ by model and effort level, which is covered in the fourth example. Sterling figures use approx. £0.75 per $1 and exclude VAT.

ScenarioToken assumptionsClaude Sonnet 5.5GPT-6.1 Sol
1. Agentic coding session3M cached-read tokens, 300K fresh input, 60K output$0.60 + $0.60 + $0.60 = $1.80 (approx. £1.35)$0.30 + $0.60 + $0.60 = $1.50 (approx. £1.13)
2. 10,000 support ticketsPer ticket: 2,000 cached system-prompt tokens, 1,000 fresh input, 400 output$4 + $20 + $40 = $64 (approx. £48)$2 + $20 + $40 = $62 (approx. £46.50)
3. One 400K-token contract pack400K fresh input, 5K output, no caching$0.80 + $0.05 = $0.85 (approx. £0.64)$1.60 + $0.075 = $1.68 (approx. £1.26), with the long-context surcharge
4. Artificial Analysis index task (measured)Max effort; Sonnet ~193K output tokens, Sol ~38K$7.60 (approx. £5.70)$0.72 (approx. £0.54)

Examples 1 to 3 are our illustrative calculations from list prices (cache-write charges, which are $2.50 per million on both, are excluded for simplicity). Example 4 is measured by Artificial Analysis.

The pattern is clear. When the workload is cache-heavy (examples 1 and 2), GPT-6.1 Sol's $0.10 cache reads give it a modest 3% to 17% edge at equal token counts. When a single prompt is very long (example 3), Sonnet 5.5 is about half the price because OpenAI surcharges anything over 272K tokens. And when you let both models think as hard as they like (example 4), token volume swamps the rate card: Sonnet 5.5 at max effort cost more than ten times as much per task.

That last point cuts both ways. Anthropic says Sonnet 5.5 typically costs up to 30% less per task than Sonnet 5, and that at Low or Medium effort it beats Sonnet 5's best scores on several benchmarks for about a tenth of the cost. On AutomationBench, Sonnet 5.5 at Xhigh cost $0.48 per task for 36.83%, broadly matching OpenAI's reported ~36% for Sol at max. The biggest lever on your bill is the effort setting, not the vendor. Batch processing halves costs on both platforms for anything that is not latency-sensitive.

How Do Their Safety Profiles and Safeguards Differ?

Both companies treat these mid-tier models as capable enough to need frontier-grade cyber safeguards, but OpenAI classifies GPT-6.1 Sol at a higher formal risk tier.

Claude Sonnet 5.5. Anthropic's system card says Sonnet 5.5 is broadly less capable than Opus 5.5 and crosses no new Responsible Scaling Policy thresholds. Because its cyber capability is now comparable to Opus 5's, it is the first Sonnet model with cyber safeguards and fallbacks: routine bug-finding and fixing is unaffected, but higher-risk cyber tasks visibly fall back to Sonnet 5, with a Cyber Verification Program for approved defenders. Biology safeguards are unchanged from Sonnet 5. It is also the first Sonnet launched with anti-distillation classifiers and "preserved thinking", which ties Claude's reasoning to the account that created it. On alignment, Anthropic reports it matches or beats Sonnet 5 on most measures, with the lowest sandbox-escape and boundary-probing rates of any model tested, but flags that its thinking is "more illegible" than many earlier models and that multi-turn tests showed regressions in areas such as tracking and surveillance.

GPT-6.1 Sol. OpenAI's system card addendum treats GPT-6.1 Sol as Critical in cybersecurity and High in biological and chemical capability under its Preparedness Framework, below High for AI self-improvement, and ships it with the same safeguards stack as GPT-6 Astra. Reported results include a 1.50% coding-deception rate (versus 0.51% for GPT-6 Astra), 35.1% on ExploitGym and 21.5% arbitrary code execution on ExploitBench. Context matters here: OpenAI also shelved a planned GPT-6.1 Astra release; see why GPT-6.1 Astra was cancelled and our explainer on Astra's Critical cyber rating.

For most business users, the practical difference is how refusals feel. Sonnet 5.5 hands higher-risk cyber prompts to an older model rather than refusing outright; GPT-6.1 Sol inherits Astra's trusted-access approach. Security teams doing legitimate offensive work should apply to the relevant verification programme on either platform rather than fight the defaults.

Matthew Berman walks through what his team built with Claude Sonnet 5.5, headlined by the strategy game Crownfall.

Claude Code vs Codex: Which Ecosystem Fits?

For many teams, the harness matters as much as the model: Sonnet 5.5 is the everyday workhorse inside Claude Code, and GPT-6.1 Sol is the new default-class model inside OpenAI's Codex.

Claude Code and the Claude apps. Sonnet 5.5 runs at Medium effort by default in Claude Code and the Claude apps, and Simon Willison notes it now powers the free tier on claude.ai. On the API, it is available on Anthropic's own platform plus Amazon Bedrock, Google Cloud and Microsoft Foundry, which matters for UK and EU organisations with existing cloud commitments. Developers moving from Sonnet 5 should read the migration guide: forced tool use now returns an error, non-default temperature settings return a 400, and turning off up-front thinking now uses a between_tools setting. See our Claude Code guide and Claude Code pricing for plan costs.

Codex and ChatGPT Work. GPT-6.1 Sol is available to ChatGPT Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, but not yet in standard ChatGPT chat. DevDay also brought Codex cloud tasks that keep running with your laptop closed, two-way voice and an /agents view in the CLI, a Code Review feature and a Security Cloud scanner, plus Ultrafast generation in Codex. On the API, tool calling requires the Responses API. For the pre-DevDay state of play, see OpenAI Codex August 2026 updates.

If your team already lives in one ecosystem, the switching cost is usually higher than the model gap. If you are starting fresh, run a one-week trial of each agent on the same backlog.

Which Should You Choose for Your Use Case?

Use casePickWhy
Agentic coding in the terminal or IDEClaude Sonnet 5.5Leads Terminal-Bench 4.0 and SciCode independently; strong CursorBench
High-volume repo maintenance on a budgetGPT-6.1 Sol75.2% DeepSWE claim, far fewer tokens per task
Business workflow automation across SaaS toolsClaude Sonnet 5.5Tops Zapier AutomationBench; leads AutomationBench-AA
Reports, decks, spreadsheets, long writingClaude Sonnet 5.5Over 250 Elo ahead on GDPval-AA; clearer writing
Document Q&A and extraction from PDFsGPT-6.1 SolLeads GDP.pdf and AA-Omniscience
Cache-heavy chat or support botsGPT-6.1 Sol$0.10 cache reads, half Sonnet's rate
Single prompts over 272K tokensClaude Sonnet 5.5No long-context surcharge across 1M tokens
Lowest latency at standard priceClaude Sonnet 5.5About 2x the tokens per second of Sol
Absolute fastest generation, cost no objectGPT-6.1 Sol (Ultrafast, when live)Up to 300 tokens per second at 6x price
Multi-cloud deployment (AWS, Google, Microsoft)Claude Sonnet 5.5Available on Bedrock, Google Cloud and Microsoft Foundry at launch

If you need more capability than either offers, step up to Claude Opus 5.5 or GPT-6 Astra. If you are comparing last week's models, our GPT-6 Sol vs Claude Opus 5.5 and GPT-6 Sol vs Astra vs Luna guides still apply, and the original GPT-6 Sol and Luna review covers the model GPT-6.1 Sol replaces. Want to try Anthropic's lineup directly? See the Claude Opus 5.5 and Claude Sonnet 5 tool pages, or browse creator coverage on our videos page.

Sources

The Bottom Line

Claude Sonnet 5.5 and GPT-6.1 Sol are the most evenly matched pair of mid-tier models yet: same price, same class of context window, same 128K output ceiling, released a day apart. On the first independent evidence, Sonnet 5.5 is the more capable and faster model, with clear leads on agentic coding, knowledge work and workflow automation. GPT-6.1 Sol is the more economical model, with half-price cache reads, better factual reliability and dramatically lower token use at high effort.

Our recommendation: use Sonnet 5.5 where output quality and turnaround time drive value, and cap its effort level to keep costs in line. Use GPT-6.1 Sol for high-volume, cache-heavy or extraction-style work where every penny per task counts. And because neither vendor benchmarked against the other's new model, run your own evaluation before you commit a production workload.

Last updated: 29/09/2026. Based on Anthropic's and OpenAI's official announcements, documentation and system cards, Artificial Analysis's independent evaluations and Zapier's AutomationBench leaderboard. GPT-6.1 Sol launched today, so independent results will fill in over the coming weeks; LMArena rankings were not yet available for either model.

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

Is Claude Sonnet 5.5 better than GPT-6.1 Sol?
On independent testing, mostly yes. Artificial Analysis scores Claude Sonnet 5.5 at 56 on its Intelligence Index versus 52 for GPT-6.1 Sol (max effort), and Sonnet 5.5 leads on Terminal-Bench 4.0 (64% vs 56%), GDPval-AA v2.1 (1844 vs 1575 Elo), AA-Briefcase, AutomationBench-AA and SciCode. GPT-6.1 Sol leads on GDP.pdf (31% vs 26%), AA-Omniscience factual reliability (42 vs 32) and narrowly on CritPt, and it used about a fifth of the output tokens to run the same suite, so it is far cheaper per task at max effort.
How much do Claude Sonnet 5.5 and GPT-6.1 Sol cost?
Both list at $2 per million input tokens and $10 per million output tokens (approx. £1.50 and £7.50). The differences are in the details: GPT-6.1 Sol charges $0.10 per million for cached input versus $0.20 for Sonnet 5.5, but OpenAI doubles input and cache rates and raises output rates 1.5x for prompts above 272,000 tokens, whereas Anthropic has no long-context surcharge on Sonnet 5.5's 1M-token window. Both offer 50% batch discounts.
Which is faster, Claude Sonnet 5.5 or GPT-6.1 Sol?
At standard speed, Claude Sonnet 5.5. Artificial Analysis measured 138 output tokens per second for Sonnet 5.5 at max effort versus 67 for GPT-6.1 Sol at max effort, and Anthropic says Sonnet 5.5 is over 30% faster than Sonnet 5. OpenAI's answer is the paid Ultrafast tier, which it says reaches up to 300 tokens per second (up to 8x faster in Codex, 6x in the API) at 6x standard API pricing, but Ultrafast for GPT-6.1 Sol was still listed as coming in the following days at launch.
Which is better for coding: Claude Code with Sonnet 5.5 or Codex with GPT-6.1 Sol?
There is no shared vendor coding benchmark: Anthropic publishes Terminal-Bench 4.0 (70.6%), CursorBench 4.0 (55.5%) and FrontierCode (46.2%) for Sonnet 5.5, whilst OpenAI publishes DeepSWE v1.1 (75.2%) for GPT-6.1 Sol. On Artificial Analysis's independent runs, Sonnet 5.5 leads on Terminal-Bench 4.0 (64% vs 56%) and SciCode (61% vs 54%). In practice, pick the agent harness your team prefers and test both models on your own repository; GPT-6.1 Sol is the cheaper option at high effort.
What are the safety differences between Claude Sonnet 5.5 and GPT-6.1 Sol?
Anthropic says Sonnet 5.5 crosses no new Responsible Scaling Policy thresholds but is the first Sonnet with cyber safeguards, falling back to Sonnet 5 on higher-risk cyber tasks. OpenAI treats GPT-6.1 Sol as Critical in cybersecurity and High in biological and chemical capability under its Preparedness Framework and ships it with the same safeguards stack as GPT-6 Astra. OpenAI's system card reports a 1.50% coding-deception rate for GPT-6.1 Sol versus 0.51% for GPT-6 Astra.

Key takeaways

Same price, different bills

Both list at $2/$10 per million tokens, but GPT-6.1 Sol has half-price cache reads and a long-context surcharge above 272K tokens; Sonnet 5.5 has neither.

Sonnet 5.5 wins most independent evals

Artificial Analysis: 56 vs 52 on the Intelligence Index, with Sonnet ahead on Terminal-Bench 4.0, GDPval-AA, AutomationBench-AA and SciCode.

GPT-6.1 Sol wins on efficiency

It used 38K output tokens per index task versus 193K for Sonnet 5.5 at max effort, so its cost per task was $0.72 versus $7.60.

AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.