For the second time in a week, Anthropic and OpenAI have shipped rival models within a day of each other at an identical price. Claude Sonnet 5.5 landed on Monday 28 September 2026; GPT-6.1 Sol arrived on stage at OpenAI DevDay on Tuesday 29 September, billed as "near-Astra intelligence" at a fifth of GPT-6 Astra's token price. Both cost $2 in and $10 out per million tokens. This is the head-to-head: what each vendor actually published, what independent evaluators have measured so far, and what the price parity means once caching, effort settings and token counts are factored in.
If you want the background on each model first, read our Claude Sonnet 5.5 review and GPT-6.1 Sol review, plus the OpenAI DevDay 2026 recap.
Method note: we only put two numbers side by side when both come from the same benchmark version and, ideally, the same evaluator. Anthropic's launch table compares Sonnet 5.5 against the older GPT-6 Sol, not GPT-6.1 Sol, and OpenAI's launch charts compare GPT-6.1 Sol against Claude Opus 5.5, not Sonnet 5.5. Where one vendor has not published a score, we say "not published" rather than guessing. Sterling conversions use approx. £0.75 per $1.
Matthew Berman tests Claude Sonnet 5.5 on launch day by building full playable projects, including the Age of Empires-style Crownfall.
Claude Sonnet 5.5 or GPT-6.1 Sol: What Is the Verdict?
Claude Sonnet 5.5 is the stronger and faster model; GPT-6.1 Sol is the cheaper model to actually run. That is the short answer from the first independent data, and it holds across most of the evaluations published so far.
- Capability: Artificial Analysis ranks Sonnet 5.5 at 56 on its Intelligence Index (v4.3.2), second only to Claude Opus 5.5 (58), with GPT-6.1 Sol at 52, one point behind GPT-6 Astra (53).
- Where Sonnet 5.5 wins: agentic terminal coding, scientific coding, long-horizon knowledge work (GDPval-AA, AA-Briefcase) and multi-tool business workflows (AutomationBench).
- Where GPT-6.1 Sol wins: question answering over complex PDFs (GDP.pdf), factual reliability (AA-Omniscience), and raw token efficiency.
- Speed: Sonnet 5.5 generates roughly twice as many tokens per second at standard speed; OpenAI counters with a paid Ultrafast tier (up to 300 tokens per second) that was "coming in days" for Sol at launch.
- Cost: identical list prices, but Sol's cached input is half the price, and on Artificial Analysis's max-effort runs Sol cost $0.72 per task versus $7.60 for Sonnet 5.5, because Sonnet thought for about five times as many tokens.
The practical upshot: if you are paying per task at high volume, GPT-6.1 Sol is hard to beat on price-performance. If your bottleneck is the quality of agentic coding or long professional deliverables, and you can control Sonnet 5.5's effort level, Sonnet 5.5 is the better tool. Neither vendor's numbers alone settle it, so treat this as a shortlist for your own testing.
How Do the Specs and Pricing Compare?
Claude Sonnet 5.5 and GPT-6.1 Sol share the same headline API price and roughly the same context window, but differ on cache pricing, long-context charges, speed tiers and knowledge cutoff.
| Spec | Claude Sonnet 5.5 | GPT-6.1 Sol |
|---|---|---|
| Developer | Anthropic | OpenAI |
| Release | 28/09/2026 | 29/09/2026 (DevDay) |
| API model ID | claude-sonnet-5-5 | gpt-6.1-sol |
| Input (per 1M tokens) | $2 (approx. £1.50) | $2 (approx. £1.50) |
| Output (per 1M tokens) | $10 (approx. £7.50) | $10 (approx. £7.50) |
| Cached input / cache reads | $0.20 | $0.10 |
| Cache writes | $2.50 (5-minute), $4 (1-hour) | $2.50 |
| Batch discount | 50% off input and output | 50% off (Batch/Flex) |
| Long-context surcharge | None across the 1M window | Above 272K tokens: 2x input and cache, 1.5x output |
| Faster tiers | No fast mode listed for Sonnet 5.5 | Fast mode at 2x price; Ultrafast (6x price) announced, "coming" for Sol |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K (300K on Batch, beta) | 128K |
| Knowledge cutoff | June 2026 | 30/04/2026 |
| Reasoning control | Adaptive thinking, effort levels; default High on API, Medium in Claude apps and Claude Code | low, medium (default), high, xhigh, max; no none/minimal |
| Inputs | Text and images | Text and images |
| Where to use it | Claude apps (incl. free tier), Claude Code, Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry | OpenAI API, ChatGPT Work and Codex (Plus, Pro, Business, Enterprise, Edu); not yet in standard ChatGPT chat |
Sources: Claude Platform docs, Anthropic, OpenAI developer docs, VentureBeat. Checked 29/09/2026.
Two pricing details matter more than the matching headline. First, GPT-6.1 Sol halves the cached-input price from GPT-6 Sol's $0.20 to $0.10, which is 5% of its standard input rate; Sonnet 5.5 keeps Anthropic's usual 10% ratio at $0.20. For agents that re-read the same large context on every turn, cache reads are often the biggest line item. Second, OpenAI surcharges prompts above 272,000 tokens, whereas Anthropic prices Sonnet 5.5's whole 1M-token window at the same per-token rate. For whole-codebase or large document-pack prompts, that flips the cost advantage. For the wider market, see our AI API pricing comparison and Claude API pricing guide.
How Do Claude Sonnet 5.5 and GPT-6.1 Sol Compare on Benchmarks?
The only benchmark set where both models have been tested by the same evaluator, on the same versions, is Artificial Analysis's Intelligence Index suite. Both vendors' own launch tables compare against other rivals, so they cannot be combined directly. Here are the independent numbers first.
| Artificial Analysis evaluation | Claude Sonnet 5.5 (max) | GPT-6.1 Sol (max) | Leader |
|---|---|---|---|
| Intelligence Index v4.3.2 | 56 | 52 | Sonnet 5.5 |
| Terminal-Bench 4.0 (agentic coding) | 64% | 56% | Sonnet 5.5 |
| SciCode (scientific coding) | 61% | 54% | Sonnet 5.5 |
| GDPval-AA v2.1 (knowledge work, Elo) | 1844 | 1575 | Sonnet 5.5 |
| AA-Briefcase v1.1 (long-horizon work, Elo) | 1811 | 1564 | Sonnet 5.5 |
| AutomationBench-AA (SaaS workflows) | 71% | 65% | Sonnet 5.5 |
| Humanity's Last Exam | 55% | 53% | Sonnet 5.5 (narrow) |
| GDP.pdf (PDF question answering) | 26% | 31% | GPT-6.1 Sol |
| AA-Omniscience (knowledge and hallucination, index) | 32 | 42 | GPT-6.1 Sol |
| CritPt (frontier physics) | 31% | 32% | GPT-6.1 Sol (narrow) |
| AA-LCR v1.1 (long-context reasoning) | 83% | 83% | Tie |
| Output tokens per index task | 193K | 38K | GPT-6.1 Sol (fewer) |
| Cost per index task | $7.60 | $0.72 | GPT-6.1 Sol |
| Cost to run whole index | $8,977 | $1,082 | GPT-6.1 Sol |
Source: Artificial Analysis model comparison, accessed 29/09/2026. Sonnet 5.5 run with adaptive reasoning at max effort and default fallbacks. Anthropic notes that AA's GDPval-AA and AA-Briefcase runs used a pre-release deployment with a since-fixed structured-output bug, which it expects to understate Sonnet 5.5 slightly, if at all.
Next, the vendor-published numbers. Each row shows what each company actually reported; "not published" means the vendor has not released that score for its model.
| Vendor-reported benchmark | Claude Sonnet 5.5 (Anthropic) | GPT-6.1 Sol (OpenAI) |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% | Not published |
| CursorBench 4.0 | 55.5% | Not published |
| FrontierCode 1.1 (Main) | 46.2% (max) | Not published |
| DeepSWE v1.1 | Not published | 75.2% (high) |
| AutomationBench 1.0.6 (Zapier) | 44.75% (max, Zapier leaderboard) | About 36% (max, OpenAI-reported) |
| OSWorld (computer use) | 80.1% on OSWorld 2.1, partial credit | 71.4% on OSWorld 2.0 (max) |
| GDP.pdf | Not published | 32.0% (high) |
| Terminal-Bench-Science 0.1 | Not published | 57.0% (max) |
| Humanity's Last Exam (with tools) | 64.5% | Not published |
| Chartography (no tools) | 61.6% | Not published |
| LMArena (text / WebDev) | Not yet ranked | Not yet ranked |
Sources: Anthropic; OpenAI's launch figures as reported by Vellum, Handy AI and OrcaRouter; Zapier AutomationBench. OSWorld versions differ, so those two numbers are not directly comparable. Some third-party write-ups cite a DeepSWE score for Sonnet 5.5; we could not trace it to an Anthropic source, so we leave it out.
Three honest caveats. First, Anthropic's own Terminal-Bench 4.0 figure for Sonnet 5.5 (70.6%) is higher than Artificial Analysis's independent run (64%), a reminder that harness and effort settings move scores by several points. Second, Sonnet 5.5's lead is bought partly with tokens: at max effort it produced about five times as much output per task as GPT-6.1 Sol. Third, both vendors have a history of shipping quick fixes after launch (Anthropic notes OpenAI recently fixed an image-understanding bug in GPT-6 Sol), so expect these numbers to move over the next few weeks. For the wider leaderboard, see our frontier AI benchmarks round-up.
Which Is Better for Coding, Sonnet 5.5 or GPT-6.1 Sol?
Claude Sonnet 5.5 has the stronger independent coding results, whilst GPT-6.1 Sol has the stronger vendor claim on repository-scale software engineering. The two companies simply chose different headline tests, so the independent Terminal-Bench 4.0 and SciCode runs are the fairest direct comparison, and Sonnet 5.5 leads both by seven to eight points.
Anthropic's pitch is that Sonnet 5.5 is a huge jump over Claude Sonnet 5: 70.6% on Terminal-Bench 4.0 against Sonnet 5's 10.3%, and 55.5% on CursorBench 4.0, within about two points of Claude Opus 5.5. It positions the model for "well-scoped everyday tasks" and bug fixing rather than the open-ended architectural work it still reserves for Opus 5.5. Customer quotes back this up: Lovable reported a third fewer tool calls and roughly half the shell runs per task, and Base44 said that across 118 real app builds, Sonnet 5.5 reached Opus 5-level quality in 3.6 iterations versus 7.7.
OpenAI's headline for GPT-6.1 Sol is 75.2% on DeepSWE v1.1 at high effort, which it says matches GPT-6 Astra at roughly a fifth of the cost and beats GPT-6 Sol's best score by 6.4 points. It is a meaningful claim, but Anthropic has not published a DeepSWE number for Sonnet 5.5, so there is no like-for-like vendor comparison.
One Sonnet-specific quirk worth knowing: Anthropic notes Sonnet 5.5 scores lower at Max effort than at Xhigh on FrontierCode, because at Max it more often ran Claude Code's multi-subagent code-review skill, leading to timeouts or out-of-scope edits. Simon Willison hit a related issue, with Max effort burning 128,000 tokens ($1.28) without producing output on a simple SVG test, whilst Xhigh finished in 41 seconds for under 6 cents. The lesson for developers: do not assume Max is best for Sonnet 5.5; Xhigh or High is often the sweet spot.
Our read: for agentic terminal work and IDE-style multi-file edits, Sonnet 5.5 has the better evidence today. For large, well-defined repository maintenance where cost per task matters, GPT-6.1 Sol at high effort is a very credible alternative. Test both on a slice of your own backlog. Both models also run in third-party tools such as Cursor and GitHub Copilot, where model availability can vary by plan.

Which Is Better for Agents and Tool Use?
Claude Sonnet 5.5 currently leads on multi-step, multi-tool business automation, the category closest to what most companies mean by "agents". On Zapier's AutomationBench 1.0.6, which tests end-to-end workflows across 47 real tools in sales, marketing, operations, support, finance and HR, Sonnet 5.5 at max effort sits at the top of the public leaderboard with 44.75% at $1.14 per task, ahead of Opus 5.5 (42.47%) and GPT-6 Astra (41.4%). At Xhigh it scores 36.83% for $0.48 per task.
OpenAI reports GPT-6.1 Sol at about 36% at max effort, 2.2 points above Claude Opus 5.5 at medium effort and 4.8 points up on GPT-6 Sol, at roughly a third of Opus 5.5's cost. GPT-6.1 Sol was not yet on Zapier's public board when we checked, so this is a vendor figure. Artificial Analysis's own variant, AutomationBench-AA, independently puts Sonnet 5.5 ahead, 71% to 65%.
On computer use, both are strong but not comparable: Anthropic reports 80.1% for Sonnet 5.5 on OSWorld 2.1 with partial credit, while OpenAI reports 71.4% for GPT-6.1 Sol on OSWorld 2.0 at max effort, within 2.1 points of GPT-6 Astra at about one-seventh of its cost per task. Different versions and scoring mean neither number can be declared the winner.
On agent reliability, the signals point in different directions. OpenAI says GPT-6.1 Sol fails to flag a broken search tool 2.1% of the time (versus 4.9% for GPT-6 Sol) and attempts to bypass explicit restrictions far less often than its predecessor (23.5% versus 64.4%), though still more often than GPT-6 Astra (17.4%). Anthropic says Sonnet 5.5 has the lowest rate of attempted sandbox escapes and unprompted boundary probing of any Claude model it tested, and is its most prompt-injection-robust Sonnet yet. Zendesk reported Sonnet 5.5 made fewer wrong decisions and resolved support tickets 20% faster than its production models.
For fully managed agents, OpenAI also launched persistent "Dots" agents at DevDay; see OpenAI Dots agents explained. For building your own, the older OpenAI Agents API and Anthropic's Claude Code hooks remain the main building blocks.
Liam Ottley and team react live to the DevDay 2026 keynote where OpenAI announced GPT-6.1 Sol, Dots agents and the Ultrafast speed tier.
Which Is Better for Writing and Knowledge Work?
Claude Sonnet 5.5 is clearly ahead on producing professional deliverables; GPT-6.1 Sol is ahead on answering precise questions from documents and on factual reliability. That split is the most useful finding in the independent data.
On GDPval-AA v2.1, which grades real work products across 44 occupations, Sonnet 5.5 scores 1844 Elo against GPT-6.1 Sol's 1575, a gap of over 250 Elo, and it is just two points behind Opus 5.5. AA-Briefcase, a newer long-horizon knowledge-work test, tells the same story (1811 vs 1564). Anthropic says Sonnet 5.5 shares Opus 5.5's clearer writing style and is strong at "polished documents, slides, and spreadsheets". Box reported it rechecks source documents and catches errors Sonnet 5 missed, whilst running 2.4x faster; Balyasny Asset Management said it scored ahead on 2,441 finance tasks using 121K tokens per answer versus Sonnet 5's 497K.
GPT-6.1 Sol hits back on precision. On GDP.pdf, which tests answers to professional questions over complex PDFs, OpenAI reports 32.0%, level with GPT-6 Astra and above Opus 5.5 with fallbacks; Artificial Analysis independently has Sol ahead of Sonnet 5.5, 31% to 26%. On AA-Omniscience, which rewards correct answers and penalises confident wrong ones, Sol scores 42 versus Sonnet's 32. OpenAI also says Sol cut factual errors from 11.4% to 7.7% at low effort versus GPT-6 Sol.
Practical takeaway: for drafting reports, decks, analyses and long written work, start with Sonnet 5.5. For extraction, document Q&A and research where a confident wrong answer is costly, GPT-6.1 Sol has the better independent evidence. Our earlier piece on Claude's writing strengths gives the longer history.
Which Is Faster, Sonnet 5.5 or GPT-6.1 Sol?
At standard speed, Claude Sonnet 5.5 is roughly twice as fast. Artificial Analysis measured 138 output tokens per second for Sonnet 5.5 (max effort, with fallbacks) against 67 for GPT-6.1 Sol (max effort). Anthropic says Sonnet 5.5 generates output over 30% faster than Sonnet 5, making it its fastest Sonnet to date, and Atlassian reported Rovo agents running up to 30% faster.

Tokens per second is only half of speed. Because Sonnet 5.5 at max effort generated about five times as many tokens per task, its end-to-end time on Artificial Analysis's 500-token response test was slower (about 378 seconds versus 275 for Sol, both dominated by thinking time at max effort). At the lower effort settings most people actually use, both respond far faster; Anthropic sets Medium as the default in Claude Code and the Claude apps for that reason.
OpenAI's speed answer is Ultrafast, a paid tier announced at DevDay that it says reaches up to 300 generated tokens per second: up to 8x faster in Codex and up to 6x faster through the API, with API usage billed at 6x standard pricing. At launch, Ultrafast was available for GPT-6 Astra (for Pro 500 and Enterprise) and described as coming within days for GPT-6.1 Sol. OpenAI has not published a separate Sol Ultrafast price; if the 6x rule applies, that would be about $12 in / $60 out per million tokens (approx. £9 / £45), which is our calculation, not an OpenAI figure. There is also a Fast mode at 2x standard price. Anthropic lists no fast mode for Sonnet 5.5; its Fast mode is offered on Opus 5.5.
What Does Each Model Cost per Task?
Identical list prices do not mean identical bills. The worked examples below use each vendor's published rates and assume the same token counts for both models, so they isolate the pricing structure. In reality, token counts differ by model and effort level, which is covered in the fourth example. Sterling figures use approx. £0.75 per $1 and exclude VAT.
| Scenario | Token assumptions | Claude Sonnet 5.5 | GPT-6.1 Sol |
|---|---|---|---|
| 1. Agentic coding session | 3M cached-read tokens, 300K fresh input, 60K output | $0.60 + $0.60 + $0.60 = $1.80 (approx. £1.35) | $0.30 + $0.60 + $0.60 = $1.50 (approx. £1.13) |
| 2. 10,000 support tickets | Per ticket: 2,000 cached system-prompt tokens, 1,000 fresh input, 400 output | $4 + $20 + $40 = $64 (approx. £48) | $2 + $20 + $40 = $62 (approx. £46.50) |
| 3. One 400K-token contract pack | 400K fresh input, 5K output, no caching | $0.80 + $0.05 = $0.85 (approx. £0.64) | $1.60 + $0.075 = $1.68 (approx. £1.26), with the long-context surcharge |
| 4. Artificial Analysis index task (measured) | Max effort; Sonnet ~193K output tokens, Sol ~38K | $7.60 (approx. £5.70) | $0.72 (approx. £0.54) |
Examples 1 to 3 are our illustrative calculations from list prices (cache-write charges, which are $2.50 per million on both, are excluded for simplicity). Example 4 is measured by Artificial Analysis.
The pattern is clear. When the workload is cache-heavy (examples 1 and 2), GPT-6.1 Sol's $0.10 cache reads give it a modest 3% to 17% edge at equal token counts. When a single prompt is very long (example 3), Sonnet 5.5 is about half the price because OpenAI surcharges anything over 272K tokens. And when you let both models think as hard as they like (example 4), token volume swamps the rate card: Sonnet 5.5 at max effort cost more than ten times as much per task.
That last point cuts both ways. Anthropic says Sonnet 5.5 typically costs up to 30% less per task than Sonnet 5, and that at Low or Medium effort it beats Sonnet 5's best scores on several benchmarks for about a tenth of the cost. On AutomationBench, Sonnet 5.5 at Xhigh cost $0.48 per task for 36.83%, broadly matching OpenAI's reported ~36% for Sol at max. The biggest lever on your bill is the effort setting, not the vendor. Batch processing halves costs on both platforms for anything that is not latency-sensitive.
How Do Their Safety Profiles and Safeguards Differ?
Both companies treat these mid-tier models as capable enough to need frontier-grade cyber safeguards, but OpenAI classifies GPT-6.1 Sol at a higher formal risk tier.
Claude Sonnet 5.5. Anthropic's system card says Sonnet 5.5 is broadly less capable than Opus 5.5 and crosses no new Responsible Scaling Policy thresholds. Because its cyber capability is now comparable to Opus 5's, it is the first Sonnet model with cyber safeguards and fallbacks: routine bug-finding and fixing is unaffected, but higher-risk cyber tasks visibly fall back to Sonnet 5, with a Cyber Verification Program for approved defenders. Biology safeguards are unchanged from Sonnet 5. It is also the first Sonnet launched with anti-distillation classifiers and "preserved thinking", which ties Claude's reasoning to the account that created it. On alignment, Anthropic reports it matches or beats Sonnet 5 on most measures, with the lowest sandbox-escape and boundary-probing rates of any model tested, but flags that its thinking is "more illegible" than many earlier models and that multi-turn tests showed regressions in areas such as tracking and surveillance.
GPT-6.1 Sol. OpenAI's system card addendum treats GPT-6.1 Sol as Critical in cybersecurity and High in biological and chemical capability under its Preparedness Framework, below High for AI self-improvement, and ships it with the same safeguards stack as GPT-6 Astra. Reported results include a 1.50% coding-deception rate (versus 0.51% for GPT-6 Astra), 35.1% on ExploitGym and 21.5% arbitrary code execution on ExploitBench. Context matters here: OpenAI also shelved a planned GPT-6.1 Astra release; see why GPT-6.1 Astra was cancelled and our explainer on Astra's Critical cyber rating.
For most business users, the practical difference is how refusals feel. Sonnet 5.5 hands higher-risk cyber prompts to an older model rather than refusing outright; GPT-6.1 Sol inherits Astra's trusted-access approach. Security teams doing legitimate offensive work should apply to the relevant verification programme on either platform rather than fight the defaults.
Matthew Berman walks through what his team built with Claude Sonnet 5.5, headlined by the strategy game Crownfall.
Claude Code vs Codex: Which Ecosystem Fits?
For many teams, the harness matters as much as the model: Sonnet 5.5 is the everyday workhorse inside Claude Code, and GPT-6.1 Sol is the new default-class model inside OpenAI's Codex.
Claude Code and the Claude apps. Sonnet 5.5 runs at Medium effort by default in Claude Code and the Claude apps, and Simon Willison notes it now powers the free tier on claude.ai. On the API, it is available on Anthropic's own platform plus Amazon Bedrock, Google Cloud and Microsoft Foundry, which matters for UK and EU organisations with existing cloud commitments. Developers moving from Sonnet 5 should read the migration guide: forced tool use now returns an error, non-default temperature settings return a 400, and turning off up-front thinking now uses a between_tools setting. See our Claude Code guide and Claude Code pricing for plan costs.
Codex and ChatGPT Work. GPT-6.1 Sol is available to ChatGPT Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, but not yet in standard ChatGPT chat. DevDay also brought Codex cloud tasks that keep running with your laptop closed, two-way voice and an /agents view in the CLI, a Code Review feature and a Security Cloud scanner, plus Ultrafast generation in Codex. On the API, tool calling requires the Responses API. For the pre-DevDay state of play, see OpenAI Codex August 2026 updates.
If your team already lives in one ecosystem, the switching cost is usually higher than the model gap. If you are starting fresh, run a one-week trial of each agent on the same backlog.
Which Should You Choose for Your Use Case?
| Use case | Pick | Why |
|---|---|---|
| Agentic coding in the terminal or IDE | Claude Sonnet 5.5 | Leads Terminal-Bench 4.0 and SciCode independently; strong CursorBench |
| High-volume repo maintenance on a budget | GPT-6.1 Sol | 75.2% DeepSWE claim, far fewer tokens per task |
| Business workflow automation across SaaS tools | Claude Sonnet 5.5 | Tops Zapier AutomationBench; leads AutomationBench-AA |
| Reports, decks, spreadsheets, long writing | Claude Sonnet 5.5 | Over 250 Elo ahead on GDPval-AA; clearer writing |
| Document Q&A and extraction from PDFs | GPT-6.1 Sol | Leads GDP.pdf and AA-Omniscience |
| Cache-heavy chat or support bots | GPT-6.1 Sol | $0.10 cache reads, half Sonnet's rate |
| Single prompts over 272K tokens | Claude Sonnet 5.5 | No long-context surcharge across 1M tokens |
| Lowest latency at standard price | Claude Sonnet 5.5 | About 2x the tokens per second of Sol |
| Absolute fastest generation, cost no object | GPT-6.1 Sol (Ultrafast, when live) | Up to 300 tokens per second at 6x price |
| Multi-cloud deployment (AWS, Google, Microsoft) | Claude Sonnet 5.5 | Available on Bedrock, Google Cloud and Microsoft Foundry at launch |
If you need more capability than either offers, step up to Claude Opus 5.5 or GPT-6 Astra. If you are comparing last week's models, our GPT-6 Sol vs Claude Opus 5.5 and GPT-6 Sol vs Astra vs Luna guides still apply, and the original GPT-6 Sol and Luna review covers the model GPT-6.1 Sol replaces. Want to try Anthropic's lineup directly? See the Claude Opus 5.5 and Claude Sonnet 5 tool pages, or browse creator coverage on our videos page.
Sources
- Anthropic: Introducing Claude Sonnet 5.5 (28/09/2026)
- Claude Platform docs: Claude Sonnet 5.5 overview
- Anthropic: Claude Sonnet 5.5 System Card
- OpenAI: Introducing GPT-6.1 Sol (29/09/2026)
- OpenAI: DevDay 2026 recap
- OpenAI developer docs: GPT-6.1 Sol
- OpenAI Deployment Safety Hub: GPT-6.1 Sol system card addendum
- Artificial Analysis: Claude Sonnet 5.5 vs GPT-6.1 Sol
- Artificial Analysis: Claude Sonnet 5.5 reaches #2 on the Intelligence Index
- Zapier: AutomationBench leaderboard
- TechCrunch: OpenAI launches GPT-6.1 Sol
- VentureBeat: GPT-6.1 Sol and the Ultrafast tier
- VentureBeat: Anthropic launches Claude Sonnet 5.5
- SiliconANGLE: Anthropic debuts Claude Sonnet 5.5
- The Next Web: GPT-6.1 Sol at a fifth of Astra's prices
- Simon Willison: Claude Sonnet 5.5
- Vellum: GPT-6.1 Sol benchmarks explained
The Bottom Line
Claude Sonnet 5.5 and GPT-6.1 Sol are the most evenly matched pair of mid-tier models yet: same price, same class of context window, same 128K output ceiling, released a day apart. On the first independent evidence, Sonnet 5.5 is the more capable and faster model, with clear leads on agentic coding, knowledge work and workflow automation. GPT-6.1 Sol is the more economical model, with half-price cache reads, better factual reliability and dramatically lower token use at high effort.
Our recommendation: use Sonnet 5.5 where output quality and turnaround time drive value, and cap its effort level to keep costs in line. Use GPT-6.1 Sol for high-volume, cache-heavy or extraction-style work where every penny per task counts. And because neither vendor benchmarked against the other's new model, run your own evaluation before you commit a production workload.
Last updated: 29/09/2026. Based on Anthropic's and OpenAI's official announcements, documentation and system cards, Artificial Analysis's independent evaluations and Zapier's AutomationBench leaderboard. GPT-6.1 Sol launched today, so independent results will fill in over the coming weeks; LMArena rankings were not yet available for either model.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.





