Anthropic launched Claude Opus 5.5 on 22 September 2026, 18 days after GPT-6 Astra reached general availability on 4 September. Anthropic put Astra straight into its launch table. That makes this the first flagship head-to-head with numbers on both sides. It also means one vendor chose the benchmarks. Read the scores with that in mind.
This comparison uses only the published launch figures, list prices checked on 23 September 2026 and our own workload cost model. For the full single-model reviews, see our Claude Opus 5.5 review and our GPT-6 Astra review.
Choosing between Claude and ChatGPT?
Our free 20-page guide compares Claude, ChatGPT, Gemini and Grok on price, features and what each is good at. It predates this week's launches, but the platform-level trade-offs still apply.

Matt Wolfe's same-day reaction to both flagship launches landing within 90 minutes of each other.
The Verdict
- Benchmarks: six benchmarks carry scores for both models in Anthropic's table. Opus 5.5 leads four. Astra leads two.
- Biggest Opus 5.5 margins: GDPval-AA v2.1 knowledge work (1846 vs 1542 Elo), Humanity's Last Exam with tools (67.7% vs 57.2%) and Terminal-Bench 4.0 (66.4% vs 57.9%).
- Astra's wins: AutomationBench (41.4% vs 40.0%, a narrow gap) and Terminal-Bench-Science 0.1 (64.6% vs 58.7%, a clear gap).
- Price: Opus 5.5 is 40% of Astra on input and output tokens, and 20% on cached input.
- Context: both offer up to 1M tokens. Anthropic charges no long-context premium on Opus 5.5.
- Safety: OpenAI rates Astra "Critical" for cyber. Opus 5.5 is Anthropic's best-scoring model on its behavioural audit.
If you run coding agents or professional knowledge work, Opus 5.5 wins on both score and cost. If your work is scientific terminal tasks, hard maths or vetted security research, Astra has the stronger evidence.
Shared Benchmarks: Whose Table Is It?
Every head-to-head number below comes from Anthropic's Opus 5.5 launch table of 22 September 2026. Claude figures are at max effort with production safeguards switched on. Astra figures are listed "as reported by OpenAI", not re-run by Anthropic. OpenAI's own launches have not compared Astra with Opus 5.5. Its 22 September GPT-6 Sol announcement compared against Opus 5 and Fable 5.1, because Opus 5.5 had launched only about 90 minutes earlier.
| Benchmark | Claude Opus 5.5 | GPT-6 Astra | Leader |
|---|---|---|---|
| FrontierCode v1.1 Main (coding) | 54.4% | 53.3% | Opus 5.5 (+1.1) |
| Humanity's Last Exam (with tools) | 67.7% | 57.2% | Opus 5.5 (+10.5) |
| AutomationBench (Zapier) | 40.0% | 41.4% | Astra (+1.4) |
| Terminal-Bench-Science 0.1 | 58.7% | 64.6% | Astra (+5.9) |
| CursorBench 4.0 | 57.8% | — | No Astra score |
Source: Anthropic Opus 5.5 launch table, 22 September 2026. Astra figures as reported by OpenAI. Terminal-Bench 4.0 and GDPval-AA v2.1 are charted below. OSWorld 2.0 is excluded: Anthropic reports a partial-credit score for Opus 5.5 (81.8%) and OpenAI reports Astra on its own offline harness (72.6%). The two are not comparable.
Terminal-Bench 4.0 needs its own caveat. Anthropic compares Opus 5.5 at xhigh effort against Astra at high effort, each model's highest reported score. Anthropic gives a standard error of ±2.6 points for Opus 5.5. The 8.5-point lead is well outside that margin, but it is not a same-settings comparison.
Terminal-Bench 4.0: agentic coding in the terminal
Percentage of tasks solved. Opus 5.5 at xhigh effort; GPT-6 Astra at high effort. Higher is better.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| Claude Opus 5.5 (xhigh, ±2.6 SE) | Anthropic | 66.4% |
| GPT-6 Astra (high, as reported by OpenAI) | OpenAI | 57.9% |
| Claude Fable 5.1 | Anthropic | 55.8% |
| Claude Opus 5 | Anthropic | 52.3% |
| GPT-5.6 Sol (as reported by OpenAI) | OpenAI | 37.3% |
Source: Anthropic Claude Opus 5.5 launch table, 22 September 2026 (OpenAI figures as reported by OpenAI). As of 23 September 2026.
GDPval-AA v2.1 grades professional work across 44 occupations on an Elo scale. This is the widest gap in the table. Astra even scores below OpenAI's previous GPT-5.6 Sol here. Anthropic also claims Opus 5.5 at medium effort beats Astra at max effort on GDPval for about a fifth of the cost per task.
GDPval-AA v2.1: professional knowledge work
Elo rating across 44 occupations. Higher is better.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| Claude Opus 5.5 | Anthropic | 1,846 |
| Claude Fable 5.1 | Anthropic | 1,735 |
| Claude Opus 5 | Anthropic | 1,708 |
| GPT-5.6 Sol (as reported by OpenAI) | OpenAI | 1,588 |
| GPT-6 Astra (as reported by OpenAI) | OpenAI | 1,542 |
Source: Anthropic Claude Opus 5.5 launch table, 22 September 2026 (OpenAI figures as reported by OpenAI). As of 23 September 2026.
Anthropic also warns that at this level of capability, benchmark margins are a weaker guide to real-world differences than they used to be. The one-point gaps on FrontierCode and AutomationBench should be read as ties. For more context on how every frontier model scores this month, see our September 2026 frontier benchmarks roundup.
Where Astra Wins
This is where Opus 5.5 loses. Astra's two shared-benchmark wins are real. Terminal-Bench-Science 0.1 tests research workflows in the terminal, and Astra's 5.9-point lead is clear. AutomationBench, run by Zapier on business workflows, is closer at 1.4 points. Anthropic notes Zapier ran it without fallback models, so safeguard interventions counted as Opus 5.5 failures.
OpenAI's own card also reports results that Anthropic did not publish for Opus 5.5. None of these can be compared directly, but they show where Astra is strongest:
- GPQA Diamond: 96.0. Graduate-level science questions. For reference, OpenAI lists Fable 5.1 and Opus 5 at 93.7.
- FrontierMath Tier 4 v2: 97.6. The hardest maths tier. OpenAI lists Fable 5.1 at 87.8 and Opus 5 at 73.2.
- ExploitBench: 100.0 and SRE-Bench: 88.0. Security and site reliability engineering.
- Long context: 100% on MRCR v2 8-needle between 256K and 512K tokens, and 96.3% between 512K and 1M (via DataCamp).
- OSWorld 2.0 (OpenAI's offline harness): 72.6%. Not comparable with Anthropic's partial-credit score for Opus 5.5.
One headline figure needs a warning. Astra's 99.9% on ARC-AGI-3 was achieved only with an expensive stateful adapter harness. On ARC Prize's standard stateless harness it scores roughly 17% to 63%, depending on reasoning tier. Do not quote 99.9% as a like-for-like result.
Price per Token
| Rate (USD per million tokens) | Claude Opus 5.5 | GPT-6 Astra | Opus 5.5 as % of Astra |
|---|---|---|---|
| Input | $4.00 | $10.00 | 40% |
| Output | $20.00 | $50.00 | 40% |
| Cached input (read) | $0.20 | $1.00 | 20% |
| Cache write | $5.00 (5-minute), $8.00 (1-hour) | $12.50 | 40% (5-minute) |
| Batch (input / output) | $2.00 / $10.00 | Not in our sources | — |
| Fast mode (input / output) | $8.00 / $40.00 | 2x standard price for up to 2.5x speed | — |
| Context window | 1M, no long-context premium | Up to 1M | — |
Checked 23 September 2026. Sources: Anthropic pricing documentation (Claude); OpenAI announcement via VentureBeat, Vellum and DataCamp (GPT-6 Astra).
The cached-input gap matters most for agents. Coding agents re-read the same repository context on every turn, so cache reads dominate their bills. Opus 5.5's cached rate is one fifth of Astra's. Opus 5.5 in fast mode ($8/$40) is still cheaper than Astra at standard rates ($10/$50). For the full Anthropic rate card, see our Claude API pricing guide. For every vendor side by side, see the AI API pricing comparison.
Weighing up AI costs for your team?
Get our free 20-page guide comparing Claude, ChatGPT, Gemini and Grok on price, features and what each is good at, a useful baseline before you commit to an API or subscription.

What It Costs per Month
Per-token rates hide how workloads mix input, output and cache hits. The AI Tools Review workload cost model applies each model's standard list price to three fixed monthly volumes:
- Chatbot: 10M input tokens (80% cache hits), 2M output.
- Coding agent: 50M input tokens (90% cache hits), 5M output.
- Bulk extraction: 100M input tokens (no caching), 10M output.
| Workload (per month) | Claude Opus 5.5 | GPT-6 Astra | Monthly saving with Opus 5.5 |
|---|---|---|---|
| Chatbot | $49.60 | $128.00 | $78.40 |
| Coding agent | $129.00 (approx £102) | $345.00 (approx £273) | $216.00 |
| Bulk extraction | $600.00 | $1,500.00 | $900.00 |
Source: AI Tools Review workload cost model, computed 23 September 2026 from list prices checked the same day. Excludes cache writes, batch discounts, fast mode and tool fees. GBP at £0.79 per $1.
On all three workloads Opus 5.5 costs 40% of Astra's bill or less. The chart below adds Anthropic's previous models on the coding-agent workload. Astra is the most expensive of the four, above even Fable 5.1, whose cheaper cached rate ($0.25) offsets its matching $10/$50 headline price.
Monthly cost: coding-agent workload
50M input tokens (90% cached) and 5M output per month at list prices. Sorted cheapest first. Lower is better.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| Claude Opus 5.5 | Anthropic | $129.00 |
| Claude Opus 5 | Anthropic | $172.50 |
| Claude Fable 5.1 | Anthropic | $311.25 |
| GPT-6 Astra | OpenAI | $345.00 |
Source: AI Tools Review workload cost model, from Anthropic and OpenAI list prices (excludes cache writes, batch, fast mode and tool fees). As of 23 September 2026.
One important caveat. This model measures per-token list cost. The two models do not use the same number of tokens for the same task, because tokenisers, verbosity and reasoning length all differ. Real cost per task can land either side of these figures. Anthropic's own cost-per-task claims point the same way as our model: it says Opus 5.5 matches Astra on Terminal-Bench 4.0 at about 40% of the cost, and beats it on FrontierCode at default effort for about 20% of the cost per task. Those are vendor claims, not independent measurements.
Safety & Access
The two labs have taken different routes with capable models, and it affects what you can actually do with each.
GPT-6 Astra reaches "Critical" for cyber under its Preparedness Framework. By default it refuses to create proof-of-concept exploits. Expanded security access runs through OpenAI's Daybreak programme. The UK AI Security Institute found Astra could evade monitoring under adversarial prompting, a regression in chain-of-thought monitorability. Our Astra review covers the Critical rating in full.
Claude Opus 5.5 is the first Opus model to ship with Fable-5.1-class safeguards on cyber, biology and distillation. Most cybersecurity tasks are re-routed to Claude Opus 4.8. Full-capability biology work needs Anthropic's Life Sciences Verification Program. Security researchers can apply to the Cyber Verification Program, which is expanding to three tiers. Anthropic reports Opus 5.5 has its best automated behavioural audit score to date, and made about 85% fewer attempts to cross containment boundaries than Opus 5 or Mythos 5.1. Thinking cannot be switched off, and zero data retention is available. More detail is in our Opus 5.5 review.
In practice, both models gate offensive security work behind a vetting programme. Neither is a drop-in tool for exploit development. If that is your job, the choice comes down to which programme accepts you.
Which One to Choose
Choose Claude Opus 5.5 if:
- You run coding agents. It leads Terminal-Bench 4.0 and FrontierCode, and its cached input is a fifth of Astra's price.
- You produce professional documents, analysis or reports. The GDPval-AA gap (1846 vs 1542 Elo) is the largest in the table.
- You need tool-using research and reasoning. It leads Humanity's Last Exam with tools by 10.5 points.
- Cost matters at scale. Our bulk-extraction workload saves $900 a month against Astra.
- You want the model with the stronger published alignment results.
Choose GPT-6 Astra if:
- Your work is scientific computing in the terminal. It leads Terminal-Bench-Science 0.1 by 5.9 points.
- You need frontier maths or graduate-level science. FrontierMath Tier 4 (97.6) and GPQA Diamond (96.0) are its strongest published results, though Anthropic has not published Opus 5.5 scores to compare.
- You do vetted security or SRE work and qualify for the Daybreak programme.
- You depend on very long documents and want OpenAI's published MRCR retrieval results across the full 1M window.
- You are already on Amazon Bedrock or ChatGPT Pro, Business or Enterprise, where Astra and Astra Pro are available.
Who Should Choose Neither
Many teams do not need a flagship. GPT-6 Sol and Claude Sonnet 5 both cost $2 input, $0.20 cached and $10 output per million tokens. That is half of Opus 5.5 and a fifth of Astra. On our coding-agent workload, either comes to $69 a month against $129 for Opus 5.5 and $345 for Astra.
On AutomationBench, the one benchmark with clean results for all three, OpenAI reports GPT-6 Sol at 33.2% at xhigh effort. Opus 5.5 scores 40.0% and Astra 41.4%. That gap matters for complex agents. It does not matter for customer support, drafting or routine extraction. Start with Sol or Sonnet 5 and move up only where your own tests show a failure. Our GPT-6 Sol vs Opus 5.5 comparison and GPT-6 Sol vs Astra vs Luna guide cover that decision. Consumer users should compare subscriptions instead, starting with our Claude pricing guide.
Sources
- Anthropic: Claude Opus 5.5 announcement and benchmark table
- Anthropic: Claude API pricing documentation
- OpenAI: GPT-6 Astra
- DataCamp: GPT-6 Astra benchmarks and pricing
- OpenAI: Introducing GPT-6 Sol and Luna
- VentureBeat: OpenAI releases GPT-6 Sol and Luna
- Vellum: GPT-6 Sol and Luna benchmarks explained
- TechCrunch: OpenAI launches GPT-6 Sol and Luna
Last updated: 23 September 2026. Head-to-head benchmark figures come from Anthropic's Opus 5.5 launch table, with GPT-6 Astra figures as reported by OpenAI. Astra-only figures come from OpenAI via DataCamp. Prices checked 23 September 2026. We will update this comparison when independent labs publish results for both models.
Still deciding between Claude and ChatGPT?
Download our free 20-page guide comparing Claude, ChatGPT, Gemini and Grok on price, features and what each is good at, so you can pick the right platform before choosing a model.






