On 22 September 2026, Anthropic launched Claude Opus 5.5 and, about 90 minutes later, OpenAI launched GPT-6 Sol and GPT-6 Luna. GPT-6 Astra had shipped three weeks earlier. Each launch came with a benchmark table. None of those tables was built to be read alongside the others.
This page is a benchmark matrix with one rule: a number only sits next to another number if both came from the same table. Where that is not true, we say "not comparable" rather than guess. For model-by-model detail, see our Claude Opus 5.5 review and GPT-6 Astra review.
Choosing between Claude, ChatGPT, Gemini and Grok?
Benchmarks are one input. Our free 20-page guide compares the big four on price, features and what each is actually good at, so you can match a model to your work.

The Verdict by Category
| Category | Leader on available evidence | Key figures | Reported by |
|---|---|---|---|
| Agentic coding | Claude Opus 5.5 | Terminal-Bench 4.0 66.4%, FrontierCode 54.4%, CursorBench 57.8% | Anthropic |
| Knowledge work | Claude Opus 5.5 | GDPval-AA v2.1 1846 Elo (Astra 1542) | Anthropic |
| Business automation | GPT-6 Astra ≈ Claude Opus 5.5 | AutomationBench 41.4% vs 40.0% | Anthropic table (Zapier benchmark) |
| Scientific research | GPT-6 Astra | Terminal-Bench-Science 0.1 64.6% (Opus 5.5 58.7%) | Anthropic table |
| Maths & science reasoning | GPT-6 Astra | FrontierMath Tier 4 v2 97.6, GPQA Diamond 96.0 | OpenAI (no Opus 5.5 figure) |
| Independent composite | GPT-6 Astra = Claude Fable 5.1 | AA Intelligence Index 53, 53; Sol 48; Grok 4.7 46 | Artificial Analysis (Opus 5.5 not yet indexed) |
Sources: Anthropic, 22 September 2026; OpenAI via DataCamp; Artificial Analysis. Checked 23 September 2026.
Two things stand out. First, GPT-6 Sol does not lead any category. It is not built to: at $2/$10 per million tokens it is priced against mid-tier models, not flagships (see our AI API pricing comparison). Second, the maths row has no Claude 5.5 challenger at all, because Anthropic did not publish GPQA or FrontierMath for Opus 5.5 in its launch table. An empty cell is not a loss. It is missing data.
Matrix A: Anthropic's Opus 5.5 Launch Table
This is the most useful single table right now because it holds five models across nine benchmarks. Claude figures are at max effort with production safeguards on. The GPT figures are, in Anthropic's words, "as reported by OpenAI". Anthropic did not rerun them.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 Main (coding) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 (coding) | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench (Zapier) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 (partial credit) | 81.8% | 80.7% | 74.0% | — | — |
| Chartography (with tools) | 89.0% | 88.4% | 83.4% | — | — |
Source: Anthropic, 22 September 2026. Bold marks the row leader. Claude models at max effort with safeguards on; GPT figures as reported by OpenAI. Terminal-Bench 4.0: Opus 5.5 at xhigh effort, GPT-6 Astra at high; standard error ±2.6 pts for Opus 5.5. — = not reported in this table.
Opus 5.5 leads seven of nine rows; Astra leads two. Anthropic also claims cost advantages: at default effort Opus 5.5 beats Astra on FrontierCode at about 20% of the cost per task, and matches Astra on Terminal-Bench 4.0 at about 40%. Those are Anthropic's claims, not independent measurements.
Terminal-Bench 4.0: agentic coding
Anthropic's launch table only. Opus 5.5 at xhigh, GPT-6 Astra at high. Higher is better.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| Claude Opus 5.5 (xhigh, ±2.6 SE) | Anthropic | 66.4% |
| GPT-6 Astra (high, as reported by OpenAI) | OpenAI | 57.9% |
| Claude Fable 5.1 | Anthropic | 55.8% |
| Claude Opus 5 | Anthropic | 52.3% |
| GPT-5.6 Sol (as reported by OpenAI) | OpenAI | 37.3% |
Source: Anthropic Claude Opus 5.5 launch table, 22 September 2026. As of 23 September 2026.
GDPval-AA v2.1: professional knowledge work (Elo)
Anthropic's launch table only. Claude at max effort. Higher is better.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| Claude Opus 5.5 | Anthropic | 1,846 |
| Claude Fable 5.1 | Anthropic | 1,735 |
| Claude Opus 5 | Anthropic | 1,708 |
| GPT-5.6 Sol (as reported by OpenAI) | OpenAI | 1,588 |
| GPT-6 Astra (as reported by OpenAI) | OpenAI | 1,542 |
Source: Anthropic Claude Opus 5.5 launch table, 22 September 2026. As of 23 September 2026.
Matrix B: OpenAI's GPT-6 Sol/Luna Launch Table
OpenAI compared Sol and Luna against Claude Opus 5 and Fable 5 / 5.1, not Opus 5.5. That is not evasion: Opus 5.5 launched about 90 minutes earlier. It does mean this table cannot tell you how Sol compares with Opus 5.5, except on AutomationBench (see the bridge). Effort levels are shown in brackets because OpenAI varied them by model.
| Benchmark | GPT-6 Sol | GPT-6 Luna | GPT-6 Astra | GPT-5.6 Sol | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|---|---|---|
| AutomationBench 1.0.6 | 33.2% (xhigh) | — | 30.3% (low) | — | 31.4% | — | 26.9% |
| DeepSWE v1.1 | 68.8% (max) | 66.6% (max) | — | — | — | 69.9% (xhigh) | 66.0% (med) |
| OSWorld 2.0 (offline, OpenAI harness) | 60.5% (xhigh) | 58.1% (max) | 72.6% | 57.0% (med) | — | — | 60.3% (med) |
| Agents' Last Exam | 56.4% (max) | — | 59.3% | 53.6% | — | — | — |
Source: OpenAI, 22 September 2026, via VentureBeat and Vellum. — = not reported in this table. Effort level shown where OpenAI stated it.
The story OpenAI tells here is cost, not rank. On DeepSWE, Sol (68.8%) sits just behind Fable 5 (69.9%) at what OpenAI says is about 80% lower cost per task. On AutomationBench, Sol costs $0.27 per task, which OpenAI says is 8.9x cheaper than Fable 5.1 and 11.1x cheaper than Opus 5. OpenAI also reports Sol's deception rate on an adversarial coding test fell to 1.3% from 10.4% for its predecessor. For the full Sol and Luna picture, see GPT-6 Sol and Luna explained and GPT-6 Sol vs Claude Opus 5.5.
Matrix C: OpenAI's GPT-6 Astra Card
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Fable 5.1 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|
| GPQA Diamond | 96.0 | 94.6 | 93.7 | 93.7 | 95.3 |
| FrontierMath Tier 4 v2 | 97.6 | 83.0 | 87.8 | 73.2 | — |
| ExploitBench | 100.0 | — | — | — | — |
| SRE-Bench | 88.0 | — | — | — | — |
| ScreenSpot-Pro | 92.7 | — | — | — | — |
Source: OpenAI GPT-6 Astra release, via DataCamp. All figures OpenAI-reported. — = not reported. Opus 5.5 did not exist when this card was published.
Astra's card is strongest on reasoning. FrontierMath Tier 4 at 97.6 is 9.8 points clear of Fable 5.1 and 24.4 clear of Opus 5 in the same table. GPQA Diamond is tighter: Astra 96.0, Gemini 3.8 Flash 95.3, and both Claude models 93.7. On long context, OpenAI reports 100% on MRCR v2 8-needle at 256K-512K tokens and 96.3% at 512K-1M.
ExploitBench at 100.0 is a capability flag as much as a score. OpenAI rates Astra "Critical" for cyber under its Preparedness Framework. Our Claude Opus 5.5 vs GPT-6 Astra comparison covers what that means in practice.
Benchmarks tell you who won a test. Which one fits your work?
Get our free 20-page guide comparing Claude, ChatGPT, Gemini and Grok on price, features and what each is good at. It was written before this week's launches, so pair it with the matrices on this page.

The Traps
1. OSWorld 2.0 means two different things
OSWorld 2.0 appears in both vendors' tables. Anthropic scores it with partial credit. OpenAI runs an offline version in its own harness. The proof that these are different measurements is Claude Opus 5, which appears in both: 74.0% in Anthropic's table (max effort) and 60.3% in OpenAI's (medium effort). Same model, same benchmark name, 13.7 points apart.
Now imagine the naive comparison: Opus 5.5 81.8% (Anthropic) against GPT-6 Astra 72.6% (OpenAI), and "Opus 5.5 wins computer use by 9 points". That headline would be built on two different scoring methods. The gap between harnesses for Opus 5 alone is bigger than the gap between the two models. We mark OSWorld 2.0 as not comparable across vendors.
2. ARC-AGI-3 at 99.9% needs an asterisk
GPT-6 Astra's 99.9% on ARC-AGI-3 was achieved only with an expensive stateful adapter harness. On ARC Prize's standard stateless harness it scores roughly 17% to 63% depending on reasoning tier. Quote the 99.9% without the harness and you overstate it by a wide margin.
3. Effort levels differ inside the same chart
On OpenAI's AutomationBench chart, GPT-6 Sol runs at xhigh effort (33.2%) while GPT-6 Astra runs at low (30.3%). That makes Sol look like it beats Astra. In Anthropic's table, Astra scores 41.4%. Effort settings change results by double digits, and vendors choose them.
4. Vendor-run is not independent
Every number in Matrices A, B and C was published by a model vendor. Anthropic's table takes its GPT figures from OpenAI's own reporting rather than rerunning them. The only independent composite here is Artificial Analysis, and it has not indexed Opus 5.5 yet.
5. Anthropic says margins matter less now
Anthropic itself says that at this level of capability, benchmark margins are a less reliable guide to real-world differences, and that in its own use the gap between Opus 5.5 and Fable 5.1 is narrower than its scores suggest. When a vendor tells you to discount its own lead, take it seriously.
The One Bridge: AutomationBench
Zapier's AutomationBench is the only benchmark where the two vendors' tables agree on shared anchors. Claude Opus 5 scores 26.9% and Fable 5.1 scores 31.4% in both Anthropic's and OpenAI's tables. Matching anchors mean the scores come from a consistent run, so we can place GPT-6 Sol (from OpenAI) on the same scale as Opus 5.5 (from Anthropic). That gives the only fair six-model ranking available this week.
AutomationBench: business workflow automation
Combined from Anthropic's and OpenAI's tables, which agree on Opus 5 (26.9) and Fable 5.1 (31.4). Higher is better.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| GPT-6 Astra (Anthropic's table) | OpenAI | 41.4% |
| Claude Opus 5.5 (Anthropic's table) | Anthropic | 40.0% |
| GPT-6 Sol (xhigh, OpenAI's table) | OpenAI | 33.2% |
| Claude Fable 5.1 (both tables) | Anthropic | 31.4% |
| GPT-5.6 Sol (Anthropic's table) | OpenAI | 28.8% |
| Claude Opus 5 (both tables) | Anthropic | 26.9% |
Source: Anthropic Opus 5.5 launch table and OpenAI GPT-6 Sol launch (AutomationBench 1.0.6, via VentureBeat/Vellum), 22 September 2026. As of 23 September 2026.
The result is a clear tier. Astra and Opus 5.5 are 1.4 points apart at the top. Sol is about 7 points behind Opus 5.5 but, per OpenAI, runs at $0.27 per task. One caveat: our sources do not state the effort level behind Astra's 41.4%, and OpenAI's own chart shows Astra at low effort scoring 30.3%. We used 41.4% because it is the figure in the same table as Opus 5.5.
Independent: Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index (v4.3.x)
Independent composite score. Opus 5.5 not yet indexed. Higher is better.
- OpenAI
- Anthropic
- xAI
Data table
| Model | Vendor | Value |
|---|---|---|
| GPT-6 Astra | OpenAI | 53 |
| Claude Fable 5.1 | Anthropic | 53 |
| GPT-6 Sol (max effort) | OpenAI | 48 |
| Grok 4.7 | xAI | 46 |
Source: Artificial Analysis Intelligence Index v4.3.x. As of 23 September 2026.
This is the one ranking not produced by a vendor. Astra and Fable 5.1 tie at 53. Sol at max effort scores 48, and Artificial Analysis measured it at 131.2 output tokens per second with a blended price of $1.54 per million tokens. Grok 4.7 scores 46. There is no Opus 5.5 score in any source we could verify, so we do not show one. We will add it when it is published.
Gemini and Grok: What We Have, and What We Don't
Gemini. Our only figure for Google's current model is from OpenAI's Astra card: Gemini 3.8 Flash scores 95.3 on GPQA Diamond, ahead of both Claude models in that table (93.7) and behind Astra (96.0). That is strong for a Flash-tier model. Google's Pro tier is stuck: Gemini 3.5 Pro was announced at I/O on 19 May 2026 and is still unreleased after missing June, July and August targets. The live Pro model is Gemini 3.1 Pro, which appears in none of these tables. Gemini 4 has confirmed pre-training but no release.
Grok. Grok 4.7 launched on 21 September 2026. Its only comparable number is 46 on the Artificial Analysis index, two points behind GPT-6 Sol. Neither Anthropic nor OpenAI included it in their tables. Independent testers describe it as slow and verbose. For how the four families compare beyond benchmarks, see Claude vs ChatGPT vs Gemini vs Grok.
Where Each Leader Loses
- Do not choose Claude Opus 5.5 on benchmarks if your work is scientific research workflows (Astra 64.6% vs 58.7% on Terminal-Bench-Science) or maths-heavy reasoning, where Anthropic published no GPQA or FrontierMath figure. It also has no independent composite score yet.
- Do not choose GPT-6 Astra on benchmarks if your work is coding or professional documents. It trails Opus 5.5 on Terminal-Bench 4.0 (57.9% vs 66.4%), GDPval-AA (1542 vs 1846) and Humanity's Last Exam with tools (57.2% vs 67.7%). It also costs $10/$50 per million tokens against Opus 5.5's $4/$20 (see Claude API pricing).
- Do not choose GPT-6 Sol expecting flagship scores. It trails Opus 5.5 by about 7 points on AutomationBench and Astra on Agents' Last Exam (56.4% vs 59.3%). Its case is price.
- Do not use this matrix at all if your task is not on it. None of these benchmarks tests your data, your prompts or your latency budget. Run your own evaluation on the two or three models that lead your category.
Methodology
- Same-table rule. Numbers are only compared when they come from the same published table. Cross-table comparison is allowed only where shared anchor models score identically in both (AutomationBench: Opus 5 and Fable 5.1).
- Attribution. Every figure is labelled by who reported it: Anthropic, OpenAI, or Artificial Analysis. Competitor figures inside a vendor table are marked "as reported by" the rival.
- Effort levels. Shown where the vendor stated them. Where not stated, we say so.
- Exclusions. OSWorld 2.0 is not compared across vendors. ARC-AGI-3 is not charted because the headline figure depends on a non-standard harness. Figures from secondary sources are taken as reported by VentureBeat, Vellum and DataCamp.
- No estimates. Blank cells stay blank. We do not impute an Opus 5.5 Artificial Analysis score or a Gemini 3.5 Pro result.
- Checked: 23 September 2026.
Sources
Last updated: 23 September 2026. Benchmark figures are vendor-reported unless marked as Artificial Analysis. We will add Claude Opus 5.5's independent index score when it is published.
Get the free Claude vs ChatGPT, Gemini & Grok guide
A 20-page guide comparing the big four AI families on price, features and what each is good at. Use it with this matrix to shortlist the models worth testing on your own work.






