AI Tools Review
Frontier AI Benchmarks Compared: Opus 5.5, GPT-6 & More

Analysis

Frontier AI Benchmarks Compared: Opus 5.5, GPT-6 & More

AI Tools Review Editorial Team23 September 2026
  • Benchmarks
  • Claude Opus 5.5
  • GPT-6 Astra
  • GPT-6 Sol

On 22 September 2026, Anthropic launched Claude Opus 5.5 and, about 90 minutes later, OpenAI launched GPT-6 Sol and GPT-6 Luna. GPT-6 Astra had shipped three weeks earlier. Each launch came with a benchmark table. None of those tables was built to be read alongside the others.

This page is a benchmark matrix with one rule: a number only sits next to another number if both came from the same table. Where that is not true, we say "not comparable" rather than guess. For model-by-model detail, see our Claude Opus 5.5 review and GPT-6 Astra review.

Free Guide

Choosing between Claude, ChatGPT, Gemini and Grok?

Benchmarks are one input. Our free 20-page guide compares the big four on price, features and what each is actually good at, so you can match a model to your work.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

The Verdict by Category

CategoryLeader on available evidenceKey figuresReported by
Agentic codingClaude Opus 5.5Terminal-Bench 4.0 66.4%, FrontierCode 54.4%, CursorBench 57.8%Anthropic
Knowledge workClaude Opus 5.5GDPval-AA v2.1 1846 Elo (Astra 1542)Anthropic
Business automationGPT-6 Astra ≈ Claude Opus 5.5AutomationBench 41.4% vs 40.0%Anthropic table (Zapier benchmark)
Scientific researchGPT-6 AstraTerminal-Bench-Science 0.1 64.6% (Opus 5.5 58.7%)Anthropic table
Maths & science reasoningGPT-6 AstraFrontierMath Tier 4 v2 97.6, GPQA Diamond 96.0OpenAI (no Opus 5.5 figure)
Independent compositeGPT-6 Astra = Claude Fable 5.1AA Intelligence Index 53, 53; Sol 48; Grok 4.7 46Artificial Analysis (Opus 5.5 not yet indexed)

Sources: Anthropic, 22 September 2026; OpenAI via DataCamp; Artificial Analysis. Checked 23 September 2026.

Two things stand out. First, GPT-6 Sol does not lead any category. It is not built to: at $2/$10 per million tokens it is priced against mid-tier models, not flagships (see our AI API pricing comparison). Second, the maths row has no Claude 5.5 challenger at all, because Anthropic did not publish GPQA or FrontierMath for Opus 5.5 in its launch table. An empty cell is not a loss. It is missing data.

Matrix A: Anthropic's Opus 5.5 Launch Table

This is the most useful single table right now because it holds five models across nine benchmarks. Claude figures are at max effort with production safeguards on. The GPT figures are, in Anthropic's words, "as reported by OpenAI". Anthropic did not rerun them.

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0 (agentic coding)66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 Main (coding)54.4%50.3%48.0%53.3%47.5%
CursorBench 4.0 (coding)57.8%51.8%46.6%41.7%
GDPval-AA v2.1 (knowledge work, Elo)18461735170815421588
AutomationBench (Zapier)40.0%31.4%26.9%41.4%28.8%
Humanity's Last Exam (with tools)67.7%65.6%63.6%57.2%
Terminal-Bench-Science 0.158.7%52.6%29.0%64.6%22.4%
OSWorld 2.0 (partial credit)81.8%80.7%74.0%
Chartography (with tools)89.0%88.4%83.4%

Source: Anthropic, 22 September 2026. Bold marks the row leader. Claude models at max effort with safeguards on; GPT figures as reported by OpenAI. Terminal-Bench 4.0: Opus 5.5 at xhigh effort, GPT-6 Astra at high; standard error ±2.6 pts for Opus 5.5. — = not reported in this table.

Opus 5.5 leads seven of nine rows; Astra leads two. Anthropic also claims cost advantages: at default effort Opus 5.5 beats Astra on FrontierCode at about 20% of the cost per task, and matches Astra on Terminal-Bench 4.0 at about 40%. Those are Anthropic's claims, not independent measurements.

Terminal-Bench 4.0: agentic coding

Anthropic's launch table only. Opus 5.5 at xhigh, GPT-6 Astra at high. Higher is better.

  • OpenAI
  • Anthropic

Data table

Terminal-Bench 4.0: agentic coding: data table
ModelVendorValue
Claude Opus 5.5 (xhigh, ±2.6 SE)Anthropic66.4%
GPT-6 Astra (high, as reported by OpenAI)OpenAI57.9%
Claude Fable 5.1Anthropic55.8%
Claude Opus 5Anthropic52.3%
GPT-5.6 Sol (as reported by OpenAI)OpenAI37.3%

Source: Anthropic Claude Opus 5.5 launch table, 22 September 2026. As of 23 September 2026.

GDPval-AA v2.1: professional knowledge work (Elo)

Anthropic's launch table only. Claude at max effort. Higher is better.

  • OpenAI
  • Anthropic

Data table

GDPval-AA v2.1: professional knowledge work (Elo): data table
ModelVendorValue
Claude Opus 5.5Anthropic1,846
Claude Fable 5.1Anthropic1,735
Claude Opus 5Anthropic1,708
GPT-5.6 Sol (as reported by OpenAI)OpenAI1,588
GPT-6 Astra (as reported by OpenAI)OpenAI1,542

Source: Anthropic Claude Opus 5.5 launch table, 22 September 2026. As of 23 September 2026.

Matrix B: OpenAI's GPT-6 Sol/Luna Launch Table

OpenAI compared Sol and Luna against Claude Opus 5 and Fable 5 / 5.1, not Opus 5.5. That is not evasion: Opus 5.5 launched about 90 minutes earlier. It does mean this table cannot tell you how Sol compares with Opus 5.5, except on AutomationBench (see the bridge). Effort levels are shown in brackets because OpenAI varied them by model.

BenchmarkGPT-6 SolGPT-6 LunaGPT-6 AstraGPT-5.6 SolFable 5.1Fable 5Opus 5
AutomationBench 1.0.633.2% (xhigh)30.3% (low)31.4%26.9%
DeepSWE v1.168.8% (max)66.6% (max)69.9% (xhigh)66.0% (med)
OSWorld 2.0 (offline, OpenAI harness)60.5% (xhigh)58.1% (max)72.6%57.0% (med)60.3% (med)
Agents' Last Exam56.4% (max)59.3%53.6%

Source: OpenAI, 22 September 2026, via VentureBeat and Vellum. — = not reported in this table. Effort level shown where OpenAI stated it.

The story OpenAI tells here is cost, not rank. On DeepSWE, Sol (68.8%) sits just behind Fable 5 (69.9%) at what OpenAI says is about 80% lower cost per task. On AutomationBench, Sol costs $0.27 per task, which OpenAI says is 8.9x cheaper than Fable 5.1 and 11.1x cheaper than Opus 5. OpenAI also reports Sol's deception rate on an adversarial coding test fell to 1.3% from 10.4% for its predecessor. For the full Sol and Luna picture, see GPT-6 Sol and Luna explained and GPT-6 Sol vs Claude Opus 5.5.

Matrix C: OpenAI's GPT-6 Astra Card

BenchmarkGPT-6 AstraGPT-5.6 SolFable 5.1Opus 5Gemini 3.8 Flash
GPQA Diamond96.094.693.793.795.3
FrontierMath Tier 4 v297.683.087.873.2
ExploitBench100.0
SRE-Bench88.0
ScreenSpot-Pro92.7

Source: OpenAI GPT-6 Astra release, via DataCamp. All figures OpenAI-reported. — = not reported. Opus 5.5 did not exist when this card was published.

Astra's card is strongest on reasoning. FrontierMath Tier 4 at 97.6 is 9.8 points clear of Fable 5.1 and 24.4 clear of Opus 5 in the same table. GPQA Diamond is tighter: Astra 96.0, Gemini 3.8 Flash 95.3, and both Claude models 93.7. On long context, OpenAI reports 100% on MRCR v2 8-needle at 256K-512K tokens and 96.3% at 512K-1M.

ExploitBench at 100.0 is a capability flag as much as a score. OpenAI rates Astra "Critical" for cyber under its Preparedness Framework. Our Claude Opus 5.5 vs GPT-6 Astra comparison covers what that means in practice.

Free Guide

Benchmarks tell you who won a test. Which one fits your work?

Get our free 20-page guide comparing Claude, ChatGPT, Gemini and Grok on price, features and what each is good at. It was written before this week's launches, so pair it with the matrices on this page.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

The Traps

1. OSWorld 2.0 means two different things

OSWorld 2.0 appears in both vendors' tables. Anthropic scores it with partial credit. OpenAI runs an offline version in its own harness. The proof that these are different measurements is Claude Opus 5, which appears in both: 74.0% in Anthropic's table (max effort) and 60.3% in OpenAI's (medium effort). Same model, same benchmark name, 13.7 points apart.

Now imagine the naive comparison: Opus 5.5 81.8% (Anthropic) against GPT-6 Astra 72.6% (OpenAI), and "Opus 5.5 wins computer use by 9 points". That headline would be built on two different scoring methods. The gap between harnesses for Opus 5 alone is bigger than the gap between the two models. We mark OSWorld 2.0 as not comparable across vendors.

2. ARC-AGI-3 at 99.9% needs an asterisk

GPT-6 Astra's 99.9% on ARC-AGI-3 was achieved only with an expensive stateful adapter harness. On ARC Prize's standard stateless harness it scores roughly 17% to 63% depending on reasoning tier. Quote the 99.9% without the harness and you overstate it by a wide margin.

3. Effort levels differ inside the same chart

On OpenAI's AutomationBench chart, GPT-6 Sol runs at xhigh effort (33.2%) while GPT-6 Astra runs at low (30.3%). That makes Sol look like it beats Astra. In Anthropic's table, Astra scores 41.4%. Effort settings change results by double digits, and vendors choose them.

4. Vendor-run is not independent

Every number in Matrices A, B and C was published by a model vendor. Anthropic's table takes its GPT figures from OpenAI's own reporting rather than rerunning them. The only independent composite here is Artificial Analysis, and it has not indexed Opus 5.5 yet.

5. Anthropic says margins matter less now

Anthropic itself says that at this level of capability, benchmark margins are a less reliable guide to real-world differences, and that in its own use the gap between Opus 5.5 and Fable 5.1 is narrower than its scores suggest. When a vendor tells you to discount its own lead, take it seriously.

The One Bridge: AutomationBench

Zapier's AutomationBench is the only benchmark where the two vendors' tables agree on shared anchors. Claude Opus 5 scores 26.9% and Fable 5.1 scores 31.4% in both Anthropic's and OpenAI's tables. Matching anchors mean the scores come from a consistent run, so we can place GPT-6 Sol (from OpenAI) on the same scale as Opus 5.5 (from Anthropic). That gives the only fair six-model ranking available this week.

AutomationBench: business workflow automation

Combined from Anthropic's and OpenAI's tables, which agree on Opus 5 (26.9) and Fable 5.1 (31.4). Higher is better.

  • OpenAI
  • Anthropic

Data table

AutomationBench: business workflow automation: data table
ModelVendorValue
GPT-6 Astra (Anthropic's table)OpenAI41.4%
Claude Opus 5.5 (Anthropic's table)Anthropic40.0%
GPT-6 Sol (xhigh, OpenAI's table)OpenAI33.2%
Claude Fable 5.1 (both tables)Anthropic31.4%
GPT-5.6 Sol (Anthropic's table)OpenAI28.8%
Claude Opus 5 (both tables)Anthropic26.9%

Source: Anthropic Opus 5.5 launch table and OpenAI GPT-6 Sol launch (AutomationBench 1.0.6, via VentureBeat/Vellum), 22 September 2026. As of 23 September 2026.

The result is a clear tier. Astra and Opus 5.5 are 1.4 points apart at the top. Sol is about 7 points behind Opus 5.5 but, per OpenAI, runs at $0.27 per task. One caveat: our sources do not state the effort level behind Astra's 41.4%, and OpenAI's own chart shows Astra at low effort scoring 30.3%. We used 41.4% because it is the figure in the same table as Opus 5.5.

Independent: Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index (v4.3.x)

Independent composite score. Opus 5.5 not yet indexed. Higher is better.

  • OpenAI
  • Anthropic
  • xAI

Data table

Artificial Analysis Intelligence Index (v4.3.x): data table
ModelVendorValue
GPT-6 AstraOpenAI53
Claude Fable 5.1Anthropic53
GPT-6 Sol (max effort)OpenAI48
Grok 4.7xAI46

Source: Artificial Analysis Intelligence Index v4.3.x. As of 23 September 2026.

This is the one ranking not produced by a vendor. Astra and Fable 5.1 tie at 53. Sol at max effort scores 48, and Artificial Analysis measured it at 131.2 output tokens per second with a blended price of $1.54 per million tokens. Grok 4.7 scores 46. There is no Opus 5.5 score in any source we could verify, so we do not show one. We will add it when it is published.

Gemini and Grok: What We Have, and What We Don't

Gemini. Our only figure for Google's current model is from OpenAI's Astra card: Gemini 3.8 Flash scores 95.3 on GPQA Diamond, ahead of both Claude models in that table (93.7) and behind Astra (96.0). That is strong for a Flash-tier model. Google's Pro tier is stuck: Gemini 3.5 Pro was announced at I/O on 19 May 2026 and is still unreleased after missing June, July and August targets. The live Pro model is Gemini 3.1 Pro, which appears in none of these tables. Gemini 4 has confirmed pre-training but no release.

Grok. Grok 4.7 launched on 21 September 2026. Its only comparable number is 46 on the Artificial Analysis index, two points behind GPT-6 Sol. Neither Anthropic nor OpenAI included it in their tables. Independent testers describe it as slow and verbose. For how the four families compare beyond benchmarks, see Claude vs ChatGPT vs Gemini vs Grok.

Where Each Leader Loses

  • Do not choose Claude Opus 5.5 on benchmarks if your work is scientific research workflows (Astra 64.6% vs 58.7% on Terminal-Bench-Science) or maths-heavy reasoning, where Anthropic published no GPQA or FrontierMath figure. It also has no independent composite score yet.
  • Do not choose GPT-6 Astra on benchmarks if your work is coding or professional documents. It trails Opus 5.5 on Terminal-Bench 4.0 (57.9% vs 66.4%), GDPval-AA (1542 vs 1846) and Humanity's Last Exam with tools (57.2% vs 67.7%). It also costs $10/$50 per million tokens against Opus 5.5's $4/$20 (see Claude API pricing).
  • Do not choose GPT-6 Sol expecting flagship scores. It trails Opus 5.5 by about 7 points on AutomationBench and Astra on Agents' Last Exam (56.4% vs 59.3%). Its case is price.
  • Do not use this matrix at all if your task is not on it. None of these benchmarks tests your data, your prompts or your latency budget. Run your own evaluation on the two or three models that lead your category.

Methodology

  • Same-table rule. Numbers are only compared when they come from the same published table. Cross-table comparison is allowed only where shared anchor models score identically in both (AutomationBench: Opus 5 and Fable 5.1).
  • Attribution. Every figure is labelled by who reported it: Anthropic, OpenAI, or Artificial Analysis. Competitor figures inside a vendor table are marked "as reported by" the rival.
  • Effort levels. Shown where the vendor stated them. Where not stated, we say so.
  • Exclusions. OSWorld 2.0 is not compared across vendors. ARC-AGI-3 is not charted because the headline figure depends on a non-standard harness. Figures from secondary sources are taken as reported by VentureBeat, Vellum and DataCamp.
  • No estimates. Blank cells stay blank. We do not impute an Opus 5.5 Artificial Analysis score or a Gemini 3.5 Pro result.
  • Checked: 23 September 2026.

Sources

Last updated: 23 September 2026. Benchmark figures are vendor-reported unless marked as Artificial Analysis. We will add Claude Opus 5.5's independent index score when it is published.

Free Guide

Get the free Claude vs ChatGPT, Gemini & Grok guide

A 20-page guide comparing the big four AI families on price, features and what each is good at. Use it with this matrix to shortlist the models worth testing on your own work.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

What is the best AI model in September 2026?
It depends on the task. On Anthropic's 22 September table, Claude Opus 5.5 leads agentic coding (Terminal-Bench 4.0 66.4%, FrontierCode 54.4%, CursorBench 57.8%) and knowledge work (GDPval-AA v2.1 1846 Elo). GPT-6 Astra leads scientific research (Terminal-Bench-Science 64.6%) and, on OpenAI-reported figures, maths and science reasoning (FrontierMath Tier 4 97.6, GPQA Diamond 96.0). On Artificial Analysis's independent Intelligence Index, GPT-6 Astra and Claude Fable 5.1 tie at 53; Opus 5.5 has not been indexed yet.
Is Claude Opus 5.5 better than GPT-6 Sol on benchmarks?
Only one benchmark lets you compare them fairly: AutomationBench, where Opus 5.5 scores 40.0% and GPT-6 Sol (xhigh) 33.2%. OpenAI compared Sol against Opus 5 and Fable 5.1, not Opus 5.5, because Opus 5.5 launched about 90 minutes before Sol. There is no other shared benchmark from the same harness.
Why can't I compare OSWorld 2.0 scores between Anthropic and OpenAI?
The two vendors score it differently. Anthropic reports OSWorld 2.0 with partial credit; OpenAI reports an offline version in its own harness. The same model, Claude Opus 5, scores 74.0% in Anthropic's table and 60.3% (at medium effort) in OpenAI's. A 14-point gap for one model shows the tables cannot be mixed.
Does Claude Opus 5.5 have an Artificial Analysis Intelligence Index score?
Not in any source we could verify as of 23 September 2026. The current index shows GPT-6 Astra 53, Claude Fable 5.1 53, GPT-6 Sol (max) 48 and Grok 4.7 46. Treat any Opus 5.5 index figure you see elsewhere with caution until Artificial Analysis publishes it.
Where do Gemini and Grok rank?
We have little comparable data. Gemini 3.8 Flash scores 95.3 on GPQA Diamond in OpenAI's GPT-6 Astra table, and Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index. Gemini 3.5 Pro was announced in May 2026 but is still unreleased, and Gemini 4 has not shipped.

Key takeaways

No single winner

Opus 5.5 leads seven of nine rows in Anthropic's table; Astra leads AutomationBench and Terminal-Bench-Science. Independent composite: Astra and Fable 5.1 tie at 53.

Tables do not mix

Claude Opus 5 scores 74.0% on OSWorld 2.0 in Anthropic's table and 60.3% in OpenAI's. Same model, same benchmark name, 14 points apart.

One clean bridge

AutomationBench is the only benchmark with consistent figures across both vendors, giving a fair six-model ranking that includes GPT-6 Sol and Opus 5.5.

AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.