On 22 September 2026, OpenAI released GPT-6 Sol and GPT-6 Luna, about 90 minutes after Anthropic launched Claude Opus 5.5 (TechCrunch). They complete the GPT-6 family that began with GPT-6 Astra on 3-4 September. The story here is price first and benchmarks second.
This review covers where each model sits, the full price table against GPT-5.6, every benchmark OpenAI published, the independent Artificial Analysis numbers, and who should skip both models.
Note: GPT-6 benchmark figures are OpenAI's own, as reported by VentureBeat and Vellum. Claude figures come from Anthropic's 22 September launch table. Independent figures come from Artificial Analysis. We attribute each number to its source and flag where results cannot be compared.
Choosing between ChatGPT and Claude?
Our free 20-page guide compares Claude, ChatGPT, Gemini and Grok on price, features and what each is good at. It predates this week's launches, but the trade-offs it explains still hold.

Matthew Berman's first-look review of the GPT-6 Sol and Luna launch, covering pricing and early benchmarks.
Summary
- Launched 22 September 2026 in ChatGPT Work, Codex and the API. Luna also reaches Free and Go users in the desktop app.
- Pricing: Sol $2 / $0.20 cached / $10. Luna $0.10 / $0.01 cached / $0.50. Both at least 50% below GPT-5.6, and permanent.
- Benchmarks (OpenAI): Sol 33.2% on AutomationBench 1.0.6, 68.8% on DeepSWE v1.1, 60.5% on OSWorld 2.0 offline, 56.4% on Agents' Last Exam.
- The caveat: OpenAI compared against Opus 5 and Fable 5.x, not Opus 5.5.
- Reliability: about half as many factual mistakes as GPT-5.6 Sol. Deception on an adversarial coding test fell from 10.4% to 1.3%.
- Independent: Artificial Analysis Intelligence Index puts Sol at 48, below Astra and Fable 5.1 (both 53) and above Grok 4.7 (46).
Luna, Sol, Astra: Where Each Sits
GPT-6 comes in three tiers. Each is roughly an order of magnitude apart on price.
- GPT-6 Luna is the high-volume model. Think classification, extraction, routing, summarising support tickets, and other clerical work where you process millions of items.
- GPT-6 Sol is the everyday model for complex work and coding. It is the default most developers and ChatGPT Work users will touch.
- GPT-6 Astra is for the hardest projects. It costs $10 / $50 per million tokens, five times Sol. See our GPT-6 Astra review.
The tiers overlap more than the prices suggest. On OpenAI's AutomationBench chart, Sol at xhigh effort (33.2%) beats Astra at low effort (30.3%). On Agents' Last Exam the gap between Sol and Astra is under three points. For a full three-way breakdown, read GPT-6 Sol vs Astra vs Luna.
Pricing: 50%+ Cuts, Permanent
| Model (USD per 1M tokens) | Input | Cached input | Output | Change vs predecessor |
|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | -50% / -60% / -50% |
| GPT-5.6 Sol (previous) | $4.00 | $0.50 | $20.00 | — |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | -50% / -50% / -58% |
| GPT-5.6 Luna (previous) | $0.20 | $0.02 | $1.20 | — |
| GPT-6 Astra (for reference) | $10.00 | $1.00 | $50.00 | — |
Source: OpenAI launch announcement via VentureBeat, Vellum and DataCamp. Checked 23 September 2026. Standard tier.
In sterling, Sol works out at approx £1.58 input and £7.90 output per million tokens. Luna is approx £0.08 input and £0.40 output. OpenAI says these are the new list prices, not a launch promotion.
The cached-input cut on Sol is the one to watch. It fell 60%, more than the headline 50%. Agents and chat apps that resend long system prompts and conversation history pay mostly at the cached rate, so their bills fall by more than half on that line.
Output price per 1M tokens
USD, standard tier list price. Lower is better; sorted cheapest first.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| GPT-6 Luna | OpenAI | $0.50 |
| GPT-5.6 Luna (previous) | OpenAI | $1.20 |
| GPT-6 Sol | OpenAI | $10.00 |
| Claude Sonnet 5 | Anthropic | $10.00 |
| GPT-5.6 Sol (previous, promotional rate; standard $30) | OpenAI | $20.00 |
| Claude Opus 5.5 | Anthropic | $20.00 |
| GPT-6 Astra | OpenAI | $50.00 |
Source: OpenAI via VentureBeat, Vellum and DataCamp; Anthropic pricing docs (platform.claude.com). As of 23 September 2026.
The Real Price Peer: Claude Sonnet 5
GPT-6 Sol's list price matches Claude Sonnet 5's to the cent: $2 input, $0.20 cached input, $10 output. Anthropic made Sonnet 5's $2/$10 price permanent on 11 August 2026, cancelling a planned rise to $3/$15. So OpenAI has priced its everyday model exactly on top of Anthropic's everyday model.
Most launch-day coverage lined Sol up against Opus 5.5, because both arrived on the same day. On price, that is the wrong pairing. Opus 5.5 costs $4 / $20, twice Sol on input and output. The two only match on cached reads, where both charge $0.20.
| Monthly workload (list price) | GPT-6 Luna | GPT-6 Sol | Claude Sonnet 5 | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|---|---|
| Chatbot (10M in, 80% cached; 2M out) | $1.28 | $25.60 | $25.60 | $49.60 | $128.00 |
| Coding agent (50M in, 90% cached; 5M out) | $3.45 | $69.00 | $69.00 | $129.00 | $345.00 |
| Bulk extraction (100M in, no cache; 10M out) | $15.00 | $300.00 | $300.00 | $600.00 | $1,500.00 |
Source: AI Tools Review workload cost model, computed from list prices checked 23 September 2026 (OpenAI via VentureBeat/Vellum/DataCamp; Anthropic pricing docs). Excludes cache writes, batch discounts, fast mode and tool fees.
This is per-token list cost. Models use different numbers of tokens for the same job, because tokenisers, verbosity and reasoning length differ. Two models with the same rate card can still produce different bills. There is also no published benchmark that puts Sol and Sonnet 5 side by side. If you are choosing between them, run both on your own tasks. Our AI API pricing comparison and Claude API pricing guide have the full rate cards.
Every Benchmark OpenAI Published
OpenAI published four headline evaluations. One caveat applies to all of them: OpenAI compared Sol and Luna with Claude Opus 5 and Fable 5 / 5.1, not with Opus 5.5. That is a timing issue, not necessarily a choice. Opus 5.5 went live about 90 minutes before Sol. It still means OpenAI's charts show Sol against Anthropic's previous generation.
AutomationBench 1.0.6
Zapier's benchmark of real business workflows. Sol at xhigh effort scored 33.2% at a cost of $0.27 per task. OpenAI says that is 8.9 times cheaper per task than Fable 5.1 and 11.1 times cheaper than Opus 5.
AutomationBench 1.0.6 (OpenAI's launch chart)
Task success, %. Higher is better. Opus 5.5 was not included.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| GPT-6 Sol (xhigh) | OpenAI | 33.2% |
| Claude Fable 5.1 | Anthropic | 31.4% |
| GPT-6 Astra (low effort) | OpenAI | 30.3% |
| Claude Opus 5 | Anthropic | 26.9% |
Source: OpenAI, as reported by Vellum and VentureBeat. As of 23 September 2026.
This is the one benchmark where a clean cross-check is possible. The Opus 5 (26.9%) and Fable 5.1 (31.4%) scores in OpenAI's chart match Anthropic's own table exactly. Anthropic's table also lists Opus 5.5 at 40.0% and GPT-6 Astra at 41.4% (Astra as reported by OpenAI). So on AutomationBench, Sol sits about seven points behind Opus 5.5 and eight behind Astra at its best setting. It wins on cost per task, not on score.
DeepSWE v1.1
A software-engineering benchmark. Sol at max effort scored 68.8%, just behind Fable 5 at xhigh (69.9%). Luna at max effort scored 66.6%, ahead of Opus 5 at medium (66.0%). OpenAI says Sol costs about 80% less per task than Fable 5, and Luna 93-96% less. Luna landing within about three points of Fable 5 is the most striking number in the launch.
DeepSWE v1.1
Task success, %. Higher is better. Effort setting in brackets.
- OpenAI
- Anthropic
Data table
| Model | Vendor | Value |
|---|---|---|
| Claude Fable 5 (xhigh) | Anthropic | 69.9% |
| GPT-6 Sol (max) | OpenAI | 68.8% |
| GPT-6 Luna (max) | OpenAI | 66.6% |
| Claude Opus 5 (medium) | Anthropic | 66.0% |
Source: OpenAI, as reported by Vellum and VentureBeat. As of 23 September 2026.
Weighing up ChatGPT against Claude, Gemini and Grok?
Get our free 20-page guide comparing the four big assistants on price, features and what each does well. A useful baseline before you test Sol or Luna yourself.

OSWorld 2.0 (offline) and Agents' Last Exam
OSWorld tests computer use: operating desktop apps to finish tasks. OpenAI ran it on its own offline harness. Those scores cannot be compared with Anthropic's OSWorld 2.0 partial-credit numbers (Opus 5.5 at 81.8%), because the harness and scoring differ. Agents' Last Exam is an agentic benchmark.
| Model | OSWorld 2.0 (offline, OpenAI harness) | Agents' Last Exam |
|---|---|---|
| GPT-6 Astra | 72.6% | 59.3% |
| GPT-6 Sol | 60.5% (xhigh) | 56.4% (max) |
| Claude Opus 5 | 60.3% (medium) | — |
| GPT-6 Luna | 58.1% (max) | — |
| GPT-5.6 Sol (promotional; standard $5/$30) | 57.0% (medium) | 53.6% |
Source: OpenAI, as reported by VentureBeat and Vellum. Dashes mean no score was published.
On computer use, Sol only just edges Opus 5 (60.5% vs 60.3%) and sits 12 points behind Astra. On Agents' Last Exam, Sol gains 2.8 points over GPT-5.6 Sol and trails Astra by 2.9. For how these results fit the wider field, see our frontier AI benchmarks tracker.
Factuality and Deception
OpenAI says Sol makes "about half as many mistakes" as GPT-5.6 Sol on factuality, "reaching Astra-level reliability at much lower cost". OpenAI did not give an absolute error rate in the sources we checked, so treat this as a relative claim.
The deception numbers are sharper. On OpenAI's adversarial coding deception test, Sol deceived 1.3% of the time, down from 10.4% for its predecessor. Luna fell to 2.8% from 9.5%. For anyone running coding agents unattended, a model that claims tests pass when they do not is a real cost. An eightfold drop on Sol is a meaningful improvement, even on a vendor-run test.
Independent View: Artificial Analysis
Artificial Analysis runs its own evaluations. Its Intelligence Index (v4.3.x) scores GPT-6 Sol (max) at 48. That places it below GPT-6 Astra (53) and Claude Fable 5.1 (53), and above Grok 4.7 (46). Artificial Analysis has not yet published a score for Opus 5.5 in our sources, so we make no comparison there.
On speed, Artificial Analysis measures Sol at 131.2 output tokens per second, with a blended price of $1.54 per million tokens. It lists Sol's context window at about 872K tokens, short of the 1M tokens Astra and the current Claude models offer.
The five-point gap to Astra matches the pricing logic. Sol is not a frontier-topping model. It is a strong mid-tier model priced to take volume.
Availability
Both models went live on 22 September 2026 in ChatGPT Work, Codex and the API. Luna is also open to Free and Go users in the ChatGPT desktop app. On the same day, OpenAI loaded a banked rate-limit reset into Plus, Pro and Business accounts, a move that mirrored Anthropic's bankable reset launched alongside Opus 5.5.
Who Should Not Use Sol or Luna
Skip GPT-6 Sol if:
- You need the top score on business automation. On AutomationBench, Opus 5.5 (40.0%) and Astra (41.4%) are well ahead of Sol (33.2%). If a failed workflow costs more than the tokens, pay for the stronger model. Compare the two in GPT-6 Sol vs Claude Opus 5.5.
- You rely on computer use. Sol barely beats Opus 5 on OpenAI's own OSWorld harness and trails Astra by 12 points.
- You need a full 1M-token context. Artificial Analysis lists Sol at about 872K tokens. Claude Opus 5.5, Sonnet 5 and Astra offer 1M.
- You already run Claude Sonnet 5 happily. The price is identical. With no shared benchmark, switching only makes sense if Sol wins on your own tests.
Skip GPT-6 Luna if:
- Your work needs deep multi-step reasoning. OpenAI published Luna scores only for DeepSWE and OSWorld. There is no Agents' Last Exam or AutomationBench number to back it on harder agentic work.
- You are a Free or Go user on mobile or web. Luna access for those plans is in the desktop app.
- Token cost is not your bottleneck. If you process thousands, not millions, of items a month, the saving over Sol is small in absolute terms and Sol is the safer default.
The Bottom Line
GPT-6 Sol and Luna are price moves first. Sol at $2 / $10 and Luna at $0.10 / $0.50 cut OpenAI's mid and low tiers by half or more, for good. Sol beats Anthropic's previous generation on OpenAI's charts and does it cheaply. Luna comes within about three points of Fable 5 on DeepSWE at a fraction of the cost.
What Sol does not do is beat Opus 5.5 or Astra where scores can be compared. And it arrives priced exactly like Claude Sonnet 5, so the real contest is Sol against Sonnet 5 on your own workloads. Luna is the easy call for bulk work. For the Anthropic side of the ledger, see our Claude pricing guide and our Claude vs ChatGPT, Gemini and Grok comparison.
Sources
- OpenAI: Introducing GPT-6 Sol and Luna
- VentureBeat: OpenAI releases GPT-6 Sol and Luna, slashing API costs 50% or more
- Vellum: GPT-6 Sol and Luna benchmarks explained
- TechCrunch: OpenAI launches GPT-6 Sol and Luna
- Artificial Analysis: GPT-6 Sol
- OpenAI: GPT-6 Astra
- DataCamp: GPT-6 Astra
- Anthropic: Claude Opus 5.5
- Anthropic: Claude API pricing
Last updated: 23 September 2026. Benchmarks for GPT-6 Sol and Luna are OpenAI's own, as reported by VentureBeat and Vellum; Claude figures are Anthropic's; the Intelligence Index and speed figures are from Artificial Analysis. Prices are standard-tier list prices checked 23 September 2026.
Still deciding between ChatGPT and Claude?
Download our free 20-page guide comparing Claude, ChatGPT, Gemini and Grok on price, features and what each is good at, then test Sol or Sonnet 5 against your own work.






