Quick answer:
OpenAI previewed GPT-5.6 Sol Ultrafast on 13 August 2026. The limited API tier is powered by Cerebras and is claimed to run up to 14 times faster than Standard processing, reaching up to 750 output tokens per second. Pricing and wider access have not been published. Standard API prices remain Sol $5/$30, Terra $2/$12 and Luna $0.20/$1.20 per million input/output tokens.
Matt Wolfe's complete guide to GPT-5.6.
GPT-5.6 launched globally on 9 July 2026. OpenAI has since cut Terra and Luna API prices, changed which models power everyday ChatGPT conversations, and previewed a much faster way to serve Sol.
Ultrafast is the latest change, but it is not a new model checkpoint or a general release. It is a limited API service tier for selected customers. Developers still need to choose between capability, cost and latency, then test the option that fits their own workload.
Note: this analysis uses OpenAI's 13 August Ultrafast preview, 6 August ChatGPT announcement, launch and pricing updates, model documentation, system card and published evaluations. Ultrafast speed figures and the 62% and 68% factual-error reductions below come from OpenAI, not independent testing. API prices are billed in US dollars.
GPT-5.6 Sol Ultrafast: up to 750 output tokens per second
OpenAI previewed GPT-5.6 Sol Ultrafast on 13 August 2026. It is a new API service tier powered by Cerebras, not a new model. OpenAI says it can run Sol at up to 14 times the speed of Standard processing and generate up to 750 output tokens per second.
Those numbers need careful interpretation. Both are maximum vendor claims, and OpenAI has not published a full latency distribution, prompt-length breakdown or independent comparison. Actual performance will depend on the request, output length, capacity and application architecture.
Access is currently limited to selected preview customers. OpenAI has not published token prices or a date for wider availability, so teams cannot yet make a reliable public cost comparison with Standard or Sol Fast mode.
| Sol service option | Published speed claim | Availability | Published price |
|---|---|---|---|
| Standard | Baseline | Generally available | $5 input / $30 output per 1M tokens |
| Fast | Up to 2.5 times Standard | Generally available | Twice the Standard rate |
| Ultrafast preview | Up to 14 times Standard; up to 750 output tokens per second | Selected customers | Not published |
The clearest use cases are interactive workloads where waiting has a direct operational cost, such as incident response, time-sensitive research, customer support and live coding assistance. For batch jobs, ordinary agents and work dominated by tool or network delays, faster token generation may not improve the whole workflow enough to justify an unknown premium.
Treat Ultrafast as an evaluation candidate rather than a default recommendation. Measure end-to-end task time, output quality, throughput under load and total cost once pricing becomes available. OpenAI's official preview announcement is the source for the current speed and availability claims.
6 August ChatGPT update: Sol for paid users, Luna for free
OpenAI updated the ChatGPT experience for GPT-5.6 on 6 August. For Plus and Pro users, a revised version of Sol now handles both quick responses and deeper reasoning. A slider on web, mobile and desktop lets users choose how much thought the model applies, so changing effort should feel more like extending one answer than switching between models with different styles.
OpenAI says the revised Sol is more direct, uses less unnecessary formatting and is more willing to correct a faulty premise. In the company's internal evaluation of financial, medical and legal prompts that required factual detail, answers containing at least one factual error were 68% less common with the updated Sol than with GPT-5.5 Instant. Luna recorded a 62% reduction in the same test. OpenAI has not published enough detail here for those figures to be treated as independently verified benchmark results.
Free and Go users are moving to Luna as the default during the week of 6 August. OpenAI says unlimited text chats and a Think button will follow in the week beginning 10 August, subject to abuse protections. File uploads, image generation and other tools will still have limits, so “unlimited” applies to text chat rather than the whole product.
This revised Sol is limited to the ChatGPT chat experience. OpenAI explicitly says the Sol version used by ChatGPT Work and Codex is not changing as part of the release. Developers should therefore avoid reading the ChatGPT behavioural update as a new API model or a reason to rerun production evaluations without a documented API checkpoint change.
Sol Fast mode and the 30 July price cuts
OpenAI cut Terra's API prices by 20% and Luna's by 80% on 30 July 2026. Sol's standard rates did not change. The reductions also apply through AWS, according to OpenAI's announcement.
| Model | Launch input/output | Current input/output | Change |
|---|---|---|---|
| Sol | $5 / $30 | $5 / $30 | Unchanged |
| Terra | $2.50 / $15 | $2 / $12 | 20% lower |
| Luna | $1 / $6 | $0.20 / $1.20 | 80% lower |
OpenAI also introduced Sol Fast mode, which runs at up to 2.5 times the standard speed for twice the standard token price. It replaces Priority Processing, although existing API calls that use the priority service tier continue to work.
ChatGPT and Codex subscription prices did not change. OpenAI says Terra and Luna now consume fewer usage credits, but the underlying quota budgets remain the same.
GPT-5.6 Sol vs Terra vs Luna
The three-model structure is easier to understand when treated as a workload decision rather than a simple best-to-worst ranking.
| Model | Best for | API input/output | Main trade-off |
|---|---|---|---|
| GPT-5.6 Sol | Complex coding, research, design and professional work | $5 / $30 | Highest cost |
| GPT-5.6 Terra | Everyday production agents, analysis and coding | $2 / $12 | Less headroom on the hardest tasks |
| GPT-5.6 Luna | Extraction, classification and high-volume transformations | $0.20 / $1.20 | Lowest overall capability |
GPT-5.6 Sol
Sol is the flagship and the target of most of OpenAI's headline claims. It is the model to choose when a failed answer costs more than the extra tokens: a difficult codebase change, a detailed research synthesis, a complex financial model or a polished client presentation. OpenAI also positions Sol as its best coding model and its strongest option for computer use and frontend design.
Sol supports reasoning levels from none through max. Pro mode spends more model work on a single final answer when accuracy matters more than speed. In ChatGPT Work and Codex, ultra goes further by coordinating parallel agents for work that divides cleanly into separate streams. That can reduce elapsed time, but it increases token use and should be reserved for genuinely complex tasks.
GPT-5.6 Terra
Terra now costs 40% of Sol for both input and output, yet OpenAI describes its performance as competitive with GPT-5.5. In OpenAI's published coding results, Terra scores 63.4% on SWE-Bench Pro against Sol's 64.6%, and 87.4% on Terminal-Bench 2.1 against Sol's 88.8%. The gaps are small, but Luna's larger price cut means Terra is no longer the automatic value choice.
That makes Terra a sensible candidate for routine software changes, document analysis, structured research, internal agents and tool-heavy business workflows. Sol remains the safer choice for the hardest edge cases, but routing ordinary work to Terra could cut a large production bill without a dramatic quality drop.
GPT-5.6 Luna
Luna is not intended to win flagship benchmark charts. Its value is that it retains the 1.05 million-token context window and core GPT-5.6 features at one twenty-fifth of Sol's input and output price. It suits repeatable work such as extracting fields, classifying content, reformatting data, producing first-pass summaries and handling straightforward customer-service operations.
Luna still posts 62.7% on SWE-Bench Pro and 84.7% on Terminal-Bench 2.1 in OpenAI's evaluation, which is stronger than its price might suggest. However, its longer-context performance falls away more sharply than Sol or Terra in OpenAI's own testing. A large advertised context window tells you what fits into the request, not how reliably every detail will be used.
What's new in GPT-5.6
Programmatic tool calling
Instead of repeatedly returning to the model between every small tool action, GPT-5.6 can write lightweight JavaScript that coordinates eligible tools, processes intermediate results and passes data between calls in a hosted runtime. OpenAI says this reduces model round trips and token use on bounded workflows. It is particularly relevant to agents that need to query several systems, filter results and assemble one final answer.

Multi-agent work
Multi-agent support is available in beta through the Responses API. One GPT-5.6 instance can coordinate concurrent subagents and combine their findings. This is useful when the task has independent parts, such as comparing several markets or reviewing separate parts of a codebase. It is a poor fit for work where every step depends closely on the one before it.
Better prompt caching and persisted reasoning
Developers can now mark reusable prompt prefixes with explicit cache breakpoints, rather than relying only on automatic caching. OpenAI applies a 30-minute minimum cache life. Cache reads retain a 90% discount, whilst cache writes cost 1.25 times the normal input rate. Persisted reasoning can also carry eligible reasoning items across turns, improving multi-turn quality and cache efficiency when a workflow continues over several calls.
Stronger design and knowledge work
OpenAI places unusual emphasis on layout, visual hierarchy and design judgement in this release. GPT-5.6 can inspect rendered interfaces, revise them and produce more polished websites, presentations, documents and spreadsheets. The claim is broader than code generation: the model is meant to recognise when the output looks unfinished and continue refining it.
That matters for professional work because the last 10% often consumes most of the human editing time. A technically correct slide deck with weak hierarchy is not ready to present. A working frontend with poor spacing is not ready to ship. GPT-5.6 is designed to narrow that gap, although output quality will still depend heavily on good reference material and clear constraints.



GPT-5.6 benchmarks
The final general-availability numbers answer one major question left open during the preview: OpenAI has now published SWE-Bench Pro results for all three models. Sol reaches 64.6%, Terra 63.4% and Luna 62.7%, compared with GPT-5.5 at 59.4% in the same launch table.
| OpenAI evaluation | Sol | Terra | Luna | GPT-5.5 |
|---|---|---|---|---|
| SWE-Bench Pro | 64.6% | 63.4% | 62.7% | 59.4% |
| Terminal-Bench 2.1 | 88.8% | 87.4% | 84.7% | 85.6% |
| BrowseComp | 90.4% | 87.5% | 83.3% | 84.4% |
| OSWorld 2.0 | 62.6% | 50.2% | 45.6% | 47.5% |
These results make two points. First, Sol is the strongest all-round model in the family, especially for computer use and browsing. Second, Terra stays remarkably close on coding benchmarks. Luna sometimes falls behind GPT-5.5, but does so at a much lower token price.
Benchmark caveats still apply. Most figures above come from OpenAI's own launch evaluation, model prompts and harnesses can affect results, and a few percentage points do not guarantee better performance on your codebase or workflow. The right migration test uses representative tasks, measures total tokens and latency, and checks the final work rather than trusting model confidence.
For the wider market comparison, see our Claude vs ChatGPT, Gemini and Grok guide. Teams considering open weights should also compare the cost and control offered by GLM 5.2.
The Artificial Analysis verdict: independent numbers
Artificial Analysis published independent measurements for all three models on 9 July 2026. GPT-5.6 Sol (max) scores 59 on the Artificial Analysis Intelligence Index v4.1, one point below Claude Fable 5 (with fallback) at 60. Its measured cost was $1.04 per Intelligence Index task against Fable 5's $2.75. Terra (max) scores 55 and Luna (max) 51, with measured costs of $0.55 and $0.21 per task.
Those cost-per-task figures are historical measurements made with the launch prices. Artificial Analysis has not yet republished them using OpenAI's 30 July rates. The benchmark scores remain useful, and the lower token prices improve Terra and Luna's likely inference economics, but the table below should not be read as a fresh rerun.
On the new Artificial Analysis Coding Agent Index (which pairs each model with its own agentic harness across DeepSWE, Terminal-Bench v2 and SWE-Atlas-QnA) Sol (max) in Codex leads outright at 80 points, topping all three evaluations (tying Grok 4.5 in Grok Build on SWE-Atlas-QnA). It also does so at a per-task cost roughly 40% below Claude Fable 5 (max) and 10% below Claude Opus 4.8 (max) in Claude Code. Terra and Luna score 77 and 75, with ~60% and ~80% per-task cost reductions against Sol.
| Artificial Analysis metric | Sol (max) | Terra (max) | Luna (max) | Fable 5 |
|---|---|---|---|---|
| Intelligence Index v4.1 | 59 | 55 | 51 | 60 |
| Coding Agent Index | 80 | 77 | 75 | 77 |
| Cost per Intelligence Index task | $1.04 | $0.55 | $0.21 | $2.75 |
Intelligence Index v4.1: the full field
GPT-5.6 Sol, Terra and Luna (marked) against every scored frontier model. Sol sits one point below Claude Fable 5 at roughly a third of its cost per task.
Coding Agent Index: each model in its harness
Sol in Codex leads outright at 80, with Terra and Luna bracketing Claude Fable 5 and GPT-5.5.
The Pareto analysis is the strategically interesting part. Across reasoning-effort levels, every GPT-5.6 model pushes past GPT-5.5 on the intelligence-versus-cost frontier, and Luna and Sol are always on the frontier ahead of Terra. In practice, for any Terra effort level there is a Luna or Sol configuration that is either more intelligent at no extra cost or equally intelligent for less, which makes Terra a curious middle child: sensible as a drop-in default, rarely the optimal pick. Luna (max) also matches or exceeds the intelligence of GLM-5.2 (max) and Gemini 3.5 Flash at a lower cost per task.
On knowledge work, Sol (max) ranks second only to Fable 5 (max) on AA-Briefcase, Artificial Analysis's new benchmark of realistic project tasks, and records the highest Presentation Elo of any model, meaning its PowerPoint, Excel and document outputs are judged the most visually attractive. Fable 5 keeps the overall lead on substance: a 56% rubric score against Sol's 42%, and an Analytical Quality Elo of 1,764 versus 1,592. On GDPval-AA v2, the two are effectively level. And on AA-Omniscience, Sol posts only a minor gain over GPT-5.5, a small accuracy uplift coupled with a higher hallucination rate, an honest caveat worth carrying into production plans.
Two more details round out the picture. GPT-5.6 models are OpenAI's first with cache-write pricing: cache writes cost 1.25× the input-token rate (the 90% cache-read discount stays), matching Anthropic's approach and better reflecting what cached tokens actually cost to serve. And Sol (max) is unusually token-efficient, about 15,000 output tokens per Intelligence Index task versus GPT-5.5's 16,000, defining a new frontier of intelligence versus output tokens and using fewer tokens than Claude Opus 4.8 (max), GLM-5.2 (max) or Gemini 3.5 Flash (high) while scoring higher.
We track all of these figures (Intelligence Index, Coding Agent Index and cost per task for more than 20 frontier models) on our live Benchmarks page. (Source: Artificial Analysis, 9 July 2026.)
ChatGPT, Work and Codex access
GPT-5.6 access now differs more clearly between consumer chat, Work, Codex and the API.
- ChatGPT Free and Go: Luna is becoming the default during the week of 6 August. Unlimited text chats and a Think button are due in the week beginning 10 August, subject to abuse guardrails. Separate limits remain for files, images and other tools.
- ChatGPT Plus and Pro: the updated Sol powers quick and deeper responses. A slider controls reasoning effort across web, mobile and desktop.
- ChatGPT Work: Free and Go users can access Terra. Plus, Pro, Business and Enterprise users can choose Terra or Luna, with Sol available according to plan and workspace settings.
- Codex: Free and Go users can access Terra. Plus, Pro, Business and Enterprise users can choose Sol, Terra or Luna.
- OpenAI API: developers can call all three models. The
gpt-5.6alias points to Sol.
The 6 August update removes the earlier GPT-5.5 Instant distinction for Plus and Pro chat: OpenAI says the same updated Sol now powers both quick and deeper responses. It does not alter the version of Sol used by Work or Codex, and OpenAI did not announce an API model change.
API pricing and specifications
OpenAI bills the API in US dollars. These rates took effect on 30 July 2026.
| Model | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| Sol | $5 | $0.50 | $30 |
| Terra | $2 | $0.20 | $12 |
| Luna | $0.20 | $0.02 | $1.20 |
Cache writes cost 1.25 times the normal input rate. Sol Fast mode costs twice the standard Sol rate and can run up to 2.5 times faster.
All three models list the same headline technical limits: a 1,050,000-token context window, 128,000 maximum output tokens and a 16 February 2026 knowledge cut-off. They accept text and image input, but not audio or video input, and return text.
Requests containing more than 272,000 input tokens are charged at twice the normal input price and 1.5 times the output price for the whole request. That makes careless long-context use expensive. Retrieval, summarisation and prompt caching still matter even when the model can technically accept a million tokens.
Safety and limitations
OpenAI classifies all three GPT-5.6 models as High capability in cybersecurity and in biological and chemical risk under its Preparedness Framework. None reaches the High threshold for AI self-improvement. The company says it spent about 700,000 A100-equivalent GPU hours on automated black-box red teaming before general availability.
The system card also describes a tendency to go beyond the user's intent in some agentic coding evaluations. Absolute rates were low, but stronger autonomy makes approval boundaries more important. Give agents scoped permissions, require confirmation before irreversible actions and keep logs that somebody reviews.
- Vendor benchmarks need independent confirmation. OpenAI's results are useful, but they are not a substitute for testing your own workflow.
- A huge context window is not perfect memory. Luna's weaker long-context evaluation shows that fitting information into a prompt does not guarantee reliable retrieval.
- Sol can be unnecessary. Paying flagship rates for simple extraction or classification wastes money and can add latency without improving the result.
- Safeguards can affect benign work. OpenAI warns that stronger cyber and biology protections can create friction for legitimate requests.
- Rollout is gradual. Eligible users may not see every GPT-5.6 option immediately, and workspace administrators can restrict access.
Which GPT-5.6 model should you choose?
Choose Sol when the work is difficult, ambiguous or expensive to get wrong. It is the right starting point for major code changes, deep research, polished client deliverables, complex computer-use tasks and workflows that need maximum reasoning.
Choose Terra when you need more headroom than Luna but cannot justify Sol for every call. Its published coding scores are close to Sol's at 40% of the token price. If you currently use GPT-5.5, test Terra at the same reasoning level and one level lower, which is also OpenAI's migration advice.
Choose Luna for clear, repeatable and high-volume tasks. At 4% of Sol's token price, it now deserves to be the first model tested for extraction, labelling, reformatting, first-pass summaries and simple automated responses. Add a routing rule that escalates uncertain or high-value cases to Terra or Sol.
The best setup may use all three. Route straightforward volume to Luna, normal production work to Terra and difficult exceptions to Sol. Measure completed-task cost rather than price per token alone, because a cheaper model that needs repeated corrections can cost more overall.
The bottom line
GPT-5.6 now spans three sharply separated model prices and a growing set of service speeds. In ChatGPT, Sol gives paid users one model across quick and deeper responses, while Luna brings broader access to Free and Go users. In the API, Sol remains the quality ceiling, and Ultrafast adds a limited option for work where response time is unusually valuable.
Luna's 80% cut still changes the default recommendation. Start bounded, high-volume tasks with Luna, move uncertain or more demanding work to Terra, and reserve Sol for cases where quality justifies the premium. Consider Ultrafast only when end-to-end latency matters enough to outweigh its unknown price and limited availability.
OpenAI's 13 August Ultrafast preview, 6 August ChatGPT update, 30 July pricing announcement, serving-efficiency report, launch announcement, model guidance and system card provide the underlying access, prices, specifications and evaluation detail.
Last updated: 15 August 2026. This article now reflects OpenAI's 13 August Ultrafast preview, 6 August ChatGPT changes, the 30 July Terra and Luna price cuts and Sol Fast mode. Artificial Analysis cost-per-task figures remain labelled as historical 9 July measurements.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






