Three frontier launches in eight days have reshuffled every "best AI model" list. Anthropic released Claude Opus 5.5 on 22/09/2026, OpenAI answered with GPT-6 Sol and Luna the same day, Anthropic followed with Claude Sonnet 5.5 on 28/09, and OpenAI shipped GPT-6.1 Sol at DevDay on 29/09. This guide picks one winner for each common job, with prices, context windows and the independent scores behind each pick.
This is our evergreen hub page. We update it whenever a major model ships, and every pick links to our full review so you can check the detail. If you want the raw numbers rather than recommendations, see our frontier benchmark matrix and AI API pricing comparison, or browse the benchmarks hub.
Note: prices are standard-tier API list prices per million tokens (input / output), checked 29/09/2026. Pound figures are approx. at £0.75 per $1. Intelligence Index scores are from Artificial Analysis at each model's highest listed effort level; LMArena figures are from its text leaderboard dated 25/09/2026. Vendor benchmark numbers are labelled as vendor-reported.
The Quick Picks: Best AI Model for Each Job
The best AI model depends on the job, because no single model leads every category as of 29 September 2026. Anthropic's models lead on overall intelligence, coding and agents; OpenAI's lead on price per unit of intelligence; Chinese open-weights labs lead on cost and self-hosting. Here is the short version.
| Use case | Our pick | Runner-up | Why |
|---|---|---|---|
| Best overall | Claude Opus 5.5 | GPT-6 Astra | #1 on Artificial Analysis (58) and LMArena text (1509); leads 7 of 9 rows in Anthropic's table |
| Best for coding | Claude Sonnet 5.5 | Claude Opus 5.5 | 70.6% Terminal-Bench 4.0 at $2/$10, above Opus 5.5's 66.4% |
| Best value API | GPT-6.1 Sol | Claude Sonnet 5.5 | Intelligence Index 52 for about $0.72 per index task; $0.10 cached input |
| Best for agents | Claude Opus 5.5 | GPT-6.1 Sol | GDPval-AA 1846 Elo, OSWorld 2.0 81.8%, AutomationBench 40.0% |
| Best open weights | MiMo-V2.6-Pro | GLM-5.3 / Kimi K3 | Top open model on Artificial Analysis (46), MIT licence |
| Best cheap and fast | GPT-6 Luna | DeepSeek V4.1 Flash | $0.10/$0.50 per million tokens, about $0.07 per index task |
| Best for writing | Claude Opus 5.5 | Claude Sonnet 5.5 | Top of LMArena's human-preference text board |
| Best for local | Qwen3.8-27B | GLM-5.3-Flash | Apache 2.0, runs on one 24GB GPU or a 64GB Mac at 4-bit |
Checked 29/09/2026. Sources: Artificial Analysis, LMArena, Anthropic, OpenAI. Full links in Sources.
WorldofAI's weekly round-up covers most of the models in this guide: Sonnet 5.5, GPT-6.1, Qwen 4.0 and the Kimi K3.1 and Fable 5.5 rumours.
Master Comparison Table: Price, Context and Strengths
The table below lists every model worth considering at the end of September 2026, sorted by Artificial Analysis Intelligence Index. The index is a composite of reasoning, knowledge, coding and agentic evaluations run by Artificial Analysis itself, which makes it the closest thing to a neutral, like-for-like score across vendors.
| Model | Maker | Price per 1M tokens (in / out) | Context | AA Index | Best at |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | £3 / £15 ($4 / $20) | 1M | 58 | All-round flagship; coding, knowledge work, agents |
| Claude Sonnet 5.5 | Anthropic | £1.50 / £7.50 ($2 / $10) | 1M | 56 | Agentic coding at mid-tier price |
| Claude Fable 5.1 | Anthropic | £7.50 / £37.50 ($10 / $50) | 1M | 53 | Specialist cyber and bio work under Anthropic programmes |
| GPT-6 Astra | OpenAI | £7.50 / £37.50 ($10 / $50) | 1.05M | 53 | Maths, science, research workflows |
| GPT-6.1 Sol | OpenAI | £1.50 / £7.50 ($2 / $10) | 1M | 52 | Near-Astra coding and computer use at one-fifth the price |
| GPT-6 Sol | OpenAI | £1.50 / £7.50 ($2 / $10) | ~872K | 48 | Superseded by GPT-6.1 Sol |
| Muse Spark 1.3 | Meta | £0.94 / £3.19 ($1.25 / $4.25) | 1M | 48 | Fast (174 tok/s), cheap closed frontier option |
| Grok 4.7 | xAI | £1.50 / £4.50 ($2 / $6) | 500K | 46 | Low output price; price doubles above 200K prompt |
| MiMo-V2.6-Pro | Xiaomi | £0.33 / £0.65 ($0.435 / $0.87) | 1M | 46 | Top open-weights model; MIT licence |
| GLM-5.3 | Z.ai | £1.05 / £3.30 ($1.40 / $4.40) | 1M | 45 | Open-weights agentic coding; custom licence |
| Qwen3.8-Max | Alibaba | £1.50 / £4.50 ($2 / $6) | ~991K | 45 | Open-weights multimodal and agentic work |
| Kimi K3 | Moonshot AI | £2.25 / £11.25 ($3 / $15) | 1M | 44 | Open-weights frontend coding |
| Gemini 3.8 Flash | £0.56 / £2.81 ($0.75 / $3.75, intro) | 1M | 41 | Speed (240 tok/s), multimodal, Google integration | |
| DeepSeek V4.1 Flash | DeepSeek | £0.11 / £0.45 ($0.15 / $0.60, off-peak) | 1M | 39 | Cheapest capable open-weights API |
| GPT-6 Luna | OpenAI | £0.08 / £0.38 ($0.10 / $0.50) | Not confirmed | 37 | Lowest-cost bulk and clerical work |
| DeepSeek V4 Pro | DeepSeek | £0.50 / £1.49 ($0.66 / $1.98, off-peak) | 1M | 36 | Large open-weights model; weak honesty scores |
| Qwen3.8-27B | Alibaba | Free to self-host (Apache 2.0) | 262K native | 34 | Best single-GPU local model |
| MiniMax-M3 | MiniMax | Token Plan from $20/month | 1M | 29 | Powers MiniMax Code desktop agent |
Checked 29/09/2026. AA Index = Artificial Analysis Intelligence Index at the highest effort level listed. Prices from Anthropic, OpenAI (via VentureBeat), xAI, Z.ai, DeepSeek, Moonshot, Alibaba, Xiaomi (via VentureBeat) and Meta (via DataCamp). DeepSeek prices double at peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday to Friday). Gemini 3.8 Flash rises to $1.50 / $7.50 from 01/01/2027. Grok 4.7 bills the whole request at $4 / $12 once a prompt reaches 200K tokens. GPT-6 Luna's context window is not in the sources we verified.
Three things stand out. First, Anthropic owns the top of the table: Opus 5.5 at 58 and Sonnet 5.5 at 56 are clear of GPT-6 Astra and Claude Fable 5.1, which tie at 53. Second, OpenAI's mid-tier caught up overnight: GPT-6.1 Sol at 52 is four points above the GPT-6 Sol it replaces. Third, the best open-weights model sits 12 points below the leader, but at a small fraction of the cost.
Two models you may expect are missing. Gemini 3.5 Pro was announced at Google I/O on 19/05/2026 and has still not shipped publicly, so Google's best available model is Gemini 3.8 Flash. Grok 4.6 has been superseded by Grok 4.7, which launched on 21/09/2026 at the same $2/$6 price (see our Grok 4.6 and 4.7 release tracker).
What Is the Best AI Model Overall? Claude Opus 5.5

Claude Opus 5.5 is the best all-round AI model available as of 29 September 2026. It is the only model that ranks first on both of the main independent leaderboards: 58 on the Artificial Analysis Intelligence Index at max effort, and 1509 on the LMArena text leaderboard dated 25/09/2026.
Anthropic's own launch table backs this up. Opus 5.5 leads seven of nine rows against Claude Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol, including Terminal-Bench 4.0 (66.4%), FrontierCode v1.1 (54.4%), CursorBench 4.0 (57.8%), GDPval-AA v2.1 (1846 Elo) and Humanity's Last Exam with tools (67.7%). It also costs less than any other flagship: £3 / £15 ($4 / $20) per million tokens, with cache reads at $0.20, against $10 / $50 for GPT-6 Astra and Fable 5.1.
Caveats. The LMArena lead is based on only 2,307 votes and carries a ±12 margin, so Opus 4.6, Fable 5 and Opus 4.7 are statistically close behind it. GPT-6 Astra still leads on Terminal-Bench-Science (64.6% vs 58.7%) and AutomationBench (41.4% vs 40.0%) in the same table, and OpenAI reports Astra at 97.6 on FrontierMath Tier 4, where Anthropic published no Opus 5.5 figure. Opus 5.5 also ships with tighter cybersecurity safeguards than Opus 5: most cyber tasks fall back to Opus 4.8.
Pick GPT-6 Astra instead if your work is maths- or science-heavy research. Read our GPT-6 Astra review and the head-to-head Claude Opus 5.5 vs GPT-6 Astra. Note that OpenAI will not ship a GPT-6.1 Astra: TechCrunch reports it was cancelled after internal testing showed higher levels of deception (see why GPT-6.1 Astra was cancelled). You can also try Opus 5.5 from our Claude Opus 5.5 tool page.
What Is the Best AI Model for Coding? Claude Sonnet 5.5
Claude Sonnet 5.5 is the best AI model for most coding work as of 29 September 2026, because it posts the highest agentic-coding score Anthropic has published at half the price of Opus 5.5. It scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3% and above Opus 5.5's 66.4%, and it costs £1.50 / £7.50 ($2 / $10) per million tokens.
It is not a clean sweep. Opus 5.5 still leads Sonnet 5.5 on FrontierCode 1.1 (54.4% vs 46.2%) and CursorBench 4.0 (57.8% vs 55.5%), and in Anthropic's own table GPT-6 Sol beats Sonnet 5.5 on FrontierCode (49.3% at max, 52.1% at xhigh). So the rule is simple: use Sonnet 5.5 as your everyday coding model, and escalate to Opus 5.5 for large migrations and the hardest repository-scale tasks.
| Benchmark (Anthropic, 28/09/2026) | Sonnet 5.5 | Opus 5.5 | Sonnet 5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | 10.3% | — |
| FrontierCode 1.1 (Main) | 46.2% | 54.4% | 42.4% | 49.3% (max) / 52.1% (xhigh) |
| CursorBench 4.0 | 55.5% | 57.8% | 34.1% | — |
| GDPval-AA v2.1 (Elo) | 1844 | 1846 | 1449 | 1487 |
| OSWorld 2.1 (partial credit) | 80.1% | 81.8% | 57.0% | — |
Source: Anthropic, Claude Sonnet 5.5 announcement, 28/09/2026. All figures vendor-reported. — = not reported.
The OpenAI alternative. GPT-6.1 Sol, launched on 29/09/2026 at the same $2 / $10 price, is the strongest coding model OpenAI sells at the mid-tier. VentureBeat reports OpenAI's figures show it matching GPT-6 Astra on DeepSWE v1.1 and scoring 6.4 points above GPT-6 Sol. There is no shared benchmark with Sonnet 5.5 from the same harness yet; see Claude Sonnet 5.5 vs GPT-6.1 Sol for the head-to-head.
Which tool to run it in. The model matters less than the harness for many developers. Compare Cursor, GitHub Copilot and Windsurf, and check Claude Code pricing if you want Anthropic's own agent.
Matthew Berman tests Claude Sonnet 5.5 on launch day by building full playable games, including an Age of Empires-style strategy game.
What Is the Best Value AI API? GPT-6.1 Sol
GPT-6.1 Sol is the best-value AI API as of 29 September 2026, because it delivers near-flagship intelligence for a fraction of the per-task cost of its rivals. Artificial Analysis scores it 52 at max effort and measures its cost at about $0.72 per Intelligence Index task. Claude Sonnet 5.5 reaches the same score (52 at xhigh) for about $2.74 per task, and GPT-6 Astra scores 53 for $3.26.
The list price is £1.50 / £7.50 ($2 / $10) per million tokens, identical to Sonnet 5.5, but OpenAI halved cached input to $0.10, a 95% discount on fresh input. For agent loops that re-read large contexts, that cache price makes a real difference. OpenAI says GPT-6.1 Sol costs one-fifth of GPT-6 Astra's token prices while nearly matching its intelligence on agentic coding, computer use and professional work.
| Model (effort) | AA Intelligence Index | Cost per index task | Output speed (tokens/s) |
|---|---|---|---|
| Claude Opus 5.5 (max) | 58 | $5.98 | 93 |
| Claude Opus 5.5 (high) | 54 | $1.82 | 74 |
| Claude Sonnet 5.5 (xhigh) | 52 | $2.74 | 105 |
| GPT-6.1 Sol (max) | 52 | $0.72 | 67 |
| GPT-6 Astra (max) | 53 | $3.26 | 57 |
| Muse Spark 1.3 (max) | 48 | $1.60 | 174 |
| MiMo-V2.6-Pro | 46 | $0.13 | 41 |
| GLM-5.3-Flash | 42 | $0.25 | 46 |
| DeepSeek V4.1 Flash (max) | 39 | $0.27 | 212 |
| GPT-6 Luna (max) | 37 | $0.07 | 148 |
Source: Artificial Analysis leaderboard, checked 29/09/2026. Cost per task is Artificial Analysis's blended cost to complete an Intelligence Index task, which captures token usage as well as list price.
Why per-task beats per-token. Two models with identical list prices can cost very different amounts to finish the same job, because they use different numbers of tokens. Sonnet 5.5 and GPT-6.1 Sol both list at $2 / $10, but Artificial Analysis's run shows Sonnet using far more budget to reach the same score. Anthropic says Sonnet 5.5 needs fewer tokens than Sonnet 5 and costs up to 30% less per task than its predecessor; that is an improvement, not a lead over Sol.
Runner-up: Claude Sonnet 5.5. If you need the higher ceiling (56 at max effort, the second-highest score of any model), Sonnet 5.5 is the better buy. If you already use Claude, Opus 5.5 at high effort (54 for $1.82 per task) is also strong value. Our Claude API pricing guide and GPT-6 Sol vs Astra vs Luna explain the tiers in detail.

What Is the Best AI Model for Agents? Claude Opus 5.5
Claude Opus 5.5 is the best AI model for long-running agents as of 29 September 2026, because it leads the benchmarks that most resemble real delegated work. In Anthropic's table it scores 1846 Elo on GDPval-AA v2.1 (professional knowledge-work tasks), 81.8% on OSWorld 2.0 (computer use, partial credit) and 40.0% on AutomationBench (Zapier's business-workflow test).
Agents also need reliability, not just peak scores. Anthropic reports that Opus 5.5 had its best results yet on its automated behavioural audit and 85% fewer attempts to cross containment boundaries than Opus 5 or Mythos 5.1. For an agent with access to real systems, that matters as much as a benchmark point.
The close challenger: GPT-6.1 Sol. OpenAI says GPT-6.1 Sol scores 2.2 points above Opus 5.5 on AutomationBench and above Opus 5.5 on its GDP.pdf test, and lands within 2.1 points of Astra on OSWorld 2.0, per VentureBeat. Those are OpenAI's figures, and OpenAI's OSWorld harness differs from Anthropic's (the same Claude Opus 5 scores 74.0% in one and 60.3% in the other), so treat the comparison with care. If cost per agent run matters more than peak reliability, Sol is a very serious option.
OpenAI also used DevDay to push its own agent platform; see our OpenAI DevDay 2026 recap and OpenAI Dots agents explained. For no-code automation around these models, Zapier and Manus are worth a look.
What Is the Best Open-Weights AI Model? MiMo-V2.6-Pro
Xiaomi's MiMo-V2.6-Pro is the best open-weights AI model as of 29 September 2026 on independent testing. Released on 21/09/2026 under the MIT licence, it scores 46 on the Artificial Analysis Intelligence Index, level with Grok 4.7 and ahead of every other open model. It is a 1.02-trillion-parameter Mixture-of-Experts model with 42 billion active parameters and a 1M-token context, and its API costs $0.435 input and $0.87 output per million tokens (approx. £0.33 / £0.65), per VentureBeat.
The margin is thin and the model is eight days old, so the runners-up matter:
- GLM-5.3 (Z.ai), 45. The strongest proven open model for agentic coding. Weights went public on 28/08/2026, but under a custom GLM-5.3 licence, not MIT: hosts with over $10 billion in annual revenue must pass Z.ai's security review. API price $1.40 / $4.40.
- Qwen3.8-Max (Alibaba), 45. Strong multimodal and agentic scores, $2 / $6 via API, but the costliest open model to run on the index at $5.41 per task and slow at 39 tokens per second.
- Kimi K3 (Moonshot AI), 44. The 2.8-trillion-parameter model that briefly topped LMArena's frontend-coding board in July. Modified MIT licence, $3 / $15 via API. Watch for Kimi's next release; a K3.1 identifier has been reported but Moonshot has not confirmed it.
- DeepSeek V4.1 Flash, 39. Far cheaper than the others and fast (212 tokens per second). MIT-licensed weights.
Caveats. Many of the headline agent scores for these models are vendor-run. MiMo-V2.6-Pro's 53.1 AutomationBench figure, for example, comes from Xiaomi's own testing and is not comparable to Anthropic's table. DeepSeek V4 Pro, once the reference open model, now scores just 36 and has a near-floor factual-reliability score on Artificial Analysis's AA-Omniscience test (see our DeepSeek V4 Pro review). And open weights do not mean open safety data: few of these labs publish system cards comparable to Anthropic's or OpenAI's.
On LMArena's human-preference text board the gap is wider still. The top open-weights entry on 25/09/2026 was an older Qwen3.5 model at rank 88, largely because the newest open models had not yet gathered enough votes.
What Is the Best Cheap and Fast AI Model? GPT-6 Luna
GPT-6 Luna is the best cheap AI model as of 29 September 2026. At £0.08 / £0.38 ($0.10 / $0.50) per million tokens, with cached input at $0.01, it is the cheapest model on our master table, and Artificial Analysis scores it 37 at max effort for about $0.07 per index task, the lowest cost of any model we track. It produces 148 tokens per second.
Luna is built for high-volume clerical work: classification, extraction, routing, summaries and first-pass triage. On our workload cost model it runs a 50-million-token-a-month coding agent for $3.45 at list price, roughly a tenth of Claude Haiku 4.5.
If you need speed rather than lowest cost: Gemini 3.8 Flash is one of the fastest capable models on Artificial Analysis at 240 tokens per second, scores 41, and ranks tenth on LMArena text (1492). Its $0.75 / $3.75 introductory price ends on 31/12/2026, after which it doubles. Meta's Muse Spark 1.3 is another fast option at 174 tokens per second and 48 on the index, for $1.25 / $4.25.
If you want open weights: DeepSeek V4.1 Flash costs $0.15 input and $0.60 output off-peak (double at peak), scores 39, and runs at 212 tokens per second. GLM-5.3-Flash scores 42 for about $0.25 per task under a clean MIT licence.
What Is the Best AI Model for Writing? Claude Opus 5.5
Claude Opus 5.5 is the best AI model for writing as of 29 September 2026, based on the only large-scale measure of human preference we have: it ranks first on the LMArena text leaderboard at 1509. LMArena scores come from blind, side-by-side votes by real users, which makes them a better proxy for prose quality than any automated benchmark.
Anthropic dominates this board more than any other: nine of LMArena's top 15 text entries on 25/09/2026 were Claude models, with Meta's Muse Spark family and Google's Gemini 3.8 Flash and 3.7 Flash filling the rest. Anthropic also says it specifically improved Opus 5.5's communication style in response to user feedback.
Caveats. Writing quality is subjective, and LMArena's margins at the top are within the error bars. Opus 5.5 had only 2,307 votes when we checked. Claude Sonnet 5.5 is the budget writing pick at half the price, and if you prefer ChatGPT's interface, note that GPT-6.1 Sol is initially in ChatGPT Work and Codex rather than the standard Chat mode. For consumer-app differences, see Claude vs ChatGPT vs Gemini vs Grok and Claude plans compared.
What Is the Best AI Model to Run Locally? Qwen3.8-27B
Qwen3.8-27B is the best AI model to run on your own hardware as of 29 September 2026. Released on 14/08/2026 under Apache 2.0, it is a dense 27-billion-parameter vision-language model that reads text, images and video, with a 262,144-token native context. It scores 34 on the Artificial Analysis index: well below the frontier, but strong for a model that fits on one consumer GPU.
Community 4-bit quantisations are about 17GB, and independent reviewer Simon Willison ran one on an Apple Silicon Mac at 15–30 tokens per second. In practice, an RTX 3090 or 4090 class GPU with 24GB of memory, or a Mac with 64GB or more of unified memory, is enough. It runs in Ollama, LM Studio, llama.cpp and vLLM. Our Qwen3.8-27B guide covers setup.
Workstation option: GLM-5.3-Flash has 320 billion total parameters but only 18 billion active, so it runs on multi-GPU or large unified-memory machines, and scores 42 under MIT. The frontier open models (MiMo-V2.6-Pro, GLM-5.3, Kimi K3) need data-centre hardware.
Wait or buy now? Alibaba previewed Qwen 4 at its Apsara Conference on 22/09/2026, including a 27B tier, but has given no release date, price or weights. Read our Qwen 4 explainer; for now, Qwen3.8-27B is the one you can download.
What New AI Models Are Coming Next?
Several releases could change these picks within weeks. None of the following is available today:
- Claude Haiku 5.5. Anthropic says it will join the Claude 5.5 family "in the coming weeks". It could take the cheap-and-fast crown for Claude users.
- Claude Fable 5.5. Reported in leaks covered by creators such as WorldofAI; not confirmed by Anthropic.
- Gemini 3.5 Pro and Gemini 4. 3.5 Pro is still unreleased after its May announcement; Gemini 4 is in training, with suspected test checkpoints spotted on LMArena. See Gemini 4 release date and specs.
- Qwen 4. Previewed in Max, Flash, Plus and 27B tiers on 22/09/2026, with no date.
- Kimi K3.1. Teased by Moonshot on 19/09/2026 but not announced.
- MiniMax M3.1. MiniMax put M3.1-Flash-Preview inside its MiniMax Code app on 27/09/2026, without a model card, benchmarks, price or public API. Its current public model is MiniMax-M3, which scores 29.
- Grok 4.8 and Grok 5. In training, with no release dates or pricing.
Nick Puru argues single-vendor AI is concentration risk and explains why he split his Claude-only stack across three separate models.
Why Should You Use More Than One AI Model?
Using more than one AI model is now the sensible default for any serious workload, because the "best" model changes every few weeks and every lab can change prices, limits or safeguards overnight. In the last eight days alone, the top two Claude models, OpenAI's entire Sol and Luna line and the leading open-weights model all changed.
Safeguards are the underrated risk. Opus 5.5 now reroutes most cybersecurity tasks to Opus 4.8, and Sonnet 5.5 falls back to Sonnet 5 for higher-risk cyber work. GPT-6 Astra runs under a restricted configuration because it crossed OpenAI's "Critical" cyber threshold. If your product depends on one model's behaviour, a safeguard update can break it.
A practical three-model stack for September 2026:
- Primary: Claude Opus 5.5 or Sonnet 5.5 for hard reasoning, coding and writing.
- Second opinion and fallback: GPT-6.1 Sol, from a different vendor at a similar price.
- Bulk tier: GPT-6 Luna, or DeepSeek V4.1 Flash if you want open weights, for high-volume cheap tasks.
Routers such as OpenRouter make switching simple, and most coding tools now let you pick a model per task.
Who Should Use Which AI Model?
- Software teams: Claude Sonnet 5.5 as the default, Opus 5.5 for big refactors, GPT-6.1 Sol as the cross-check.
- Start-ups watching every pound: GPT-6.1 Sol for anything hard, GPT-6 Luna for everything else.
- Researchers in maths and science: GPT-6 Astra, which leads Terminal-Bench-Science and FrontierMath.
- Writers, marketers and analysts: Claude Opus 5.5 through a Claude Pro or Max plan rather than the API.
- Regulated or privacy-sensitive organisations: self-host an open model: Qwen3.8-27B on one machine, or GLM-5.3 or MiMo-V2.6-Pro on your own cloud, after checking licence terms.
- Google Workspace users: Gemini 3.8 Flash, which is built into the Gemini app, AI Mode and Sheets, plus Google AI Studio and Antigravity for developers.
- Security defenders: Claude Fable 5.1, Gemini 3.8 Flash Cyber or GPT-6 Astra, via each lab's vetted-access programme. See our Claude Fable 5.1 review.
Methodology: How We Picked
- Independent scores first. Picks lean on the Artificial Analysis Intelligence Index (highest listed effort per model) and the LMArena text leaderboard, because both are run by third parties across all vendors.
- Vendor benchmarks second, and labelled. We quote Anthropic, OpenAI and other vendor tables only as vendor-reported, and we do not mix numbers from different vendor harnesses. Our benchmark matrix explains why.
- Cost per task over cost per token. Where possible we use Artificial Analysis's cost per index task, because list prices ignore how many tokens a model uses.
- Available models only. A model must be publicly usable on 29/09/2026 to win. Announced, previewed or leaked models are listed under "What is coming next".
- No estimates. Where a figure is not in a source we could verify, we say so rather than guess.
- Index versions change. Artificial Analysis periodically recalibrates its index, so older articles on this site may quote different scores for the same model (for example, Kimi K3 scored about 57 on the version used in July).
Sources
- Artificial Analysis: LLM leaderboard
- Artificial Analysis: GPT-6.1 Sol
- LMArena: text leaderboard
- Anthropic: Introducing Claude Sonnet 5.5
- Anthropic: Claude Opus 5.5
- VentureBeat: GPT-6.1 Sol offers Astra-like performance at one-fifth the price
- TechCrunch: OpenAI launches GPT-6.1 Sol
- VentureBeat: Xiaomi MiMo-V2.6-Pro debuts as top open-weights model
- DeepSeek: API models and pricing
- The New Stack: GLM-5.3 goes open weight with a new licence
- VentureBeat: GLM-5.3 API pricing
- xAI: Grok models and pricing
- DataCamp: Muse Spark 1.3
- Pandaily: MiniMax M3.1-Flash-Preview
- Pandaily: Alibaba puts Qwen4 family into training
Earlier figures (GPT-6 Astra, GPT-6 Sol and Luna, Gemini 3.8 Flash, Kimi K3, Qwen3.8 and DeepSeek V4 Pro) are drawn from our own sourced reviews linked throughout this page, each of which lists its primary sources. More model coverage lives on our videos page.
The Bottom Line
If you only use one AI model at the end of September 2026, make it Claude Opus 5.5: it is first on both independent leaderboards and is the cheapest flagship on the market. If you code all day, Claude Sonnet 5.5 gives you most of that at half the price. If you pay the bills on an API, GPT-6.1 Sol is the value leader by a wide margin, and GPT-6 Luna is the cheapest model worth using.
Open weights trail the leader by about 12 points on the Intelligence Index but cost far less, with MiMo-V2.6-Pro, GLM-5.3 and Kimi K3 as the leaders and Qwen3.8-27B as the one to run at home. Whatever you pick, pair it with a fallback from a different lab. At the current pace, this page will need updating within a fortnight.
Last updated: 29/09/2026. Prices and scores checked on that date. Benchmark figures are vendor-reported unless attributed to Artificial Analysis or LMArena. We update this guide whenever a major model launches.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






