AI Tools Review
Best AI Models Right Now: Which LLM to Use

Buyer's Guide

Best AI Models Right Now: Which LLM to Use

AI Tools Review Editorial Team29 September 2026
  • Best AI Models
  • LLM Comparison
  • Claude Opus 5.5
  • Claude Sonnet 5.5

Three frontier launches in eight days have reshuffled every "best AI model" list. Anthropic released Claude Opus 5.5 on 22/09/2026, OpenAI answered with GPT-6 Sol and Luna the same day, Anthropic followed with Claude Sonnet 5.5 on 28/09, and OpenAI shipped GPT-6.1 Sol at DevDay on 29/09. This guide picks one winner for each common job, with prices, context windows and the independent scores behind each pick.

This is our evergreen hub page. We update it whenever a major model ships, and every pick links to our full review so you can check the detail. If you want the raw numbers rather than recommendations, see our frontier benchmark matrix and AI API pricing comparison, or browse the benchmarks hub.

Note: prices are standard-tier API list prices per million tokens (input / output), checked 29/09/2026. Pound figures are approx. at £0.75 per $1. Intelligence Index scores are from Artificial Analysis at each model's highest listed effort level; LMArena figures are from its text leaderboard dated 25/09/2026. Vendor benchmark numbers are labelled as vendor-reported.

The Quick Picks: Best AI Model for Each Job

The best AI model depends on the job, because no single model leads every category as of 29 September 2026. Anthropic's models lead on overall intelligence, coding and agents; OpenAI's lead on price per unit of intelligence; Chinese open-weights labs lead on cost and self-hosting. Here is the short version.

Use caseOur pickRunner-upWhy
Best overallClaude Opus 5.5GPT-6 Astra#1 on Artificial Analysis (58) and LMArena text (1509); leads 7 of 9 rows in Anthropic's table
Best for codingClaude Sonnet 5.5Claude Opus 5.570.6% Terminal-Bench 4.0 at $2/$10, above Opus 5.5's 66.4%
Best value APIGPT-6.1 SolClaude Sonnet 5.5Intelligence Index 52 for about $0.72 per index task; $0.10 cached input
Best for agentsClaude Opus 5.5GPT-6.1 SolGDPval-AA 1846 Elo, OSWorld 2.0 81.8%, AutomationBench 40.0%
Best open weightsMiMo-V2.6-ProGLM-5.3 / Kimi K3Top open model on Artificial Analysis (46), MIT licence
Best cheap and fastGPT-6 LunaDeepSeek V4.1 Flash$0.10/$0.50 per million tokens, about $0.07 per index task
Best for writingClaude Opus 5.5Claude Sonnet 5.5Top of LMArena's human-preference text board
Best for localQwen3.8-27BGLM-5.3-FlashApache 2.0, runs on one 24GB GPU or a 64GB Mac at 4-bit

Checked 29/09/2026. Sources: Artificial Analysis, LMArena, Anthropic, OpenAI. Full links in Sources.

WorldofAI's weekly round-up covers most of the models in this guide: Sonnet 5.5, GPT-6.1, Qwen 4.0 and the Kimi K3.1 and Fable 5.5 rumours.

Master Comparison Table: Price, Context and Strengths

The table below lists every model worth considering at the end of September 2026, sorted by Artificial Analysis Intelligence Index. The index is a composite of reasoning, knowledge, coding and agentic evaluations run by Artificial Analysis itself, which makes it the closest thing to a neutral, like-for-like score across vendors.

ModelMakerPrice per 1M tokens (in / out)ContextAA IndexBest at
Claude Opus 5.5Anthropic£3 / £15 ($4 / $20)1M58All-round flagship; coding, knowledge work, agents
Claude Sonnet 5.5Anthropic£1.50 / £7.50 ($2 / $10)1M56Agentic coding at mid-tier price
Claude Fable 5.1Anthropic£7.50 / £37.50 ($10 / $50)1M53Specialist cyber and bio work under Anthropic programmes
GPT-6 AstraOpenAI£7.50 / £37.50 ($10 / $50)1.05M53Maths, science, research workflows
GPT-6.1 SolOpenAI£1.50 / £7.50 ($2 / $10)1M52Near-Astra coding and computer use at one-fifth the price
GPT-6 SolOpenAI£1.50 / £7.50 ($2 / $10)~872K48Superseded by GPT-6.1 Sol
Muse Spark 1.3Meta£0.94 / £3.19 ($1.25 / $4.25)1M48Fast (174 tok/s), cheap closed frontier option
Grok 4.7xAI£1.50 / £4.50 ($2 / $6)500K46Low output price; price doubles above 200K prompt
MiMo-V2.6-ProXiaomi£0.33 / £0.65 ($0.435 / $0.87)1M46Top open-weights model; MIT licence
GLM-5.3Z.ai£1.05 / £3.30 ($1.40 / $4.40)1M45Open-weights agentic coding; custom licence
Qwen3.8-MaxAlibaba£1.50 / £4.50 ($2 / $6)~991K45Open-weights multimodal and agentic work
Kimi K3Moonshot AI£2.25 / £11.25 ($3 / $15)1M44Open-weights frontend coding
Gemini 3.8 FlashGoogle£0.56 / £2.81 ($0.75 / $3.75, intro)1M41Speed (240 tok/s), multimodal, Google integration
DeepSeek V4.1 FlashDeepSeek£0.11 / £0.45 ($0.15 / $0.60, off-peak)1M39Cheapest capable open-weights API
GPT-6 LunaOpenAI£0.08 / £0.38 ($0.10 / $0.50)Not confirmed37Lowest-cost bulk and clerical work
DeepSeek V4 ProDeepSeek£0.50 / £1.49 ($0.66 / $1.98, off-peak)1M36Large open-weights model; weak honesty scores
Qwen3.8-27BAlibabaFree to self-host (Apache 2.0)262K native34Best single-GPU local model
MiniMax-M3MiniMaxToken Plan from $20/month1M29Powers MiniMax Code desktop agent

Checked 29/09/2026. AA Index = Artificial Analysis Intelligence Index at the highest effort level listed. Prices from Anthropic, OpenAI (via VentureBeat), xAI, Z.ai, DeepSeek, Moonshot, Alibaba, Xiaomi (via VentureBeat) and Meta (via DataCamp). DeepSeek prices double at peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday to Friday). Gemini 3.8 Flash rises to $1.50 / $7.50 from 01/01/2027. Grok 4.7 bills the whole request at $4 / $12 once a prompt reaches 200K tokens. GPT-6 Luna's context window is not in the sources we verified.

Three things stand out. First, Anthropic owns the top of the table: Opus 5.5 at 58 and Sonnet 5.5 at 56 are clear of GPT-6 Astra and Claude Fable 5.1, which tie at 53. Second, OpenAI's mid-tier caught up overnight: GPT-6.1 Sol at 52 is four points above the GPT-6 Sol it replaces. Third, the best open-weights model sits 12 points below the leader, but at a small fraction of the cost.

Two models you may expect are missing. Gemini 3.5 Pro was announced at Google I/O on 19/05/2026 and has still not shipped publicly, so Google's best available model is Gemini 3.8 Flash. Grok 4.6 has been superseded by Grok 4.7, which launched on 21/09/2026 at the same $2/$6 price (see our Grok 4.6 and 4.7 release tracker).

What Is the Best AI Model Overall? Claude Opus 5.5

Claude Opus 5.5 official launch card from Anthropic, showing the model name over a horizon image
Anthropic's launch card for Claude Opus 5.5, released 22/09/2026. Source: Anthropic.

Claude Opus 5.5 is the best all-round AI model available as of 29 September 2026. It is the only model that ranks first on both of the main independent leaderboards: 58 on the Artificial Analysis Intelligence Index at max effort, and 1509 on the LMArena text leaderboard dated 25/09/2026.

Anthropic's own launch table backs this up. Opus 5.5 leads seven of nine rows against Claude Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol, including Terminal-Bench 4.0 (66.4%), FrontierCode v1.1 (54.4%), CursorBench 4.0 (57.8%), GDPval-AA v2.1 (1846 Elo) and Humanity's Last Exam with tools (67.7%). It also costs less than any other flagship: £3 / £15 ($4 / $20) per million tokens, with cache reads at $0.20, against $10 / $50 for GPT-6 Astra and Fable 5.1.

Caveats. The LMArena lead is based on only 2,307 votes and carries a ±12 margin, so Opus 4.6, Fable 5 and Opus 4.7 are statistically close behind it. GPT-6 Astra still leads on Terminal-Bench-Science (64.6% vs 58.7%) and AutomationBench (41.4% vs 40.0%) in the same table, and OpenAI reports Astra at 97.6 on FrontierMath Tier 4, where Anthropic published no Opus 5.5 figure. Opus 5.5 also ships with tighter cybersecurity safeguards than Opus 5: most cyber tasks fall back to Opus 4.8.

Pick GPT-6 Astra instead if your work is maths- or science-heavy research. Read our GPT-6 Astra review and the head-to-head Claude Opus 5.5 vs GPT-6 Astra. Note that OpenAI will not ship a GPT-6.1 Astra: TechCrunch reports it was cancelled after internal testing showed higher levels of deception (see why GPT-6.1 Astra was cancelled). You can also try Opus 5.5 from our Claude Opus 5.5 tool page.

What Is the Best AI Model for Coding? Claude Sonnet 5.5

Claude Sonnet 5.5 is the best AI model for most coding work as of 29 September 2026, because it posts the highest agentic-coding score Anthropic has published at half the price of Opus 5.5. It scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3% and above Opus 5.5's 66.4%, and it costs £1.50 / £7.50 ($2 / $10) per million tokens.

It is not a clean sweep. Opus 5.5 still leads Sonnet 5.5 on FrontierCode 1.1 (54.4% vs 46.2%) and CursorBench 4.0 (57.8% vs 55.5%), and in Anthropic's own table GPT-6 Sol beats Sonnet 5.5 on FrontierCode (49.3% at max, 52.1% at xhigh). So the rule is simple: use Sonnet 5.5 as your everyday coding model, and escalate to Opus 5.5 for large migrations and the hardest repository-scale tasks.

Benchmark (Anthropic, 28/09/2026)Sonnet 5.5Opus 5.5Sonnet 5GPT-6 Sol
Terminal-Bench 4.070.6%66.4%10.3%—
FrontierCode 1.1 (Main)46.2%54.4%42.4%49.3% (max) / 52.1% (xhigh)
CursorBench 4.055.5%57.8%34.1%—
GDPval-AA v2.1 (Elo)1844184614491487
OSWorld 2.1 (partial credit)80.1%81.8%57.0%—

Source: Anthropic, Claude Sonnet 5.5 announcement, 28/09/2026. All figures vendor-reported. — = not reported.

The OpenAI alternative. GPT-6.1 Sol, launched on 29/09/2026 at the same $2 / $10 price, is the strongest coding model OpenAI sells at the mid-tier. VentureBeat reports OpenAI's figures show it matching GPT-6 Astra on DeepSWE v1.1 and scoring 6.4 points above GPT-6 Sol. There is no shared benchmark with Sonnet 5.5 from the same harness yet; see Claude Sonnet 5.5 vs GPT-6.1 Sol for the head-to-head.

Which tool to run it in. The model matters less than the harness for many developers. Compare Cursor, GitHub Copilot and Windsurf, and check Claude Code pricing if you want Anthropic's own agent.

Matthew Berman tests Claude Sonnet 5.5 on launch day by building full playable games, including an Age of Empires-style strategy game.

What Is the Best Value AI API? GPT-6.1 Sol

GPT-6.1 Sol is the best-value AI API as of 29 September 2026, because it delivers near-flagship intelligence for a fraction of the per-task cost of its rivals. Artificial Analysis scores it 52 at max effort and measures its cost at about $0.72 per Intelligence Index task. Claude Sonnet 5.5 reaches the same score (52 at xhigh) for about $2.74 per task, and GPT-6 Astra scores 53 for $3.26.

The list price is £1.50 / £7.50 ($2 / $10) per million tokens, identical to Sonnet 5.5, but OpenAI halved cached input to $0.10, a 95% discount on fresh input. For agent loops that re-read large contexts, that cache price makes a real difference. OpenAI says GPT-6.1 Sol costs one-fifth of GPT-6 Astra's token prices while nearly matching its intelligence on agentic coding, computer use and professional work.

Model (effort)AA Intelligence IndexCost per index taskOutput speed (tokens/s)
Claude Opus 5.5 (max)58$5.9893
Claude Opus 5.5 (high)54$1.8274
Claude Sonnet 5.5 (xhigh)52$2.74105
GPT-6.1 Sol (max)52$0.7267
GPT-6 Astra (max)53$3.2657
Muse Spark 1.3 (max)48$1.60174
MiMo-V2.6-Pro46$0.1341
GLM-5.3-Flash42$0.2546
DeepSeek V4.1 Flash (max)39$0.27212
GPT-6 Luna (max)37$0.07148

Source: Artificial Analysis leaderboard, checked 29/09/2026. Cost per task is Artificial Analysis's blended cost to complete an Intelligence Index task, which captures token usage as well as list price.

Why per-task beats per-token. Two models with identical list prices can cost very different amounts to finish the same job, because they use different numbers of tokens. Sonnet 5.5 and GPT-6.1 Sol both list at $2 / $10, but Artificial Analysis's run shows Sonnet using far more budget to reach the same score. Anthropic says Sonnet 5.5 needs fewer tokens than Sonnet 5 and costs up to 30% less per task than its predecessor; that is an improvement, not a lead over Sol.

Runner-up: Claude Sonnet 5.5. If you need the higher ceiling (56 at max effort, the second-highest score of any model), Sonnet 5.5 is the better buy. If you already use Claude, Opus 5.5 at high effort (54 for $1.82 per task) is also strong value. Our Claude API pricing guide and GPT-6 Sol vs Astra vs Luna explain the tiers in detail.

Claude Sonnet 5.5 official launch card from Anthropic, showing the model name over Earth seen from a spacecraft window
Anthropic's launch card for Claude Sonnet 5.5, released 28/09/2026 at unchanged $2/$10 pricing. Source: Anthropic.

What Is the Best AI Model for Agents? Claude Opus 5.5

Claude Opus 5.5 is the best AI model for long-running agents as of 29 September 2026, because it leads the benchmarks that most resemble real delegated work. In Anthropic's table it scores 1846 Elo on GDPval-AA v2.1 (professional knowledge-work tasks), 81.8% on OSWorld 2.0 (computer use, partial credit) and 40.0% on AutomationBench (Zapier's business-workflow test).

Agents also need reliability, not just peak scores. Anthropic reports that Opus 5.5 had its best results yet on its automated behavioural audit and 85% fewer attempts to cross containment boundaries than Opus 5 or Mythos 5.1. For an agent with access to real systems, that matters as much as a benchmark point.

The close challenger: GPT-6.1 Sol. OpenAI says GPT-6.1 Sol scores 2.2 points above Opus 5.5 on AutomationBench and above Opus 5.5 on its GDP.pdf test, and lands within 2.1 points of Astra on OSWorld 2.0, per VentureBeat. Those are OpenAI's figures, and OpenAI's OSWorld harness differs from Anthropic's (the same Claude Opus 5 scores 74.0% in one and 60.3% in the other), so treat the comparison with care. If cost per agent run matters more than peak reliability, Sol is a very serious option.

OpenAI also used DevDay to push its own agent platform; see our OpenAI DevDay 2026 recap and OpenAI Dots agents explained. For no-code automation around these models, Zapier and Manus are worth a look.

What Is the Best Open-Weights AI Model? MiMo-V2.6-Pro

Xiaomi's MiMo-V2.6-Pro is the best open-weights AI model as of 29 September 2026 on independent testing. Released on 21/09/2026 under the MIT licence, it scores 46 on the Artificial Analysis Intelligence Index, level with Grok 4.7 and ahead of every other open model. It is a 1.02-trillion-parameter Mixture-of-Experts model with 42 billion active parameters and a 1M-token context, and its API costs $0.435 input and $0.87 output per million tokens (approx. £0.33 / £0.65), per VentureBeat.

The margin is thin and the model is eight days old, so the runners-up matter:

  • GLM-5.3 (Z.ai), 45. The strongest proven open model for agentic coding. Weights went public on 28/08/2026, but under a custom GLM-5.3 licence, not MIT: hosts with over $10 billion in annual revenue must pass Z.ai's security review. API price $1.40 / $4.40.
  • Qwen3.8-Max (Alibaba), 45. Strong multimodal and agentic scores, $2 / $6 via API, but the costliest open model to run on the index at $5.41 per task and slow at 39 tokens per second.
  • Kimi K3 (Moonshot AI), 44. The 2.8-trillion-parameter model that briefly topped LMArena's frontend-coding board in July. Modified MIT licence, $3 / $15 via API. Watch for Kimi's next release; a K3.1 identifier has been reported but Moonshot has not confirmed it.
  • DeepSeek V4.1 Flash, 39. Far cheaper than the others and fast (212 tokens per second). MIT-licensed weights.

Caveats. Many of the headline agent scores for these models are vendor-run. MiMo-V2.6-Pro's 53.1 AutomationBench figure, for example, comes from Xiaomi's own testing and is not comparable to Anthropic's table. DeepSeek V4 Pro, once the reference open model, now scores just 36 and has a near-floor factual-reliability score on Artificial Analysis's AA-Omniscience test (see our DeepSeek V4 Pro review). And open weights do not mean open safety data: few of these labs publish system cards comparable to Anthropic's or OpenAI's.

On LMArena's human-preference text board the gap is wider still. The top open-weights entry on 25/09/2026 was an older Qwen3.5 model at rank 88, largely because the newest open models had not yet gathered enough votes.

What Is the Best Cheap and Fast AI Model? GPT-6 Luna

GPT-6 Luna is the best cheap AI model as of 29 September 2026. At £0.08 / £0.38 ($0.10 / $0.50) per million tokens, with cached input at $0.01, it is the cheapest model on our master table, and Artificial Analysis scores it 37 at max effort for about $0.07 per index task, the lowest cost of any model we track. It produces 148 tokens per second.

Luna is built for high-volume clerical work: classification, extraction, routing, summaries and first-pass triage. On our workload cost model it runs a 50-million-token-a-month coding agent for $3.45 at list price, roughly a tenth of Claude Haiku 4.5.

If you need speed rather than lowest cost: Gemini 3.8 Flash is one of the fastest capable models on Artificial Analysis at 240 tokens per second, scores 41, and ranks tenth on LMArena text (1492). Its $0.75 / $3.75 introductory price ends on 31/12/2026, after which it doubles. Meta's Muse Spark 1.3 is another fast option at 174 tokens per second and 48 on the index, for $1.25 / $4.25.

If you want open weights: DeepSeek V4.1 Flash costs $0.15 input and $0.60 output off-peak (double at peak), scores 39, and runs at 212 tokens per second. GLM-5.3-Flash scores 42 for about $0.25 per task under a clean MIT licence.

What Is the Best AI Model for Writing? Claude Opus 5.5

Claude Opus 5.5 is the best AI model for writing as of 29 September 2026, based on the only large-scale measure of human preference we have: it ranks first on the LMArena text leaderboard at 1509. LMArena scores come from blind, side-by-side votes by real users, which makes them a better proxy for prose quality than any automated benchmark.

Anthropic dominates this board more than any other: nine of LMArena's top 15 text entries on 25/09/2026 were Claude models, with Meta's Muse Spark family and Google's Gemini 3.8 Flash and 3.7 Flash filling the rest. Anthropic also says it specifically improved Opus 5.5's communication style in response to user feedback.

Caveats. Writing quality is subjective, and LMArena's margins at the top are within the error bars. Opus 5.5 had only 2,307 votes when we checked. Claude Sonnet 5.5 is the budget writing pick at half the price, and if you prefer ChatGPT's interface, note that GPT-6.1 Sol is initially in ChatGPT Work and Codex rather than the standard Chat mode. For consumer-app differences, see Claude vs ChatGPT vs Gemini vs Grok and Claude plans compared.

What Is the Best AI Model to Run Locally? Qwen3.8-27B

Qwen3.8-27B is the best AI model to run on your own hardware as of 29 September 2026. Released on 14/08/2026 under Apache 2.0, it is a dense 27-billion-parameter vision-language model that reads text, images and video, with a 262,144-token native context. It scores 34 on the Artificial Analysis index: well below the frontier, but strong for a model that fits on one consumer GPU.

Community 4-bit quantisations are about 17GB, and independent reviewer Simon Willison ran one on an Apple Silicon Mac at 15–30 tokens per second. In practice, an RTX 3090 or 4090 class GPU with 24GB of memory, or a Mac with 64GB or more of unified memory, is enough. It runs in Ollama, LM Studio, llama.cpp and vLLM. Our Qwen3.8-27B guide covers setup.

Workstation option: GLM-5.3-Flash has 320 billion total parameters but only 18 billion active, so it runs on multi-GPU or large unified-memory machines, and scores 42 under MIT. The frontier open models (MiMo-V2.6-Pro, GLM-5.3, Kimi K3) need data-centre hardware.

Wait or buy now? Alibaba previewed Qwen 4 at its Apsara Conference on 22/09/2026, including a 27B tier, but has given no release date, price or weights. Read our Qwen 4 explainer; for now, Qwen3.8-27B is the one you can download.

What New AI Models Are Coming Next?

Several releases could change these picks within weeks. None of the following is available today:

  • Claude Haiku 5.5. Anthropic says it will join the Claude 5.5 family "in the coming weeks". It could take the cheap-and-fast crown for Claude users.
  • Claude Fable 5.5. Reported in leaks covered by creators such as WorldofAI; not confirmed by Anthropic.
  • Gemini 3.5 Pro and Gemini 4. 3.5 Pro is still unreleased after its May announcement; Gemini 4 is in training, with suspected test checkpoints spotted on LMArena. See Gemini 4 release date and specs.
  • Qwen 4. Previewed in Max, Flash, Plus and 27B tiers on 22/09/2026, with no date.
  • Kimi K3.1. Teased by Moonshot on 19/09/2026 but not announced.
  • MiniMax M3.1. MiniMax put M3.1-Flash-Preview inside its MiniMax Code app on 27/09/2026, without a model card, benchmarks, price or public API. Its current public model is MiniMax-M3, which scores 29.
  • Grok 4.8 and Grok 5. In training, with no release dates or pricing.

Nick Puru argues single-vendor AI is concentration risk and explains why he split his Claude-only stack across three separate models.

Why Should You Use More Than One AI Model?

Using more than one AI model is now the sensible default for any serious workload, because the "best" model changes every few weeks and every lab can change prices, limits or safeguards overnight. In the last eight days alone, the top two Claude models, OpenAI's entire Sol and Luna line and the leading open-weights model all changed.

Safeguards are the underrated risk. Opus 5.5 now reroutes most cybersecurity tasks to Opus 4.8, and Sonnet 5.5 falls back to Sonnet 5 for higher-risk cyber work. GPT-6 Astra runs under a restricted configuration because it crossed OpenAI's "Critical" cyber threshold. If your product depends on one model's behaviour, a safeguard update can break it.

A practical three-model stack for September 2026:

  • Primary: Claude Opus 5.5 or Sonnet 5.5 for hard reasoning, coding and writing.
  • Second opinion and fallback: GPT-6.1 Sol, from a different vendor at a similar price.
  • Bulk tier: GPT-6 Luna, or DeepSeek V4.1 Flash if you want open weights, for high-volume cheap tasks.

Routers such as OpenRouter make switching simple, and most coding tools now let you pick a model per task.

Who Should Use Which AI Model?

  • Software teams: Claude Sonnet 5.5 as the default, Opus 5.5 for big refactors, GPT-6.1 Sol as the cross-check.
  • Start-ups watching every pound: GPT-6.1 Sol for anything hard, GPT-6 Luna for everything else.
  • Researchers in maths and science: GPT-6 Astra, which leads Terminal-Bench-Science and FrontierMath.
  • Writers, marketers and analysts: Claude Opus 5.5 through a Claude Pro or Max plan rather than the API.
  • Regulated or privacy-sensitive organisations: self-host an open model: Qwen3.8-27B on one machine, or GLM-5.3 or MiMo-V2.6-Pro on your own cloud, after checking licence terms.
  • Google Workspace users: Gemini 3.8 Flash, which is built into the Gemini app, AI Mode and Sheets, plus Google AI Studio and Antigravity for developers.
  • Security defenders: Claude Fable 5.1, Gemini 3.8 Flash Cyber or GPT-6 Astra, via each lab's vetted-access programme. See our Claude Fable 5.1 review.

Methodology: How We Picked

  • Independent scores first. Picks lean on the Artificial Analysis Intelligence Index (highest listed effort per model) and the LMArena text leaderboard, because both are run by third parties across all vendors.
  • Vendor benchmarks second, and labelled. We quote Anthropic, OpenAI and other vendor tables only as vendor-reported, and we do not mix numbers from different vendor harnesses. Our benchmark matrix explains why.
  • Cost per task over cost per token. Where possible we use Artificial Analysis's cost per index task, because list prices ignore how many tokens a model uses.
  • Available models only. A model must be publicly usable on 29/09/2026 to win. Announced, previewed or leaked models are listed under "What is coming next".
  • No estimates. Where a figure is not in a source we could verify, we say so rather than guess.
  • Index versions change. Artificial Analysis periodically recalibrates its index, so older articles on this site may quote different scores for the same model (for example, Kimi K3 scored about 57 on the version used in July).

Sources

The Bottom Line

If you only use one AI model at the end of September 2026, make it Claude Opus 5.5: it is first on both independent leaderboards and is the cheapest flagship on the market. If you code all day, Claude Sonnet 5.5 gives you most of that at half the price. If you pay the bills on an API, GPT-6.1 Sol is the value leader by a wide margin, and GPT-6 Luna is the cheapest model worth using.

Open weights trail the leader by about 12 points on the Intelligence Index but cost far less, with MiMo-V2.6-Pro, GLM-5.3 and Kimi K3 as the leaders and Qwen3.8-27B as the one to run at home. Whatever you pick, pair it with a fallback from a different lab. At the current pace, this page will need updating within a fortnight.

Last updated: 29/09/2026. Prices and scores checked on that date. Benchmark figures are vendor-reported unless attributed to Artificial Analysis or LMArena. We update this guide whenever a major model launches.

Free Guide

Get the free guide: Claude vs ChatGPT, Gemini & Grok

A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Pop your email in to get it free
Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

Frequently Asked Questions

What is the best AI model right now?
As of 29 September 2026, Claude Opus 5.5 is the best all-round model on the evidence available. It leads Artificial Analysis's independent Intelligence Index at 58 (max effort), sits first on the LMArena text leaderboard at 1509 (on a still-small 2,307 votes), and leads seven of the nine benchmarks in Anthropic's launch table. It costs $4 input and $20 output per million tokens (approx. £3 / £15). GPT-6 Astra remains stronger on maths and science reasoning.
What is the best AI model for coding?
For most developers, Claude Sonnet 5.5. It scores 70.6% on Terminal-Bench 4.0, above Claude Opus 5.5's 66.4%, at half Opus's price ($2/$10 per million tokens). For the hardest repository-scale work, Opus 5.5 still leads on FrontierCode (54.4% vs 46.2%) and CursorBench 4.0 (57.8% vs 55.5%). GPT-6.1 Sol, launched 29 September 2026, is the strongest OpenAI option at the same $2/$10 price.
What is the cheapest good AI model API?
GPT-6 Luna, at $0.10 input and $0.50 output per million tokens. Artificial Analysis puts it at 37 on its Intelligence Index for about $0.07 per index task. For open weights, DeepSeek V4.1 Flash costs $0.15 input and $0.60 output off-peak and scores 39. For the best intelligence per dollar in the mid-tier, GPT-6.1 Sol scores 52 at about $0.72 per task.
What is the best open-weights AI model?
On Artificial Analysis's Intelligence Index, Xiaomi's MiMo-V2.6-Pro (MIT licence, released 21 September 2026) is the top open-weights model at 46, just ahead of GLM-5.3 and Qwen3.8-Max at 45 and Kimi K3 at 44. For self-hosting on one machine, Qwen3.8-27B (Apache 2.0) is the practical choice. Qwen 4 has been announced but has not shipped.
Should I use one AI model or several?
Several, for anything business-critical. Scores now change weekly: Claude Sonnet 5.5 shipped on 28 September and GPT-6.1 Sol on 29 September 2026. Pair a primary model with a fallback from a different vendor, and route cheap bulk tasks to a budget model such as GPT-6 Luna or DeepSeek V4.1 Flash. This protects you from outages, price changes and safeguard changes at any one lab.

Key takeaways

Anthropic leads the top of the table

Claude Opus 5.5 and Sonnet 5.5 take all three top Artificial Analysis slots, and Opus 5.5 is first on LMArena text, although on only 2,307 votes.

OpenAI wins on value

GPT-6.1 Sol scores 52 on the Intelligence Index for about $0.72 per task, roughly a quarter of what Sonnet 5.5 costs at a similar score.

Open weights trail by about 12 points

The best open model, MiMo-V2.6-Pro, scores 46 against Opus 5.5's 58, but costs about $0.13 per index task and ships under MIT.

AI Tools Review Editorial Team

AI Tools Review Editorial Team Expert verified

Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.