AI Tools Review
Benchmarks

AI Model Benchmarks

More than 20 frontier models compared on intelligence, agentic coding and cost per task, GPT-5.6 Sol, Terra and Luna against Claude, Gemini, Grok, GLM, DeepSeek, Qwen, Kimi and more. Every figure is real and sourced; unpublished values are shown as “n/a”, never guessed.

The frontier, ranked

Artificial Analysis Intelligence Index v4.1 composite. Higher is better. GPT-5.6 models are marked.

OpenAIAnthropicGooglexAIZ.aiDeepSeekAlibabaKimiMiniMaxMetaNVIDIAMistralXiaomi

Intelligence vs cost per task

Intelligence Index against weighted average cost (USD) per Intelligence Index task, log scale. Up-and-left is better value: the shaded region is the most attractive quadrant.

OpenAIAnthropicGooglexAIZ.aiDeepSeekAlibabaKimiMiniMaxNVIDIAMistralXiaomi

Models are plotted where Artificial Analysis has published both an Intelligence Index and a cost per task. GPT-5.6 models are ringed in white.

Knowledge work: AA-Briefcase

Artificial Analysis's benchmark of realistic knowledge-work tasks in complex projects, built by industry experts. Only the published head-to-head figures are shown.

Claude Fable 5 (max)
AA-Briefcase leader
Rubric score
56%
Analytical Quality Elo
1,764
Presentation
Strong
GPT-5.6 Sol (max)
Runner-up
Rubric score
42%
Analytical Quality Elo
1,592
Presentation
Highest Presentation Elo of any model

GPT-5.6 Sol's PowerPoint, Excel and document outputs were judged the most visually attractive of any model; Claude Fable 5 leads on rubric substance and analytical quality.

Individual evaluations

The evals that are directly comparable across labs, from each provider's official launch table. Evals whose variants differ between labs (OSWorld 2.0 vs OSWorld-Verified, HLE with vs without tools) are deliberately excluded rather than mixed.

ModelSWE-bench ProTerminal-Bench 2.1GDPval-AA v2 (Elo)
Claude Fable 5
Anthropic · max, with fallback
80.3%83.4%n/a
Claude Opus 5
Anthropic · max
79.2%n/a1,861
Claude Opus 4.8
Anthropic · max
69.2%74.6%1,615
GPT-5.6 Sol
OpenAI · max
64.6%88.8%n/a
GPT-5.6 Terra
OpenAI · max
63.4%87.4%n/a
Claude Sonnet 5
Anthropic · max
63.2%80.4%1,618
GPT-5.6 Luna
OpenAI · max
62.7%84.7%n/a
GLM-5.2
Z.ai · max
62.1%n/an/a
GPT-5.5
OpenAI · xhigh
59.4%85.6%1,769
Gemini 3.5 Flash
Google
55.1%76.2%1,656
Gemini 3.1 Pro Preview
Google
54.2%70.3%1,314

Column leaders in bold. Figures are each provider's self-reported launch results (harnesses and prompt scaffolds vary by lab), plus GDPval-AA v2 Elo where published; treat cross-lab deltas of a point or two as noise.

Full comparison

ModelIntelligenceCodingCost / taskTokens / taskInput $/1MOutput $/1MContextServing (OR)
Claude Fable 5
Anthropic · max, with fallback
6077
Claude Code
$2.75n/an/an/a1M6 eps
97.88% up
GPT-5.6 SolNew
OpenAI · max
5980
Codex
$1.0415k$5$301.5M4 eps
94.25% up
Claude Opus 4.8
Anthropic · max
5673
Claude Code
$1.80n/a$5$251M7 eps
99.87% up
GPT-5.6 TerraNew
OpenAI · max
5577
Codex
$0.55n/a$2.5$151.5M4 eps
99.7% up
GPT-5.5
OpenAI · xhigh
5576
Codex
$0.8616k$5$301M5 eps
94.54% up
Grok 4.5
xAI · high
5476
Grok Build
$0.31n/a$2$6500k°4 eps
98.77% up
Claude Sonnet 5
Anthropic · max
53n/a$1.53n/a$3$151M7 eps
99.76% up
GPT-5.6 LunaNew
OpenAI · max
5175
Codex
$0.21n/a$1$61.5M5 eps
99.53% up
GLM-5.2
Z.ai · max
5158
Claude Code
$0.32n/a$1.4$4.41M24 eps
99.27% up
Gemini 3.5 Flash
Google
50n/a$0.59n/a$1.5$91.049M°6 eps
98.74% up
Gemini 3.1 Pro Preview
Google
4643
Gemini CLI
$0.29n/an/an/a1.049M°3 eps
99.38% up
Qwen3.7 Max
Alibaba
46n/a$1.06n/a$1.25$3.751M°1 eps
100% up
DeepSeek V4 Pro
DeepSeek · max
4447
Claude Code
$0.04n/a$0.435$0.871M15 eps
98.59% up
Kimi K2.6
Kimi
44n/a$0.35n/a$0.95$4262k°17 eps
99.19% up
MiniMax-M3
MiniMax
44n/a$0.12n/a$0.3$1.21M°6 eps
97.53% up
Muse Spark
Meta
43n/an/an/an/an/an/an/a
MiMo-V2.5-Pro
Xiaomi · max
42n/a$0.03n/a$0.43$0.871.049M°3 eps
99.76% up
GPT-5.4 mini
OpenAI · xhigh
40n/a$0.48n/an/an/a400k°4 eps
72.76% up
Nemotron 3 Ultra
NVIDIA
38n/a$0.25n/an/an/a512k°2 eps
97.73% up
Claude 4.5 Haiku
Anthropic
30n/a$0.24n/a$1$5200k°6 eps
99.81% up
Mistral Medium 3.5
Mistral
30n/a$1.20n/a$1.5$7.5262k°1 eps
99.97% up
gpt-oss-120b
OpenAI · high
24n/a$0.04n/an/an/a131k°17 eps
99.19% up
Claude Opus 5New
Anthropic · max
n/an/an/an/a$5$251Mn/a
Composer 2.5 Fast
Cursor
n/a52
Cursor CLI
n/an/an/an/an/an/a

Updated 10 July 2026. Source: Artificial Analysis, Intelligence Index v4.1, Coding Agent Index, cost per Intelligence Index task and model pages (9–10 July 2026), plus provider-published API pricing and official launch evaluation tables. Artificial Analysis. Cost per task is AA’s weighted average cost per Intelligence Index task; tokens per task is output-token usage on the same suite. “Serving (OR)” is observed on OpenRouter (2026-07-10): live provider endpoints and median 1-day uptime; ° marks a context length observed on OpenRouter serving endpoints rather than provider-claimed. Values shown as “n/a” are not published.

Where this data comes from

Independent measurement

The Intelligence Index, Coding Agent Index, cost per task and AA-Briefcase figures are from Artificial Analysis, which runs its evaluations independently against live provider APIs: these are not vendor claims.

Self-reported launch figures

The individual evaluations (SWE-bench Pro, Terminal-Bench 2.1, GDPval-AA v2) come from each lab's official launch tables and system cards. They are self-reported: harnesses and prompt scaffolds vary by lab, so treat cross-lab deltas of a point or two as noise.

Provider price lists

Per-token pricing is each provider's published API rate card, cross-checked against OpenRouter listings. Promotional rates are noted where they apply. Unpublished values are shown as “n/a”, never estimated.

Observed serving data (OpenRouter)

The “Serving” column is snapshotted from OpenRouter's public API: live provider endpoints per model, median 1-day uptime, and the maximum context length actually served, which sometimes differs from the claimed figure (GPT-5.6 serves at 1.05M tokens on OpenRouter against the reported 1.5M). Throughput and latency are not exposed on the public API, so we do not chart them.

Our own derivation: the AITR Value For Money Index

The AITR Value For Money Index is ours, Intelligence Index points per US dollar of cost per task (Intelligence ÷ cost). It is pure declared arithmetic on the sourced inputs above, with no editorial weighting, and it surfaces what the headline charts hide: the value frontier belongs to the cheap open-weights models.

Want the story behind the numbers? Read our deep dives on GPT-5.6 Sol, Terra & Luna, Claude Opus 4.8, GLM 5.3 and DeepSeek V4 Pro.