AI Model Benchmarks
More than 20 frontier models compared on intelligence, agentic coding and cost per task, GPT-5.6 Sol, Terra and Luna against Claude, Gemini, Grok, GLM, DeepSeek, Qwen, Kimi and more. Every figure is real and sourced; unpublished values are shown as “n/a”, never guessed.
The frontier, ranked
Artificial Analysis Intelligence Index v4.1 composite. Higher is better. GPT-5.6 models are marked.
Intelligence vs cost per task
Intelligence Index against weighted average cost (USD) per Intelligence Index task, log scale. Up-and-left is better value: the shaded region is the most attractive quadrant.
Models are plotted where Artificial Analysis has published both an Intelligence Index and a cost per task. GPT-5.6 models are ringed in white.
Knowledge work: AA-Briefcase
Artificial Analysis's benchmark of realistic knowledge-work tasks in complex projects, built by industry experts. Only the published head-to-head figures are shown.
- Rubric score
- 56%
- Analytical Quality Elo
- 1,764
- Presentation
- Strong
- Rubric score
- 42%
- Analytical Quality Elo
- 1,592
- Presentation
- Highest Presentation Elo of any model
GPT-5.6 Sol's PowerPoint, Excel and document outputs were judged the most visually attractive of any model; Claude Fable 5 leads on rubric substance and analytical quality.
Individual evaluations
The evals that are directly comparable across labs, from each provider's official launch table. Evals whose variants differ between labs (OSWorld 2.0 vs OSWorld-Verified, HLE with vs without tools) are deliberately excluded rather than mixed.
| Model | SWE-bench Pro | Terminal-Bench 2.1 | GDPval-AA v2 (Elo) |
|---|---|---|---|
Claude Fable 5 Anthropic · max, with fallback | 80.3% | 83.4% | n/a |
Claude Opus 5 Anthropic · max | 79.2% | n/a | 1,861 |
Claude Opus 4.8 Anthropic · max | 69.2% | 74.6% | 1,615 |
GPT-5.6 Sol OpenAI · max | 64.6% | 88.8% | n/a |
GPT-5.6 Terra OpenAI · max | 63.4% | 87.4% | n/a |
Claude Sonnet 5 Anthropic · max | 63.2% | 80.4% | 1,618 |
GPT-5.6 Luna OpenAI · max | 62.7% | 84.7% | n/a |
GLM-5.2 Z.ai · max | 62.1% | n/a | n/a |
GPT-5.5 OpenAI · xhigh | 59.4% | 85.6% | 1,769 |
Gemini 3.5 Flash Google | 55.1% | 76.2% | 1,656 |
Gemini 3.1 Pro Preview Google | 54.2% | 70.3% | 1,314 |
Column leaders in bold. Figures are each provider's self-reported launch results (harnesses and prompt scaffolds vary by lab), plus GDPval-AA v2 Elo where published; treat cross-lab deltas of a point or two as noise.
Full comparison
| Model | Intelligence | Coding | Cost / task | Tokens / task | Input $/1M | Output $/1M | Context | Serving (OR) |
|---|---|---|---|---|---|---|---|---|
Claude Fable 5 Anthropic · max, with fallback | 60 | 77 Claude Code | $2.75 | n/a | n/a | n/a | 1M | 6 eps 97.88% up |
GPT-5.6 SolNew OpenAI · max | 59 | 80 Codex | $1.04 | 15k | $5 | $30 | 1.5M | 4 eps 94.25% up |
Claude Opus 4.8 Anthropic · max | 56 | 73 Claude Code | $1.80 | n/a | $5 | $25 | 1M | 7 eps 99.87% up |
GPT-5.6 TerraNew OpenAI · max | 55 | 77 Codex | $0.55 | n/a | $2.5 | $15 | 1.5M | 4 eps 99.7% up |
GPT-5.5 OpenAI · xhigh | 55 | 76 Codex | $0.86 | 16k | $5 | $30 | 1M | 5 eps 94.54% up |
Grok 4.5 xAI · high | 54 | 76 Grok Build | $0.31 | n/a | $2 | $6 | 500k° | 4 eps 98.77% up |
Claude Sonnet 5 Anthropic · max | 53 | n/a | $1.53 | n/a | $3 | $15 | 1M | 7 eps 99.76% up |
GPT-5.6 LunaNew OpenAI · max | 51 | 75 Codex | $0.21 | n/a | $1 | $6 | 1.5M | 5 eps 99.53% up |
GLM-5.2 Z.ai · max | 51 | 58 Claude Code | $0.32 | n/a | $1.4 | $4.4 | 1M | 24 eps 99.27% up |
Gemini 3.5 Flash Google | 50 | n/a | $0.59 | n/a | $1.5 | $9 | 1.049M° | 6 eps 98.74% up |
Gemini 3.1 Pro Preview Google | 46 | 43 Gemini CLI | $0.29 | n/a | n/a | n/a | 1.049M° | 3 eps 99.38% up |
Qwen3.7 Max Alibaba | 46 | n/a | $1.06 | n/a | $1.25 | $3.75 | 1M° | 1 eps 100% up |
DeepSeek V4 Pro DeepSeek · max | 44 | 47 Claude Code | $0.04 | n/a | $0.435 | $0.87 | 1M | 15 eps 98.59% up |
Kimi K2.6 Kimi | 44 | n/a | $0.35 | n/a | $0.95 | $4 | 262k° | 17 eps 99.19% up |
MiniMax-M3 MiniMax | 44 | n/a | $0.12 | n/a | $0.3 | $1.2 | 1M° | 6 eps 97.53% up |
Muse Spark Meta | 43 | n/a | n/a | n/a | n/a | n/a | n/a | n/a |
MiMo-V2.5-Pro Xiaomi · max | 42 | n/a | $0.03 | n/a | $0.43 | $0.87 | 1.049M° | 3 eps 99.76% up |
GPT-5.4 mini OpenAI · xhigh | 40 | n/a | $0.48 | n/a | n/a | n/a | 400k° | 4 eps 72.76% up |
Nemotron 3 Ultra NVIDIA | 38 | n/a | $0.25 | n/a | n/a | n/a | 512k° | 2 eps 97.73% up |
Claude 4.5 Haiku Anthropic | 30 | n/a | $0.24 | n/a | $1 | $5 | 200k° | 6 eps 99.81% up |
Mistral Medium 3.5 Mistral | 30 | n/a | $1.20 | n/a | $1.5 | $7.5 | 262k° | 1 eps 99.97% up |
gpt-oss-120b OpenAI · high | 24 | n/a | $0.04 | n/a | n/a | n/a | 131k° | 17 eps 99.19% up |
Claude Opus 5New Anthropic · max | n/a | n/a | n/a | n/a | $5 | $25 | 1M | n/a |
Composer 2.5 Fast Cursor | n/a | 52 Cursor CLI | n/a | n/a | n/a | n/a | n/a | n/a |
Updated 10 July 2026. Source: Artificial Analysis, Intelligence Index v4.1, Coding Agent Index, cost per Intelligence Index task and model pages (9–10 July 2026), plus provider-published API pricing and official launch evaluation tables. Artificial Analysis. Cost per task is AA’s weighted average cost per Intelligence Index task; tokens per task is output-token usage on the same suite. “Serving (OR)” is observed on OpenRouter (2026-07-10): live provider endpoints and median 1-day uptime; ° marks a context length observed on OpenRouter serving endpoints rather than provider-claimed. Values shown as “n/a” are not published.
Where this data comes from
Independent measurement
The Intelligence Index, Coding Agent Index, cost per task and AA-Briefcase figures are from Artificial Analysis, which runs its evaluations independently against live provider APIs: these are not vendor claims.
Self-reported launch figures
The individual evaluations (SWE-bench Pro, Terminal-Bench 2.1, GDPval-AA v2) come from each lab's official launch tables and system cards. They are self-reported: harnesses and prompt scaffolds vary by lab, so treat cross-lab deltas of a point or two as noise.
Provider price lists
Per-token pricing is each provider's published API rate card, cross-checked against OpenRouter listings. Promotional rates are noted where they apply. Unpublished values are shown as “n/a”, never estimated.
Observed serving data (OpenRouter)
The “Serving” column is snapshotted from OpenRouter's public API: live provider endpoints per model, median 1-day uptime, and the maximum context length actually served, which sometimes differs from the claimed figure (GPT-5.6 serves at 1.05M tokens on OpenRouter against the reported 1.5M). Throughput and latency are not exposed on the public API, so we do not chart them.
Our own derivation: the AITR Value For Money Index
The AITR Value For Money Index is ours, Intelligence Index points per US dollar of cost per task (Intelligence ÷ cost). It is pure declared arithmetic on the sourced inputs above, with no editorial weighting, and it surfaces what the headline charts hide: the value frontier belongs to the cheap open-weights models.
Want the story behind the numbers? Read our deep dives on GPT-5.6 Sol, Terra & Luna, Claude Opus 4.8, GLM 5.3 and DeepSeek V4 Pro.
