AI Tools Review

Grok 4.5

By xAI

Released: 2026-07-08

LLM
Coding
Agents
xAI
Value
Paid
New

Grok 4.5 is xAI's frontier return, public since 8 July 2026. It scores 54 on the Intelligence Index - fourth overall - at $0.31 per measured task, the cheapest of any closed frontier-tier model, and ties GPT-5.6 Sol on SWE-Atlas-QnA in its own Grok Build coding harness. Pricing is $2/$6 per million tokens.

Visit Grok 4.5

The value leader among closed models

At $0.31 per Artificial Analysis Intelligence Index task, Grok 4.5 is the cheapest closed frontier-tier model in the current set, roughly a third of the cost of GPT-5.6 Sol per task and a ninth of Fable 5.

Cursor-trained, built for agentic coding

Reported to be trained heavily on Cursor-style coding workflows, it posts a Coding Agent Index of 76 in its own Grok Build harness and ties GPT-5.6 Sol outright on SWE-Atlas-QnA.

xAI is back at the frontier

An Intelligence Index score of 54 puts Grok 4.5 fourth overall (behind Fable 5, GPT-5.6 Sol and Opus 4.8, but a level ahead of Claude Sonnet 5) xAI's strongest showing in this cycle.

Grok 4.5 went public on 8 July 2026 and returns xAI to the intelligence frontier. Coverage has framed it as "SpaceXAI's Grok 4.5": a cheap, Cursor-trained coding model rather than an outright leader. The framing is fair. It sits fourth on the Artificial Analysis Intelligence Index at 54, yet costs just $0.31 per Index task, the lowest of any closed frontier model. This review covers its capabilities, benchmarks, pricing and the caveats around its self-hosted coding harness.

What Grok 4.5 is

Grok 4.5 benchmark scores summary card: positioning and best-for list in the campaign index-card style.
At a glance: where this model fits.

Grok 4.5 is xAI's new flagship model, made publicly available on 8 July 2026. The launch matters less for any single headline number than for what it signals: after a period in which xAI's models sat a clear tier below the leading closed labs, Grok 4.5 puts the company back on the intelligence frontier. Coverage has leaned into the "SpaceXAI" framing, a nod to the organisational gravity behind the release, and described the model, more usefully, as a cheap, Cursor-trained coding model. That second description is the one that should shape your expectations.

On the Artificial Analysis Intelligence Index v4.1, measured at high reasoning effort, Grok 4.5 scores 54. That places it fourth overall, behind Fable 5 (60), GPT-5.6 Sol (59) and Opus 4.8 (56), and a level ahead of Claude Sonnet 5 (53). Fourth place is not a marketing line anyone prints on a homepage, but it is genuinely frontier-tier territory, and it is achieved at a price point none of the three models above it comes close to matching.

So the honest one-sentence summary is this: Grok 4.5 is not the smartest model you can rent, but it is the cheapest way to rent something this smart from a closed lab. The rest of this review unpacks where that trade-off pays off, chiefly in agentic coding and high-volume workloads, and where it does not.

Capabilities and the Grok Build harness

The capability story centres on coding, and specifically on agentic coding: the model operating inside a harness that lets it read files, run commands, edit code and iterate, rather than answering one-shot questions. xAI ships its own harness for this, called Grok Build, and it is within Grok Build that the model's headline coding result was recorded: a Coding Agent Index score of 76.

That 76 puts Grok 4.5 level with GPT-5.5 running in Codex, which is respectable company. More striking is the result on SWE-Atlas-QnA, one of the three evaluations that make up the Coding Agent Index: there, Grok 4.5 ties GPT-5.6 Sol outright. Sol still leads the overall index at 80, so nobody should read this as a changing of the guard, but matching the current coding leader on one of its three constituent evals is a real signal, not noise.

The harness detail deserves emphasis, because it cuts both ways. Scoring a model inside its own vendor-built harness is standard practice, every lab tunes model and scaffold together, but it means the 76 reflects the Grok 4.5 plus Grok Build combination, not the raw model dropped into whatever tooling you already run. If your team lives in a different agentic stack, treat the number as an upper bound until you have run your own trials.

  • Coding Agent Index: 76, measured in xAI's own Grok Build harness
  • Level with GPT-5.5 in Codex on the same index
  • Ties GPT-5.6 Sol on SWE-Atlas-QnA, one of the index's three evals
  • GPT-5.6 Sol leads the overall Coding Agent Index at 80

Benchmarks: fourth on intelligence, competitive on coding

Start with general intelligence. The Artificial Analysis Intelligence Index v4.1 places Grok 4.5 at 54 when run at high reasoning effort. The ordering above it is Fable 5 on 60, GPT-5.6 Sol on 59 and Opus 4.8 on 56; immediately below sits Claude Sonnet 5 on 53. Fourth place, then, but the shape of the table matters. The gap from Grok 4.5 up to Opus 4.8 is two points; the gap up to the leaders is five or six. This is a model within touching distance of the front row, not one watching it from the stands.

On coding, the picture is narrower but arguably stronger. The Coding Agent Index score of 76 (again, recorded in the Grok Build harness) matches GPT-5.5 in Codex, and the tie with GPT-5.6 Sol on SWE-Atlas-QnA is the single most impressive line in Grok 4.5's benchmark record. Sol remains ahead on the index overall at 80, driven by the other two evaluations, so across a full spread of agentic coding work the leader is still the leader.

Two reading notes. First, all figures here are Artificial Analysis measurements from July 2026, and index versions matter: these are v4.1 Intelligence Index numbers and should not be compared against scores from earlier index revisions. Second, benchmark position and practical value are different questions. Fourth on intelligence sounds mid-table until you attach the cost column, which is where Grok 4.5's case actually lives, and where the next section goes.

  • Intelligence Index v4.1 (high effort): 54: fourth, behind Fable 5 (60), GPT-5.6 Sol (59), Opus 4.8 (56); ahead of Claude Sonnet 5 (53)
  • Coding Agent Index: 76 in Grok Build; Sol leads at 80
  • SWE-Atlas-QnA: tied with GPT-5.6 Sol

Where it sits on the Intelligence Index

Artificial Analysis Intelligence Index v4.1 across every scored frontier model: this model highlighted.

Source: Artificial Analysis (9 July 2026). Interactive, hover any bar. Explore the full benchmarks →

Price-performance: the strongest of any closed frontier model

Grok 4.5 benchmark scores specification card: intelligence, coding index, cost per task, API pricing, context window and value score.
The numbers in one card: data from our benchmarks tracker.

Here is the number that defines this release: $0.31. That is what it costs, on Artificial Analysis's methodology, to run one Intelligence Index task on Grok 4.5. The equivalent figure is $1.04 for GPT-5.6 Sol, $1.80 for Opus 4.8 and $2.75 for Fable 5. Put differently, the model five or six Index points ahead of Grok 4.5 costs roughly nine times as much per unit of benchmarked work, and even the nearest premium rival costs more than three times as much. Among closed frontier-tier models in the current set, nothing else is close.

The raw API pricing tells the same story in simpler terms: $2 per million input tokens and $6 per million output tokens. Those are the rates behind our AITR Value For Money Index, our own derivation of intelligence divided by cost per task, on which Grok 4.5 scores 174, the best result of any closed model we track. On our intelligence-versus-cost scatter, it sits on the edge of the most attractive quadrant: above-median intelligence delivered at below-median cost. Very few closed models have ever occupied that corner at launch.

The practical consequence is that Grok 4.5 changes the default calculus for cost-sensitive deployments. Workloads that previously forced a choice between a frontier model you ration and a mid-tier model you run freely can now, plausibly, get frontier-adjacent intelligence at mid-tier economics. That will not matter for teams whose model spend is a rounding error. For everyone running agents at volume, it is the whole story.

Cost per Intelligence Index task, in context

What a unit of benchmarked work actually costs across the field. Lower is better.

Source: Artificial Analysis (9 July 2026). Interactive, hover any bar. Explore the full benchmarks →

Real-world fit: a coding model first

The most load-bearing phrase in the launch coverage is "Cursor-trained". Grok 4.5 was reportedly shaped around the kind of work developers actually do in an AI-native editor: multi-file edits, iterative fixes, tool calls, test-and-retry loops. That lineage shows in the benchmark record, the Grok Build result and the SWE-Atlas-QnA tie both measure exactly this style of work, and it suggests the model's natural home is inside a coding agent or IDE integration rather than as a general-purpose chat assistant.

Deployment logistics look sensible. On OpenRouter, serving endpoints were observed offering a 500k-token context window, with four live endpoints as of our 10 July 2026 snapshot. Half a million tokens is enough headroom for large-repository work (long agent transcripts, sizeable codebases held in context, extended tool-use sessions) without the aggressive pruning smaller windows force on agent builders. Multiple live endpoints at launch also matters practically: it means real routing choice and less exposure to a single provider's rate limits.

Who should shortlist it? Teams running agentic coding at meaningful volume, where the per-task economics compound daily. Startups that want frontier-adjacent capability without frontier invoices. Builders of background agents (CI triage, automated refactoring, batch code review) where a nine-fold cost difference against Fable 5 decides whether the product is viable at all. Teams whose hardest problems demand the absolute best reasoning available should still look one row up the leaderboard, and should expect to pay accordingly.

Limitations

The most obvious limitation is the one the leaderboard states plainly: fourth is not first. At 54 on the Intelligence Index, Grok 4.5 trails Fable 5 by six points and GPT-5.6 Sol by five. On the hardest reasoning-heavy work, the tasks where those top-of-index points actually bite, the premium models remain the better tools, and no cost argument changes that for teams whose bottleneck is capability rather than budget.

The coding score needs its asterisk kept attached. The Coding Agent Index of 76 was recorded in Grok Build, xAI's own harness, and GPT-5.6 Sol leads the index overall at 80. The SWE-Atlas-QnA tie is genuine, but it is one of three evaluations, and Sol's lead comes from the other two. Anyone planning to run Grok 4.5 outside Grok Build (in their own agent framework, or a third-party editor) should validate performance in that environment before committing, because the published number does not measure it.

Two further cautions. The model has been public only since 8 July 2026, so the track record on reliability, regression behaviour and long-tail failure modes is a week old at the time of writing; our OpenRouter observations are a single 10 July snapshot and endpoint counts and context limits can shift. And while 174 is the best Value For Money Index score of any closed model on our books, it is worth remembering the index's ceiling: open-weight models reach far higher, with MiMo topping the chart at 1,400. "Best value among closed models" is a real distinction, but it is a qualified one.

Verdict

Grok 4.5 benchmark scores verdict card with our one-line assessment.
The verdict, briefly.

Grok 4.5 is the most interesting closed-model release of the summer, not because it wins anything outright, but because of where it lands on the map. Fourth on intelligence at 54, level with the coding leader on one of three agentic evals, and priced at $0.31 per Index task against rivals charging $1.04 to $2.75: that combination has no precedent among closed frontier models, and our Value For Money Index score of 174 reflects it.

The recommendation splits cleanly. If your workload is agentic coding, background automation, or anything you run at volume, Grok 4.5 should be on your shortlist today, and quite possibly at the top of it: $2 per million input tokens and $6 per million output, a 500k context window, and a Cursor-trained disposition toward exactly this kind of work make it the obvious value pick. If instead you need the strongest available reasoning for a small number of hard problems, Fable 5 and GPT-5.6 Sol remain the better instruments, and the price difference will not matter to you.

Our verdict: Grok 4.5 does not move the frontier of what AI can do, but it moves the frontier of what frontier-adjacent AI costs, and for most working teams, that is the more consequential frontier. It earns a strong recommendation as a value pick, with the usual caution owed to any model a week into public life.

Grok 4.5 benchmark scores

Artificial Analysis Intelligence Index v4.1 (high effort)54
Coding Agent Index (Grok Build harness)76
Cost per Intelligence Index task$0.31

lower is better; among the cheapest frontier options

AITR Value For Money Index174

site's own derivation: intelligence ÷ cost per task

Artificial Analysis figures, July 2026.

Where Grok 4.5 fits

Agentic coding at volume

Multi-file edits, test-and-fix loops and refactoring runs inside an agent harness. The Cursor-trained lineage and the 76 Coding Agent Index score make this the model's natural home, and $0.31 per Index task means you can afford to let agents iterate.

Background automation and CI agents

Automated code review, CI failure triage, batch refactoring and other always-on agents where per-run cost decides viability. At roughly a ninth of Fable 5's cost per task, workloads that were uneconomic on premium models become routine.

Large-repository and long-context work

The 500k-token context window observed on OpenRouter endpoints leaves room for substantial codebases, long agent transcripts and extended tool-use sessions without aggressive context pruning.

Cost-sensitive product features

Startups and product teams embedding frontier-adjacent intelligence into user-facing features at $2 per million input and $6 per million output tokens, where premium-model pricing would sink the unit economics.

High-throughput general assistance

Drafting, summarisation, analysis and internal tooling where above-median intelligence at below-median cost, the most attractive quadrant on our scatter, beats paying a multiple more for a few extra index points.

Sources & further reading

xAI Model Timeline

Grok 4.5Current

500k tokens (observed) context

Grok 4.5Current
Grok 5

Frequently Asked Questions

How intelligent is Grok 4.5 compared with rival frontier models?

It scores 54 on the Artificial Analysis Intelligence Index v4.1 at high reasoning effort: fourth overall, behind Fable 5 (60), GPT-5.6 Sol (59) and Opus 4.8 (56), and one level ahead of Claude Sonnet 5 (53). It is frontier-tier, but not the leader.

Is Grok 4.5 good at coding?

Yes, with a caveat. It scores 76 on the Coding Agent Index in xAI's own Grok Build harness, level with GPT-5.5 in Codex, and ties GPT-5.6 Sol on SWE-Atlas-QnA, one of the index's three evals. Sol still leads the index overall at 80, and the 76 reflects the model paired with xAI's own harness rather than third-party tooling.

What does Grok 4.5 cost?

API pricing is $2 per million input tokens and $6 per million output tokens. On Artificial Analysis's methodology that works out at $0.31 per Intelligence Index task: the cheapest of any closed frontier-tier model, against $1.04 for GPT-5.6 Sol, $1.80 for Opus 4.8 and $2.75 for Fable 5.

What context window does Grok 4.5 support?

Serving endpoints on OpenRouter were observed offering 500k tokens of context, across four live endpoints in our 10 July 2026 snapshot. That is ample for large-repository coding work and long agent sessions, though endpoint configurations can change over time.

Should I choose Grok 4.5 over Fable 5 or GPT-5.6 Sol?

It depends on your constraint. If cost at volume is the binding constraint (agentic coding, background automation, embedded product features), Grok 4.5's value case is the strongest of any closed model, with an AITR Value For Money Index of 174. If you need the absolute best reasoning for hard problems and cost is secondary, Fable 5 or GPT-5.6 Sol remain the stronger choices.

Specifications

pricing$2/$6 per 1M tokens
context Window500k tokens (observed)

AI Evaluation

4.6
Expert Rating
Text4.6/5
Coding4.5/5

Doesn't move the frontier of what AI can do - it moves the frontier of what frontier-adjacent AI costs. The strongest price-performance of any closed model, and a top shortlist pick for agentic coding at volume.

Pros

  • Cheapest closed frontier model per task ($0.31)
  • Ties GPT-5.6 Sol on SWE-Atlas-QnA
  • Best closed-model score on our Value For Money Index (174)

Cons

  • Fourth on intelligence, not first
  • Coding score measured in its own harness
  • Public for barely a week - short track record