Quick Answer:
Grok 4.6 launched on 12 August 2026 from SpaceXAI (formerly xAI, merged with SpaceX in February 2026). Independent testing from Artificial Analysis puts it at 61 on the Intelligence Index - tying GPT-5.6 Sol, one point behind Claude Fable 5, two behind Claude Opus 5 - a five-point jump over Grok 4.5 in just over a month. Its real edge is efficiency: it resolves tasks in roughly half the turns Opus 5 needs, at $2/$6 per million tokens, around 60% cheaper than GPT-5.6 Sol or Opus 5 at list price. The catch is a 200,000-token pricing cliff that reprices an entire long request once you cross it, and independent testing found a meaningfully higher confident-fabrication rate than its benchmark parity with GPT-5.6 Sol might suggest.
Three weeks ago, this site covered the rumoured 7 August target date for Grok 4.6 and the unconfirmed specs circulating ahead of it. That rumour cycle is over: SpaceXAI shipped the real thing on 12 August 2026, alongside a same-day model card and day-one availability in Cursor. What follows is what actually shipped, sourced to Artificial Analysis's independent evaluation and SpaceXAI's own materials rather than launch-day hype.
The headline numbers are good but not category-leading, and the most interesting part of the release is not the Intelligence Index score - it's the price and the efficiency data sitting underneath it.
WorldofAI's same-day hands-on testing using their own benchmark tool.
Executive Summary
Grok 4.6 is SpaceXAI's fourth major Grok release since the company's February 2026 merger with SpaceX, and its second since the July rebrand from xAI to SpaceXAI. On paper, it lands in a genuinely competitive spot: an Artificial Analysis Intelligence Index score of 61 puts it in a three-way tie for the frontier's upper-middle tier, exactly matching OpenAI's GPT-5.6 Sol and trailing Anthropic's Claude Fable 5 and Claude Opus 5 by one and two points respectively. That is a real, independently-verified result, not a company-reported figure.
What separates Grok 4.6 from a routine point release is less the score itself than what it costs to get there. Artificial Analysis's efficiency data shows Grok 4.6 completing tasks in roughly half the reasoning turns and a quarter of the input tokens that Claude Opus 5 uses to reach a comparable result, at a headline price of $2 input / $6 output per million tokens - about 60% below GPT-5.6 Sol and Opus 5 at list pricing. For teams running high-volume agentic workloads, that combination of "good enough" capability and meaningfully lower cost per completed task is the actual pitch, more than raw intelligence-index bragging rights.
- Best for: cost-sensitive agentic and coding workloads where near-frontier capability at a fraction of the price beats chasing the top score.
- Headline number: Artificial Analysis Intelligence Index 61, tying GPT-5.6 Sol, +5 over Grok 4.5.
- Headline price: $2/$6 per million tokens under 200K-token prompts - roughly 60% cheaper than Opus 5 or GPT-5.6 Sol.
- Main caveats: loses to GPT-5.6 Sol on at least two independent coding benchmarks, a 200K-token pricing cliff that can more than double real costs on long contexts, and a confident-fabrication rate independent testing flags as higher than its benchmark parity implies.
Lineage: From Grok 4.5 to SpaceXAI
Two corporate changes sit behind this release that are easy to miss if you only follow the model numbers. First, xAI merged into SpaceX in an all-stock deal that closed on 2 February 2026, creating a combined entity reported at roughly $1.25 trillion in valuation - one of the largest private mergers on record. Second, on 6 July 2026 the combined company rebranded as SpaceXAI; the official @xai account on X moved to @SpaceXAI, and SpaceXAI's own announcement page for this release is titled, verbatim, "Introducing Grok 4.6 | SpaceXAI." The Grok product name itself is unchanged - it is the parent company's name that shifted, and this article uses "SpaceXAI" throughout for that reason.
On the model side, Grok 4.6 follows Grok 4.5, which opened to the public on 9 July 2026 at an Intelligence Index of 54. The five-week gap between the two releases is fast even by 2026's frontier-lab cadence, and Artificial Analysis frames the jump - Intelligence Index 54 to 61, a five-point gain landing "just over a month" after Grok 4.5 shipped - as one of the quicker capability jumps SpaceXAI has delivered. It also represents a 23-point gain over Grok 4.3, giving a sense of how far the Grok line has moved across 2026 as a whole.
This site's earlier piece, Grok 4.6 Release Date: 7 August Target, Grok 4.7 Next, covered the pre-release rumour cycle when Elon Musk was targeting 7 August. The model actually shipped five days later than that target, on 12 August - a reminder that Musk's own target dates for Grok releases have consistently run optimistic across the 4.x line, and that the real specs only became knowable once SpaceXAI's own materials and Artificial Analysis's independent evaluation landed.
What Actually Shipped: Specs & Model Card
SpaceXAI's own announcement does not disclose a parameter count for Grok 4.6. A figure of 1.5 trillion parameters, reusing the "V9" foundation reported for Grok 4.5, circulates across several secondary blogs, but it traces back to none of SpaceXAI's own materials and should be treated as widely repeated rather than confirmed - this article does not present it as fact.
What is independently confirmed, via Artificial Analysis's own model page, is a 500,000-token context window, unchanged from Grok 4.5, alongside an average output speed of 65.8 tokens per second and a slower-than-average 32.30-second time to first token - Grok 4.6 is not a fast-to-first-response model relative to its peers, even though its total throughput is respectable once generation begins.
SpaceXAI published a Grok 4.6 model card as a PDF dated 12 August 2026 - the same day as the release - hosted at media.x.ai. That timing directly contradicts a press report from Forkast claiming SpaceXAI shipped "for the second consecutive release without a model card." The most likely explanation is a publication-timing gap between when that story was filed and when the card went live, though it is also possible the complaint was narrower - about the depth of autonomous-capability and agentic-safety disclosure specifically, rather than the card's existence outright. Either reading is plausible; what can be stated with confidence is that a same-day-dated model card PDF does exist at SpaceXAI's own domain. Search-result characterisations of its contents describe a "layered, defense-in-depth" safety approach combining supervised fine-tuning, RLHF, verifiable rewards and model-based grading for refusal behaviour, tested under adversarial pressure, and referencing SpaceXAI's Frontier Artificial Intelligence Framework (dated 30 June 2026) - description drawn from search summaries rather than a direct read of the PDF, so treat the specific wording as approximate rather than a verbatim quote.
Benchmarks: What Artificial Analysis Found

The independent scorecard, all figures from Artificial Analysis unless noted otherwise:
- Intelligence Index: 61 - sixth of 184 tracked models, tying GPT-5.6 Sol exactly, one point behind Claude Fable 5 (62) and two behind Claude Opus 5 (63).
- GDPval-AA v2 Elo: 1753 - second only to Claude Opus 5, with overlapping confidence intervals against Claude Fable 5 and Qwen3.8 Max.
- AA-Briefcase Elo: 1577 - what Artificial Analysis characterises as "Fable 5-tier" performance on this metric.
- τ³-Banking: 50.7% - among the top two scores tested, though independent reviewer eesel.ai points out this still means Grok 4.6 fails roughly half of the banking-agent conversations in the suite.
- Terminal-Bench v2.1: 88.4% per Artificial Analysis. A separately-reported figure, Terminal-Bench v3.0 at 26% versus GPT-5.6 Sol's 34.6%, comes from a newer harness version and is not directly comparable to the v2.1 number - the two should not be read as contradicting each other.
- DeepSWE v1.1: 65.9% versus GPT-5.6 Sol's 73% - a 7.1-point gap on a coding-specific benchmark, notable given Grok 4.6's launch marketing leaned heavily on agentic coding strength.
- AA-Omniscience accuracy: 48.2%, with a non-hallucination rate of 65.7% - covered in more detail below.
SpaceXAI's own launch page cites three additional figures - CursorBench 3.2 at 69.9%, FrontierCode 1.1 at 61.3%, and APEX-Agents at 57.5% - all self-reported and not yet independently reproduced by Artificial Analysis or another third party at the time of writing. Treat these as company-reported rather than verified.
Hacker News discussion in the day after launch landed on cautious scepticism rather than enthusiasm: one representative comment put it as "Grok 4.6 is around Opus 4.8 in real world use, but clearly below Opus 5" - a reminder that a tied Intelligence Index score does not automatically translate into tied real-world usefulness, particularly for users who have already built workflows around a specific model's quirks.
The Real Story: Turns and Tokens, Not Just Scores

This is the number that matters most if you are choosing between Grok 4.6 and a pricier frontier model for a production agentic workload. Artificial Analysis measured Grok 4.6 resolving its task set in an average of roughly 53 turns and around 0.5 billion input tokens, against Claude Opus 5 in its "max" configuration needing roughly 103 turns and around 2.0 billion input tokens for comparable coverage. That is close to a 2x reduction in turns and a 4x reduction in input-token volume - a cost and latency advantage that compounds independently of the per-token price difference, since a shorter, more token-efficient task also finishes faster and burns less of an agent's context budget along the way.
Read alongside the Intelligence Index tie with GPT-5.6 Sol, this efficiency data is the more defensible version of SpaceXAI's "cheap and capable" pitch: Grok 4.6 is not claiming to out-think Opus 5, it is claiming to get comparably useful agentic work done in meaningfully fewer steps and tokens, which is a different and in some ways more commercially relevant claim for teams running agents at volume.
Hallucination Rate and Honesty
The most substantive independent critique of Grok 4.6, flagged by reviewer eesel.ai using Artificial Analysis's own AA-Omniscience data, is its honesty profile. Grok 4.6 scores a 65.7% non-hallucination rate, meaning that when the model gets an answer wrong, roughly one in three of those wrong answers is a confidently-stated fabrication rather than an acknowledgement of uncertainty. Combined with a 48.2% AA-Omniscience accuracy score, that paints a model that is reasonably capable but does not reliably know when it doesn't know something - a meaningfully different failure mode than a model that hedges or declines to answer under uncertainty.
For any workload where a wrong-but-confident answer is worse than a flagged "I'm not sure" - customer-facing agents, anything touching financial or medical information, unsupervised multi-step automation - this is the single figure from Grok 4.6's benchmark set worth weighing most heavily against the otherwise strong efficiency and price story above.
Alex Finn's same-day take on Grok 4.6's price-to-capability ratio against Claude Fable 5.
Pricing and the 200K-Token Cliff
Standard Grok 4.6 pricing, for prompts under 200,000 tokens, is $2 per million input tokens and $6 per million output tokens, with cached input priced separately at roughly $0.50 per million tokens below that threshold. That is the figure Techmeme's launch-day aggregation and Artificial Analysis both confirm, and it is the basis for SpaceXAI's headline claim - repeated across its own announcement, Artificial Analysis and press coverage from VentureBeat and the-decoder - that Grok 4.6 runs "60%+ cheaper" than Claude Opus 5 or GPT-5.6 Sol at list price.
| Model | Input $/1M | Output $/1M |
|---|---|---|
| Grok 4.6 (<200K tokens) | $2.00 | $6.00 |
| Grok 4.6 (≥200K tokens, whole request) | $4.00 | $12.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| Claude Fable 5 | $10.00 | $50.00 |
| Kimi K3 | $2.80 | $14.00 |
| DeepSeek V4 Pro | $0.44 | $0.87 |
The line worth reading carefully is the second one. Independent analysis from rohitai.com describes a pricing mechanic that SpaceXAI's own pricing page does not spell out prominently: once a request crosses the 200,000-token threshold, the entire request - not just the tokens over the line - is billed at the higher $4/$12 rate. For a long-context agentic or coding session, exactly the kind of workload Grok 4.6 is being marketed for, that can more than double the effective cost of a request that only slightly overruns the threshold. This detail comes from a single secondary source rather than direct confirmation on SpaceXAI's own pricing page, so treat it as a flagged risk to verify against your own usage pattern rather than a fully SpaceXAI-confirmed mechanic - but it is exactly the kind of gotcha worth checking before committing a long-context production workload to the standard tier.
Real-World Reception vs the Announcement

Coverage in the hours after the 12 August launch converged on a fairly consistent read: a genuinely strong, well-priced release rather than a frontier-leading one. VentureBeat's headline framed it as SpaceXAI "overtaking Kimi K3's performance and matching GPT-5.6 Sol," which is accurate to the Intelligence Index data but notably does not claim it beats Claude's top tier. The-decoder's coverage led with the price-versus-OpenAI angle specifically. Forkast's piece, despite the model-card timing discrepancy noted above, raised a fair point about documentation depth on autonomous capability - a recurring theme across SpaceXAI's last two Grok releases, according to that outlet.
Creator coverage split along predictable lines: WorldofAI's same-day testing, using their own benchmarking tool, found Grok 4.6 beating GPT-5.6 Sol, Opus 5 and Kimi K3 - a result stronger than Artificial Analysis's own Intelligence Index data supports for the Opus 5 comparison specifically, and worth treating as one creator's own test methodology rather than a third independent confirmation. Alex Finn's framing, "Grok 4.6 is Claude Fable 5, but dirt cheap," is closer to right on price - Grok 4.6's $2/$6 pricing is roughly five to eight times cheaper than Fable 5's $10/$50 - but overstates capability parity, since Grok 4.6's 61 sits one point below Fable 5's 62 and loses outright on some coding-specific benchmarks.
Limitations
- No official parameter count: the widely-repeated "1.5 trillion parameters" figure is unconfirmed by SpaceXAI's own materials.
- Loses on some coding benchmarks: trails GPT-5.6 Sol by 7.1 points on DeepSWE v1.1 and by a wider margin on the newer Terminal-Bench v3.0 harness, despite an agentic-coding-forward launch narrative.
- Elevated confident-fabrication rate: a 65.7% non-hallucination rate means roughly a third of wrong answers are stated with unwarranted confidence rather than flagged as uncertain.
- 200K-token pricing cliff: a single secondary source describes whole-request repricing past 200,000 tokens that can more than double effective cost on long-context work - verify against your own usage before committing a production workload.
- Slow time to first token: 32.30 seconds average per Artificial Analysis, meaningfully slower to start responding than several peers, even though total throughput is reasonable once generation begins.
- Documentation-depth question mark: at least one outlet argues SpaceXAI's disclosure on autonomous-capability behaviour remains thin relative to peers, independent of whether a model card PDF technically exists.
How It Compares
Against GPT-5.6 Sol, Grok 4.6 ties on the composite Intelligence Index but loses on coding-specific benchmarks and costs meaningfully less - $2/$6 versus $5/$30 per million tokens. Against Claude Opus 5, Grok 4.6 trails by two Intelligence Index points but resolves comparable agentic tasks in roughly a quarter of the input tokens and around half the turns, at 60% lower list pricing - a genuinely different value proposition rather than a straight capability comparison. Against Claude Fable 5, the gap is one Intelligence Index point in Fable's favour at five to eight times the price, which makes Grok 4.6 the harder case to argue against on cost grounds alone. Kimi K3 remains the closest price-to-capability peer in the open-weights-adjacent tier, though Grok 4.6's proprietary, closed deployment model is a different trade-off than Kimi's more open ecosystem.
None of this happens in a vacuum: DeepSeek V4 continues to undercut all of the above on raw per-token price at the cost of a meaningfully lower Intelligence Index score, and the frontier as a whole has compressed enough in 2026 that "which model is smartest" now matters less for most production decisions than "which model finishes the most tasks per dollar without an unacceptable error rate" - the exact trade-off Grok 4.6's efficiency and hallucination data force into the open.
Who Should Use It
Choose Grok 4.6 if you run high-volume agentic or coding workloads where cost per completed task matters more than chasing the single highest benchmark score, you can tolerate a slower time-to-first-token, and you are willing to monitor for the 200K-token pricing cliff on long-context jobs.
Look elsewhere if your workload is sensitive to confident-but-wrong answers (customer-facing, financial, medical, or otherwise high-stakes contexts), you specifically need frontier-leading coding benchmark performance, or you would rather pay more for Claude Opus 5's stronger overall Intelligence Index and lower demonstrated hallucination profile.
The Bottom Line
Grok 4.6 is a genuinely strong, well-priced release rather than a frontier-leading one, and the independent data backs both halves of that sentence. Tying GPT-5.6 Sol and trailing Claude's top two models by one and two Intelligence Index points is a solid result five weeks after Grok 4.5, and the efficiency data - roughly half the turns and a quarter of the input tokens Opus 5 needs for comparable agentic work - is the more commercially meaningful story than the headline score itself.
Set against that: a real coding-benchmark gap versus GPT-5.6 Sol, a hallucination profile independent testing flags as a genuine concern, and a pricing structure with a cliff that is easy to miss until a long-context bill arrives. For cost-sensitive, high-volume agentic work, Grok 4.6 is one of the better value picks at the frontier right now. For anything where a confidently wrong answer is expensive, the extra cost of Claude Opus 5 or GPT-5.6 Sol is still the safer default.
Last updated: 13 August 2026. Sourced from SpaceXAI's official Grok 4.6 announcement (x.ai/news/grok-4-6) and model card, Artificial Analysis's independent benchmark data and analysis, and launch-day reporting from VentureBeat, the-decoder and eesel.ai.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.







