AI Tools Review
Mistral Large 4 (Le Chonk): Specs, Benchmarks, Pricing

Insights

Mistral Large 4 (Le Chonk): Specs, Benchmarks, Pricing

AI Tools Review Editorial Team7 October 2026

    Quick Answer:

    Mistral Large 4 (nicknamed Le Chonk) entered public preview on 06/10/2026. It is a natively multimodal, hybrid reasoning mixture-of-experts model with about 1.05 trillion parameters (49B to 52B active) and a vendor-stated 1M-token context window. List price is $1.36 in / $4.18 out per million tokens, halved for the first two weeks. Open weights are promised by end of October 2026; the licence is not yet published. It leads open models on cybersecurity but is not the top coding model.

    For most of 2026 the conversation about frontier open models has been about China. Kimi, GLM, Qwen and DeepSeek set the pace, whilst Europe's flagship lab, Mistral, was best known for a mid-tier model that struggled on independent leaderboards. Mistral Large 4 is the company's answer, and Matthew Berman titled his coverage, with only slight exaggeration, "Mistral is BACK!".

    This guide separates what Mistral has actually released as of 07/10/2026 from what is still a promise. We read Mistral's announcement, its model documentation, the Hugging Face placeholder page and two independent trackers, downloaded the official charts, and flag every conflict in the numbers. We have not run the model ourselves. Two creator videos that prompted this piece are embedded below.

    Matthew Berman's coverage of the Mistral Large 4 announcement, which links straight to Mistral's launch post.

    Executive Summary

    What it is: Mistral Large 4 is a very large mixture-of-experts (MoE) model that combines instruction following and reasoning in one network and accepts text and images. Mistral says it was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacentres, and that it speaks more than 160 languages, including every official language of the European Union.

    • Status: public preview since 06/10/2026, available through Mistral Studio and the API under the model ID mistral-large-4.
    • Size: roughly 1.05 trillion total parameters, 49 billion active in the experts (52 billion with embeddings and output layers) and a 1.6 billion-parameter vision encoder.
    • Context: 1M tokens according to Mistral's docs; third parties report 512K to 524K. See the dedicated section below.
    • Price: $1.36 per million input tokens and $4.18 per million output tokens at list, with a 50% launch discount for two weeks.
    • Open weights: promised by the end of October 2026; not yet available and licence terms not yet published.
    • Strongest claim: state of the art among open models for cybersecurity, finance and law, according to Mistral, backed in part by Artificial Analysis data.
    • Biggest caveat: on coding agents it trails Kimi K3 and GLM-5.3 in Mistral's own charts, and independent aggregate scores place it mid-table.

    Our view: this is a serious and genuinely useful release, particularly for European organisations that care about data residency and sovereign infrastructure. It is not, on the published evidence, the best open model overall, and until the weights and licence arrive, "open" remains a promise rather than a fact. For the models it is competing with, see our reviews of Kimi K3, GLM 5.3 and DeepSeek V4 Pro.

    What Has Actually Been Released (as of 07/10/2026)

    With a launch this fresh, it is worth being explicit about the state of play. The table separates confirmed facts from promises, and names where each came from.

    ItemStatus on 07/10/2026Source
    Public preview via API / Mistral StudioLive since 06/10/2026Mistral docs changelog and model card
    Model IDmistral-large-4Mistral docs
    ArchitectureGranular MoE, about 1.05T total, 49B active (52B with embeddings)Mistral announcement, model card
    ModalitiesText and image input, text outputMistral docs, Artificial Analysis
    Context window1M tokens (vendor); 512K to 524K (independent trackers)Mistral docs versus Vals AI and Artificial Analysis
    List price$1.36 input / $4.18 output per million tokensMistral announcement, Artificial Analysis, Vals AI
    Launch price$0.68 input / $0.07 cached / $2.09 output, for two weeksMistral docs
    Open weightsPromised by end of October 2026; expected 31/10/2026 on Hugging FaceMistral announcement, Hugging Face placeholder
    LicenceNot published ("coming soon")Hugging Face placeholder, third-party summaries
    Safety workRed-teaming with cybersecurity leaders, vetted partners and state authorities, ongoingMistral announcement

    Two details matter. First, the model is in preview, so behaviour, limits and pricing may still change before general availability. Second, Artificial Analysis currently lists the model as proprietary simply because the weights are not out. That label should flip if Mistral delivers, and we will update this page when it does.

    Lineage: How Mistral Got Here

    Mistral has spent the past two years building a family that now spans several tiers. The earlier Large releases were competent but were overtaken by rivals on coding and agentic work. The mid-tier Mistral Medium 3.5 was criticised on independent leaderboards for scoring at the level of much cheaper models. Against that background, Large 4 is a statement: a trillion-parameter flagship, a reasoning-first design and a commitment to open weights.

    The most useful sibling to understand is Mistral Small 4, which shipped earlier in the year (the changelog model ID, mistral-small-2603, points to March 2026, and coverage dates its release to 16/03/2026; the changelog page itself carries a typo in the year). It is a hybrid model that unifies instruction following, reasoning and coding in one multimodal model, replacing the need to choose between Mistral's separate reasoning (Magistral), vision (Pixtral) and agentic coding (Devstral) lines. Reported specifications are a 119B-parameter MoE with 128 experts and around 6B active parameters, a 256K context window and an Apache 2.0 licence, with configurable reasoning effort. Third-party coverage prices API access at around $0.15 per million input tokens and notes it can be self-hosted from Hugging Face or run through Ollama.

    Read together, the two releases show a strategy: Small 4 is the open, cheap, easy-to-host workhorse, whilst Large 4 is the high-end model for demanding agentic, security and enterprise work. The 256K figure in early leaks seems to have belonged to the Small 4 generation, which may explain some confusion about Large 4's context window.

    Architecture and Training

    Mistral calls Large 4 a "granular" mixture-of-experts design. In an MoE model a router sends each token to a few of many specialist sub-networks, so the total parameter count governs memory needs whilst the active count governs compute per token. With roughly 1.05 trillion parameters and 49 billion active, Large 4 stores a great deal of knowledge but pays the compute cost of a model closer to 50 billion parameters on each step. That is the same recipe behind the largest Chinese open models; for comparison, see our coverage of Qwen 3.8 Max.

    Several other specifics come from the announcement:

    • Natively multimodal: a 1.6B-parameter vision encoder is built in, rather than bolted on, and the model reads images, charts and documents. Mistral reports visual-grounding results and evaluates on ChartQA Pro and a document-understanding benchmark.
    • Hybrid instruct and reasoning: one model handles quick replies and extended thinking, rather than forcing a choice between two products.
    • Training compute: 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacentres, an important part of its sovereignty message.
    • Reinforcement learning at scale: Mistral says its post-training pipeline generates roughly 33 billion tokens per day, of which about 16 billion are trainable completion tokens after filtering.
    • Languages: more than 160, including all official EU languages, a genuine differentiator for European public-sector and enterprise buyers.

    What Mistral has not published, at least yet, is a full technical report or system card with training-data details, safety evaluations in a standard format or a licence. The announcement is a product post with charts, not a research paper, and that shapes how much weight the numbers can bear.

    The Context Window Question

    The headline specification you see depends on where you look, so it deserves its own section.

    SourceContext figureNote
    Pre-launch leak (X, June 2026)256K, vision supportedUnverified; described as an upcoming release
    Mistral model card and changelog1M tokensVendor claim; primary source
    Artificial Analysisabout 524K tokensIndependent tracker
    Vals AI512K input, 256K maximum outputIndependent tracker

    Our reading is that Mistral supports 1M tokens at the API level whilst independent evaluators tested, or capped at, a lower figure; we cannot confirm which. It is also common for long-context quality to degrade well before the nominal limit, so a claim of 1M is not a promise of reliable recall at 1M. If your use case depends on very long inputs, run a needle-in-a-haystack style test with your own documents, and watch cost, since a million-token prompt at $1.36 per million input tokens is not cheap.

    Benchmarks: What the Evidence Shows

    Almost every number below is Mistral-reported, using independent evaluators such as Artificial Analysis where noted. Vendor charts choose their own comparison sets, so read them as favourable framing and look for the places where the model does not win. We found several.

    Cybersecurity: the real strength

    This is where Mistral makes its boldest claim, and where the evidence is strongest. On the Artificial Analysis Cyber Index, Mistral Large 4 Preview records 50 successes. Mistral also reports a top-five global placement, leadership among open-weight models, 82% on a vulnerability-patching test (the highest of any model it lists) and 93% of challenges solved on Cybench.

    Bar chart titled Artificial Analysis Cyber Index showing successes: Mistral Large 4 Preview 50, Qwen3.8 2.4T A95B 13, Opus 5.5 29, GPT-6 Astra 33, GLM-5.3 36, Kimi K3 41, DeepSeek-V4.1-Flash 41, with grey safety-block bars of 63 for Qwen, 36 for Opus 5.5 and 38 for GPT-6 Astra.
    Artificial Analysis Cyber Index (successes), as published by Mistral. Grey bars show safety blocks, where a model refused. Source: Mistral AI / Artificial Analysis.

    The chart carries an important subtlety. The grey bars show "safety blocks", cases where a model refused the task. Qwen3.8 shows 63, Opus 5.5 shows 36 and GPT-6 Astra shows 38, which means part of the gap between those models and Mistral reflects refusals rather than inability. Mistral also states that its cyber refusal rate is higher than that of all open-source competitors, so the model is not simply unguarded. How a vendor should balance helpfulness for defenders against misuse is a live debate; our GLM 5.3 review covers the same tension and the WorldofAI video includes two cybersecurity demos of the model.

    Coding and agents: competitive, not leading

    On software engineering Mistral reports 61.7% on DeepSWE 1.1, 28.3% on Terminal-Bench 4.0, 59.4% on SWE-Atlas-QnA and 49.8% on the Artificial Analysis Coding Agent Index. In a human evaluation run by Surge AI, it ranked second of five models with a score of 3.74 out of 5.

    Bar chart titled Artificial Analysis DeepSWE 1.1: Mistral Large 4 Preview 62, Beam self-reported 44, Qwen3.8 Max with Claude Code 51, DeepSeek-V4-Pro-0813 with Codex 57, GLM-5.3 with Opencode 61, Kimi K3 with Kimi Code CLI 68.
    Artificial Analysis DeepSWE 1.1 as published by Mistral. The footnote says scores were evaluated privately by Artificial Analysis ahead of the harness's public launch. Source: Mistral AI / Artificial Analysis.

    The chart is honest about a weakness: Kimi K3 with its own CLI scores 68, ahead of Mistral's 62, and GLM-5.3 is close behind at 61. Note also that the competitors are run with different agent harnesses (Claude Code, Codex, Opencode and Kimi Code CLI), so the comparison mixes model and harness quality.

    Bar chart titled Code Benchmarks Terminal-Bench 4.0: Mistral Large 4 Preview 28, DeepSeek V4 Pro 0813 10, Qwen3.8 Max 17, Kimi K3 21, GLM-5.3 40.
    Terminal-Bench 4.0 as published by Mistral. GLM-5.3 leads at 40; Mistral Large 4 Preview scores 28. Source: Mistral AI / Artificial Analysis.

    Terminal-Bench is blunter still. GLM-5.3 scores 40 against Mistral's 28, though Mistral leads Kimi K3 (21), Qwen3.8 Max (17) and DeepSeek V4 Pro (10) on this chart. In short, Large 4 sits in the upper half of the open-model pack for coding agents but does not take the crown. For a dedicated look at those rivals see Kimi K3 and DeepSeek V4 Pro.

    Agentic workflows and enterprise tasks

    Mistral reports 59.9% on AutomationBench and an Elo of 1,393 on AA-Briefcase, an Artificial Analysis evaluation of office-style knowledge work. In a human comparison against GLM-5.3 across coding, CAD, finance, mathematics and physics, the company says reviewers preferred Large 4 in the CAD and STEM categories, which also implies GLM-5.3 won elsewhere. It further claims state-of-the-art open-weight results on SciCode-Verified. We could not independently reproduce any of these.

    Independent aggregate scores

    The most sobering data comes from third-party trackers, which test with their own harnesses:

    TrackerResultWhat it implies
    Artificial Analysis Intelligence Index38; ranked 64th of 225 modelsMid-table overall, well behind frontier closed models
    Artificial Analysis speedabout 116 output tokens per second; 1.46 s to first tokenReasonable for a model of this size
    Vals Index48.05% (plus or minus 1.11); 32nd of 44 modelsMid-table
    Vals: individual benchmarks80.43% on MedScribe, 78.40% on Vibe Code Bench v1.1, 10.00% on ProofBenchVery uneven by task

    These figures say that Large 4 is a good, not dominant, model on general-purpose evaluations, with unusual strengths (cybersecurity, some legal and finance tasks, where Vals ranks it sixth of 75 on one legal-agent benchmark) and clear weaknesses such as formal proofs. That profile fits Mistral's own message: it is positioning for specific enterprise workloads rather than claiming to beat every frontier model at everything.

    Safety, Security and Red-Teaming

    Mistral says it is still red-teaming Large 4 in real-world settings with cybersecurity leaders, vetted partners and state authorities, which is notable for a model whose headline strength is offensive and defensive security capability. Alongside the cyber results it reports a 93.3% attack-resistance rate on the B3 agent-security benchmark and a score of 1.691 out of 2.0 on the KORA benchmark.

    Bar chart titled B3 Agent Security Benchmark, attack resistance: Mistral Large 4 Preview 93.3, GLM-5.3 93.3, GLM-5.2 88.1, Kimi K3 88.1, DeepSeek-V4-Pro-0813 85.2, Kimi K2.6 83.8. The vertical axis starts at 80.
    B3 Agent Security Benchmark, attack resistance, as published by Mistral. The vertical axis starts at 80, which exaggerates visual gaps. Source: Mistral AI.

    Note that Mistral Large 4 ties GLM-5.3 at 93.3 rather than beating it, and the truncated axis (it starts at 80) makes the bars look more different than the numbers are. B3 measures resistance to attacks such as prompt injection on agents, an increasingly important property as models are given tools. There is no published system card with the usual categories of dangerous-capability testing (biological, chemical, autonomy), so we cannot compare Large 4's risk posture with the detailed disclosures from the likes of Anthropic. If Mistral publishes a full report alongside the open weights, that will be the moment to reassess. The wider debate about releasing capable cyber models openly is covered in our piece on open weights and the Anthropic response.

    WorldofAI's news round-up. The Mistral Large 4 segment starts around 13:00 and covers cybersecurity demos, DeepSWE 1.1 and pricing.

    Pricing and Availability

    Mistral's announcement, Artificial Analysis and Vals AI agree on a list price of $1.36 per million input tokens and $4.18 per million output tokens, roughly £1.00 and £3.10 at current exchange rates. Mistral's documentation shows a launch discount: 50% off for two weeks, giving $0.68 input, $0.07 cached input and $2.09 output per million tokens (about £0.50, £0.05 and £1.55). One analysis puts the end of the discount at 20/10/2026, though we could only confirm "two weeks" from the docs.

    Available features per the model card include structured outputs, function calling, document question-answering, chat completions, batch processing and the agents and conversations APIs with built-in tools. OpenRouter also carries a listing for the model. Vals AI reports an average cost of $13.78 to run its index, a reminder that reasoning models consume many more tokens than their rate card suggests, so measure cost per completed task rather than cost per token.

    Context for the price: it is above the lower tiers of Chinese open models but far below the premium closed frontier. For where Mistral's mid-tier sits in our tools directory see Mistral Medium 3.5.

    The Sovereignty and Open-Weights Angle

    Mistral frames Large 4 as part of a bid for sovereign, open AI at the frontier. It says the model is served from European infrastructure, independently of other digital service providers and under European law, and will be available across multiple regions worldwide. For governments, regulated industries and any firm uncomfortable with sending data to a US or Chinese provider, that is a real selling point that benchmark tables do not capture.

    The open-weights promise is just as important, and just as unproven. A Hugging Face repository named Mistral-Large-4.0-1T05-A52B exists as an "upcoming release" page with a countdown and, at the time of writing, hundreds of people waiting; it contains no licence, benchmark or serving information. Mistral has a good record on open licences (Small 4 is Apache 2.0), but an open-weight trillion-parameter model is a different proposition: you will need a multi-GPU node even with aggressive quantisation, and the licence could range from fully permissive to a restricted community licence. Until it appears, plan on API use only.

    How It Compares

    ModelWhere it is aheadWhere Mistral Large 4 is ahead or level
    Kimi K3DeepSWE 1.1 (68 versus 62)Terminal-Bench 4.0 (28 versus 21); Cyber Index (50 versus 41)
    GLM-5.3Terminal-Bench 4.0 (40 versus 28)Cyber Index (50 versus 36); DeepSWE is effectively a tie (62 versus 61); B3 security is level at 93.3
    DeepSeek V4 Pro 0813Not shown ahead on any chart hereDeepSWE (62 versus 57), Terminal-Bench (28 versus 10), B3 (93.3 versus 85.2)
    Qwen3.8 MaxNot shown ahead on any chart hereDeepSWE (62 versus 51), Terminal-Bench (28 versus 17)
    Closed frontier models (Opus 5.5, GPT-6 Astra)Aggregate indices, generallyCyber Index successes (50 versus 29 and 33), partly due to refusals

    These comparisons are drawn solely from Mistral's charts and are only as good as its selection. For deeper analysis of each rival, see GLM 5.3, Kimi K3, DeepSeek V4 Pro and Qwen 3.8 Max.

    Real-World Use Versus Benchmarks

    Leaderboards measure what is easy to measure. The reason Large 4 may matter more than its index scores suggest is workload fit. A model that does well at vulnerability patching, finance and legal agent tasks, reads documents with a native vision encoder and can be hosted in the EU suits a specific buyer: a bank, law firm, public body or security team that values locality and auditability over a few points on a coding index.

    Conversely, if your main workload is autonomous coding, the published evidence points to Kimi K3 or GLM-5.3 first, and to testing Large 4 as a second option. In both cases the cost picture should be measured on your own tasks: reasoning models can spend large numbers of tokens, and the launch discount will distort early cost comparisons.

    A sensible evaluation plan has four steps. First, build a set of 30 to 50 real tasks from your workflow. Second, run Large 4 alongside your current model and one cheaper option, recording success, tokens and minutes per task. Third, test long inputs at the sizes you actually use rather than the advertised maximum. Fourth, run the set twice, since reasoning models vary between runs. Treat the result as a routing table rather than a verdict.

    Limitations and Open Questions

    • Preview status: pricing, limits and behaviour may change before general availability.
    • Weights and licence unconfirmed: the end-of-October date is a promise, and the terms are unknown.
    • Conflicting context figures: 1M, 524K and 512K all appear.
    • Vendor-chosen comparisons: every chart was produced by Mistral and shows models it selected, some with different harnesses.
    • No system card: there is no standard safety report covering biological, chemical or autonomy risks.
    • No hands-on testing: we have not run the model, so we cannot say how it feels in daily use.
    • Refusal effects: parts of the cyber lead come from rivals refusing tasks.

    Who Should Use It

    • European enterprises and public bodies that need EU-hosted, sovereign infrastructure and multilingual coverage.
    • Security teams evaluating defensive automation, with proper governance around a model that is strong at offensive tasks.
    • Finance and legal teams experimenting with document-heavy agent workflows, subject to their own validation.
    • Self-hosters should wait for the weights and licence before planning anything.
    • Not ideal for: developers who only want the best autonomous coding model today, or anyone who needs a published safety system card.

    The Bottom Line

    Mistral Large 4 is a credible comeback rather than a total reversal. It is a huge multimodal reasoning model with strong, independently supported cyber results, competitive coding scores, EU hosting and a promise of open weights. It is also mid-table on independent aggregate indices, behind Kimi K3 and GLM-5.3 on coding agents, still in preview and not yet open in practice.

    Our advice is to try the preview whilst the launch discount lasts, benchmark it on your own tasks and wait for 31/10/2026 before making architectural decisions that depend on the weights. We will update this article when the weights, licence and any system card appear.

    Sources

    Charts: Mistral AI and Artificial Analysis, as published in the Mistral announcement. Hero: Mistral AI.

    Last updated: 07/10/2026. Sourced from Mistral's announcement and documentation, the Hugging Face placeholder page and independent trackers. We have not benchmarked Mistral Large 4 ourselves; the model is in preview and prices, limits and open-weights timing may change.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    What is Mistral Large 4?
    Mistral Large 4 (model ID mistral-large-4, nicknamed Le Chonk) is Mistral AI's largest model to date: a natively multimodal, hybrid instruct-and-reasoning mixture-of-experts model with about 1.05 trillion total parameters and 49 billion active per token (52 billion including embeddings and output layers), plus a 1.6 billion-parameter vision encoder. It entered public preview on 06/10/2026 through Mistral Studio and the API.
    What is the context window of Mistral Large 4?
    Mistral's own documentation lists a 1 million-token context window. Pre-launch leaks spoke of 256K, and independent trackers report different figures: Artificial Analysis lists roughly 524K tokens, while Vals AI lists 512K with a 256K output cap. Treat 1M as the vendor claim and test long-context behaviour yourself before relying on it.
    How much does Mistral Large 4 cost?
    Mistral's announcement and independent trackers give a list price of $1.36 per million input tokens and $4.18 per million output tokens (roughly £1.00 and £3.10). Mistral's docs show a launch discount of 50% for two weeks, at $0.68 input, $0.07 cached input and $2.09 output per million tokens. Check the live pricing page, because preview prices can change.
    Is Mistral Large 4 open weight and what is the licence?
    Mistral describes it as an open-weight model and says the weights will be released by the end of October 2026; the Hugging Face placeholder page (Mistral-Large-4.0-1T05-A52B) shows an expected date of 31/10/2026. As of 07/10/2026 the weights are not downloadable and the licence terms have not been published, so you cannot yet self-host it or confirm commercial-use terms.
    Is Mistral Large 4 better than Kimi K3, GLM-5.3 or DeepSeek V4 Pro?
    Not across the board, on the evidence published so far. Mistral's own charts show it ahead on the Artificial Analysis Cyber Index and level with GLM-5.3 on the B3 agent-security benchmark, but behind Kimi K3 on DeepSWE 1.1 (62 against 68) and behind GLM-5.3 on Terminal-Bench 4.0 (28 against 40). Independent aggregate scores are mid-table, so the strongest case for it is cybersecurity, European hosting and the coming open weights.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.