Quick Answer:
Mistral Large 4 (nicknamed Le Chonk) entered public preview on 06/10/2026. It is a natively multimodal, hybrid reasoning mixture-of-experts model with about 1.05 trillion parameters (49B to 52B active) and a vendor-stated 1M-token context window. List price is $1.36 in / $4.18 out per million tokens, halved for the first two weeks. Open weights are promised by end of October 2026; the licence is not yet published. It leads open models on cybersecurity but is not the top coding model.
For most of 2026 the conversation about frontier open models has been about China. Kimi, GLM, Qwen and DeepSeek set the pace, whilst Europe's flagship lab, Mistral, was best known for a mid-tier model that struggled on independent leaderboards. Mistral Large 4 is the company's answer, and Matthew Berman titled his coverage, with only slight exaggeration, "Mistral is BACK!".
This guide separates what Mistral has actually released as of 07/10/2026 from what is still a promise. We read Mistral's announcement, its model documentation, the Hugging Face placeholder page and two independent trackers, downloaded the official charts, and flag every conflict in the numbers. We have not run the model ourselves. Two creator videos that prompted this piece are embedded below.
Matthew Berman's coverage of the Mistral Large 4 announcement, which links straight to Mistral's launch post.
Executive Summary
What it is: Mistral Large 4 is a very large mixture-of-experts (MoE) model that combines instruction following and reasoning in one network and accepts text and images. Mistral says it was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacentres, and that it speaks more than 160 languages, including every official language of the European Union.
- Status: public preview since 06/10/2026, available through Mistral Studio and the API under the model ID
mistral-large-4. - Size: roughly 1.05 trillion total parameters, 49 billion active in the experts (52 billion with embeddings and output layers) and a 1.6 billion-parameter vision encoder.
- Context: 1M tokens according to Mistral's docs; third parties report 512K to 524K. See the dedicated section below.
- Price: $1.36 per million input tokens and $4.18 per million output tokens at list, with a 50% launch discount for two weeks.
- Open weights: promised by the end of October 2026; not yet available and licence terms not yet published.
- Strongest claim: state of the art among open models for cybersecurity, finance and law, according to Mistral, backed in part by Artificial Analysis data.
- Biggest caveat: on coding agents it trails Kimi K3 and GLM-5.3 in Mistral's own charts, and independent aggregate scores place it mid-table.
Our view: this is a serious and genuinely useful release, particularly for European organisations that care about data residency and sovereign infrastructure. It is not, on the published evidence, the best open model overall, and until the weights and licence arrive, "open" remains a promise rather than a fact. For the models it is competing with, see our reviews of Kimi K3, GLM 5.3 and DeepSeek V4 Pro.
What Has Actually Been Released (as of 07/10/2026)
With a launch this fresh, it is worth being explicit about the state of play. The table separates confirmed facts from promises, and names where each came from.
| Item | Status on 07/10/2026 | Source |
|---|---|---|
| Public preview via API / Mistral Studio | Live since 06/10/2026 | Mistral docs changelog and model card |
| Model ID | mistral-large-4 | Mistral docs |
| Architecture | Granular MoE, about 1.05T total, 49B active (52B with embeddings) | Mistral announcement, model card |
| Modalities | Text and image input, text output | Mistral docs, Artificial Analysis |
| Context window | 1M tokens (vendor); 512K to 524K (independent trackers) | Mistral docs versus Vals AI and Artificial Analysis |
| List price | $1.36 input / $4.18 output per million tokens | Mistral announcement, Artificial Analysis, Vals AI |
| Launch price | $0.68 input / $0.07 cached / $2.09 output, for two weeks | Mistral docs |
| Open weights | Promised by end of October 2026; expected 31/10/2026 on Hugging Face | Mistral announcement, Hugging Face placeholder |
| Licence | Not published ("coming soon") | Hugging Face placeholder, third-party summaries |
| Safety work | Red-teaming with cybersecurity leaders, vetted partners and state authorities, ongoing | Mistral announcement |
Two details matter. First, the model is in preview, so behaviour, limits and pricing may still change before general availability. Second, Artificial Analysis currently lists the model as proprietary simply because the weights are not out. That label should flip if Mistral delivers, and we will update this page when it does.
Lineage: How Mistral Got Here
Mistral has spent the past two years building a family that now spans several tiers. The earlier Large releases were competent but were overtaken by rivals on coding and agentic work. The mid-tier Mistral Medium 3.5 was criticised on independent leaderboards for scoring at the level of much cheaper models. Against that background, Large 4 is a statement: a trillion-parameter flagship, a reasoning-first design and a commitment to open weights.
The most useful sibling to understand is Mistral Small 4, which shipped earlier in the year (the changelog model ID, mistral-small-2603, points to March 2026, and coverage dates its release to 16/03/2026; the changelog page itself carries a typo in the year). It is a hybrid model that unifies instruction following, reasoning and coding in one multimodal model, replacing the need to choose between Mistral's separate reasoning (Magistral), vision (Pixtral) and agentic coding (Devstral) lines. Reported specifications are a 119B-parameter MoE with 128 experts and around 6B active parameters, a 256K context window and an Apache 2.0 licence, with configurable reasoning effort. Third-party coverage prices API access at around $0.15 per million input tokens and notes it can be self-hosted from Hugging Face or run through Ollama.
Read together, the two releases show a strategy: Small 4 is the open, cheap, easy-to-host workhorse, whilst Large 4 is the high-end model for demanding agentic, security and enterprise work. The 256K figure in early leaks seems to have belonged to the Small 4 generation, which may explain some confusion about Large 4's context window.
Architecture and Training
Mistral calls Large 4 a "granular" mixture-of-experts design. In an MoE model a router sends each token to a few of many specialist sub-networks, so the total parameter count governs memory needs whilst the active count governs compute per token. With roughly 1.05 trillion parameters and 49 billion active, Large 4 stores a great deal of knowledge but pays the compute cost of a model closer to 50 billion parameters on each step. That is the same recipe behind the largest Chinese open models; for comparison, see our coverage of Qwen 3.8 Max.
Several other specifics come from the announcement:
- Natively multimodal: a 1.6B-parameter vision encoder is built in, rather than bolted on, and the model reads images, charts and documents. Mistral reports visual-grounding results and evaluates on ChartQA Pro and a document-understanding benchmark.
- Hybrid instruct and reasoning: one model handles quick replies and extended thinking, rather than forcing a choice between two products.
- Training compute: 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacentres, an important part of its sovereignty message.
- Reinforcement learning at scale: Mistral says its post-training pipeline generates roughly 33 billion tokens per day, of which about 16 billion are trainable completion tokens after filtering.
- Languages: more than 160, including all official EU languages, a genuine differentiator for European public-sector and enterprise buyers.
What Mistral has not published, at least yet, is a full technical report or system card with training-data details, safety evaluations in a standard format or a licence. The announcement is a product post with charts, not a research paper, and that shapes how much weight the numbers can bear.
The Context Window Question
The headline specification you see depends on where you look, so it deserves its own section.
| Source | Context figure | Note |
|---|---|---|
| Pre-launch leak (X, June 2026) | 256K, vision supported | Unverified; described as an upcoming release |
| Mistral model card and changelog | 1M tokens | Vendor claim; primary source |
| Artificial Analysis | about 524K tokens | Independent tracker |
| Vals AI | 512K input, 256K maximum output | Independent tracker |
Our reading is that Mistral supports 1M tokens at the API level whilst independent evaluators tested, or capped at, a lower figure; we cannot confirm which. It is also common for long-context quality to degrade well before the nominal limit, so a claim of 1M is not a promise of reliable recall at 1M. If your use case depends on very long inputs, run a needle-in-a-haystack style test with your own documents, and watch cost, since a million-token prompt at $1.36 per million input tokens is not cheap.
Benchmarks: What the Evidence Shows
Almost every number below is Mistral-reported, using independent evaluators such as Artificial Analysis where noted. Vendor charts choose their own comparison sets, so read them as favourable framing and look for the places where the model does not win. We found several.
Cybersecurity: the real strength
This is where Mistral makes its boldest claim, and where the evidence is strongest. On the Artificial Analysis Cyber Index, Mistral Large 4 Preview records 50 successes. Mistral also reports a top-five global placement, leadership among open-weight models, 82% on a vulnerability-patching test (the highest of any model it lists) and 93% of challenges solved on Cybench.

The chart carries an important subtlety. The grey bars show "safety blocks", cases where a model refused the task. Qwen3.8 shows 63, Opus 5.5 shows 36 and GPT-6 Astra shows 38, which means part of the gap between those models and Mistral reflects refusals rather than inability. Mistral also states that its cyber refusal rate is higher than that of all open-source competitors, so the model is not simply unguarded. How a vendor should balance helpfulness for defenders against misuse is a live debate; our GLM 5.3 review covers the same tension and the WorldofAI video includes two cybersecurity demos of the model.
Coding and agents: competitive, not leading
On software engineering Mistral reports 61.7% on DeepSWE 1.1, 28.3% on Terminal-Bench 4.0, 59.4% on SWE-Atlas-QnA and 49.8% on the Artificial Analysis Coding Agent Index. In a human evaluation run by Surge AI, it ranked second of five models with a score of 3.74 out of 5.

The chart is honest about a weakness: Kimi K3 with its own CLI scores 68, ahead of Mistral's 62, and GLM-5.3 is close behind at 61. Note also that the competitors are run with different agent harnesses (Claude Code, Codex, Opencode and Kimi Code CLI), so the comparison mixes model and harness quality.

Terminal-Bench is blunter still. GLM-5.3 scores 40 against Mistral's 28, though Mistral leads Kimi K3 (21), Qwen3.8 Max (17) and DeepSeek V4 Pro (10) on this chart. In short, Large 4 sits in the upper half of the open-model pack for coding agents but does not take the crown. For a dedicated look at those rivals see Kimi K3 and DeepSeek V4 Pro.
Agentic workflows and enterprise tasks
Mistral reports 59.9% on AutomationBench and an Elo of 1,393 on AA-Briefcase, an Artificial Analysis evaluation of office-style knowledge work. In a human comparison against GLM-5.3 across coding, CAD, finance, mathematics and physics, the company says reviewers preferred Large 4 in the CAD and STEM categories, which also implies GLM-5.3 won elsewhere. It further claims state-of-the-art open-weight results on SciCode-Verified. We could not independently reproduce any of these.
Independent aggregate scores
The most sobering data comes from third-party trackers, which test with their own harnesses:
| Tracker | Result | What it implies |
|---|---|---|
| Artificial Analysis Intelligence Index | 38; ranked 64th of 225 models | Mid-table overall, well behind frontier closed models |
| Artificial Analysis speed | about 116 output tokens per second; 1.46 s to first token | Reasonable for a model of this size |
| Vals Index | 48.05% (plus or minus 1.11); 32nd of 44 models | Mid-table |
| Vals: individual benchmarks | 80.43% on MedScribe, 78.40% on Vibe Code Bench v1.1, 10.00% on ProofBench | Very uneven by task |
These figures say that Large 4 is a good, not dominant, model on general-purpose evaluations, with unusual strengths (cybersecurity, some legal and finance tasks, where Vals ranks it sixth of 75 on one legal-agent benchmark) and clear weaknesses such as formal proofs. That profile fits Mistral's own message: it is positioning for specific enterprise workloads rather than claiming to beat every frontier model at everything.
Safety, Security and Red-Teaming
Mistral says it is still red-teaming Large 4 in real-world settings with cybersecurity leaders, vetted partners and state authorities, which is notable for a model whose headline strength is offensive and defensive security capability. Alongside the cyber results it reports a 93.3% attack-resistance rate on the B3 agent-security benchmark and a score of 1.691 out of 2.0 on the KORA benchmark.

Note that Mistral Large 4 ties GLM-5.3 at 93.3 rather than beating it, and the truncated axis (it starts at 80) makes the bars look more different than the numbers are. B3 measures resistance to attacks such as prompt injection on agents, an increasingly important property as models are given tools. There is no published system card with the usual categories of dangerous-capability testing (biological, chemical, autonomy), so we cannot compare Large 4's risk posture with the detailed disclosures from the likes of Anthropic. If Mistral publishes a full report alongside the open weights, that will be the moment to reassess. The wider debate about releasing capable cyber models openly is covered in our piece on open weights and the Anthropic response.
WorldofAI's news round-up. The Mistral Large 4 segment starts around 13:00 and covers cybersecurity demos, DeepSWE 1.1 and pricing.
Pricing and Availability
Mistral's announcement, Artificial Analysis and Vals AI agree on a list price of $1.36 per million input tokens and $4.18 per million output tokens, roughly £1.00 and £3.10 at current exchange rates. Mistral's documentation shows a launch discount: 50% off for two weeks, giving $0.68 input, $0.07 cached input and $2.09 output per million tokens (about £0.50, £0.05 and £1.55). One analysis puts the end of the discount at 20/10/2026, though we could only confirm "two weeks" from the docs.
Available features per the model card include structured outputs, function calling, document question-answering, chat completions, batch processing and the agents and conversations APIs with built-in tools. OpenRouter also carries a listing for the model. Vals AI reports an average cost of $13.78 to run its index, a reminder that reasoning models consume many more tokens than their rate card suggests, so measure cost per completed task rather than cost per token.
Context for the price: it is above the lower tiers of Chinese open models but far below the premium closed frontier. For where Mistral's mid-tier sits in our tools directory see Mistral Medium 3.5.
The Sovereignty and Open-Weights Angle
Mistral frames Large 4 as part of a bid for sovereign, open AI at the frontier. It says the model is served from European infrastructure, independently of other digital service providers and under European law, and will be available across multiple regions worldwide. For governments, regulated industries and any firm uncomfortable with sending data to a US or Chinese provider, that is a real selling point that benchmark tables do not capture.
The open-weights promise is just as important, and just as unproven. A Hugging Face repository named Mistral-Large-4.0-1T05-A52B exists as an "upcoming release" page with a countdown and, at the time of writing, hundreds of people waiting; it contains no licence, benchmark or serving information. Mistral has a good record on open licences (Small 4 is Apache 2.0), but an open-weight trillion-parameter model is a different proposition: you will need a multi-GPU node even with aggressive quantisation, and the licence could range from fully permissive to a restricted community licence. Until it appears, plan on API use only.
How It Compares
| Model | Where it is ahead | Where Mistral Large 4 is ahead or level |
|---|---|---|
| Kimi K3 | DeepSWE 1.1 (68 versus 62) | Terminal-Bench 4.0 (28 versus 21); Cyber Index (50 versus 41) |
| GLM-5.3 | Terminal-Bench 4.0 (40 versus 28) | Cyber Index (50 versus 36); DeepSWE is effectively a tie (62 versus 61); B3 security is level at 93.3 |
| DeepSeek V4 Pro 0813 | Not shown ahead on any chart here | DeepSWE (62 versus 57), Terminal-Bench (28 versus 10), B3 (93.3 versus 85.2) |
| Qwen3.8 Max | Not shown ahead on any chart here | DeepSWE (62 versus 51), Terminal-Bench (28 versus 17) |
| Closed frontier models (Opus 5.5, GPT-6 Astra) | Aggregate indices, generally | Cyber Index successes (50 versus 29 and 33), partly due to refusals |
These comparisons are drawn solely from Mistral's charts and are only as good as its selection. For deeper analysis of each rival, see GLM 5.3, Kimi K3, DeepSeek V4 Pro and Qwen 3.8 Max.
Real-World Use Versus Benchmarks
Leaderboards measure what is easy to measure. The reason Large 4 may matter more than its index scores suggest is workload fit. A model that does well at vulnerability patching, finance and legal agent tasks, reads documents with a native vision encoder and can be hosted in the EU suits a specific buyer: a bank, law firm, public body or security team that values locality and auditability over a few points on a coding index.
Conversely, if your main workload is autonomous coding, the published evidence points to Kimi K3 or GLM-5.3 first, and to testing Large 4 as a second option. In both cases the cost picture should be measured on your own tasks: reasoning models can spend large numbers of tokens, and the launch discount will distort early cost comparisons.
A sensible evaluation plan has four steps. First, build a set of 30 to 50 real tasks from your workflow. Second, run Large 4 alongside your current model and one cheaper option, recording success, tokens and minutes per task. Third, test long inputs at the sizes you actually use rather than the advertised maximum. Fourth, run the set twice, since reasoning models vary between runs. Treat the result as a routing table rather than a verdict.
Limitations and Open Questions
- Preview status: pricing, limits and behaviour may change before general availability.
- Weights and licence unconfirmed: the end-of-October date is a promise, and the terms are unknown.
- Conflicting context figures: 1M, 524K and 512K all appear.
- Vendor-chosen comparisons: every chart was produced by Mistral and shows models it selected, some with different harnesses.
- No system card: there is no standard safety report covering biological, chemical or autonomy risks.
- No hands-on testing: we have not run the model, so we cannot say how it feels in daily use.
- Refusal effects: parts of the cyber lead come from rivals refusing tasks.
Who Should Use It
- European enterprises and public bodies that need EU-hosted, sovereign infrastructure and multilingual coverage.
- Security teams evaluating defensive automation, with proper governance around a model that is strong at offensive tasks.
- Finance and legal teams experimenting with document-heavy agent workflows, subject to their own validation.
- Self-hosters should wait for the weights and licence before planning anything.
- Not ideal for: developers who only want the best autonomous coding model today, or anyone who needs a published safety system card.
The Bottom Line
Mistral Large 4 is a credible comeback rather than a total reversal. It is a huge multimodal reasoning model with strong, independently supported cyber results, competitive coding scores, EU hosting and a promise of open weights. It is also mid-table on independent aggregate indices, behind Kimi K3 and GLM-5.3 on coding agents, still in preview and not yet open in practice.
Our advice is to try the preview whilst the launch discount lasts, benchmark it on your own tasks and wait for 31/10/2026 before making architectural decisions that depend on the weights. We will update this article when the weights, licence and any system card appear.
Sources
- Mistral AI: Introducing Mistral Large 4 (06/10/2026): specifications, benchmarks, languages, training hardware, availability and charts.
- Mistral docs: Mistral Large 4 model card: context window, launch pricing and features.
- Mistral docs: changelog: preview status and Mistral Small 4 model ID and context.
- Hugging Face: Mistral-Large-4.0-1T05-A52B: upcoming-release placeholder.
- Artificial Analysis: Mistral Large 4 Preview: Intelligence Index, speed, context and price.
- Vals AI: Mistral Large 4: Vals Index and task-level results.
- Kingy AI: Mistral Large 4 specs, benchmarks, pricing: third-party summary; no hands-on testing.
- TestingCatalog: Mistral Small 4 under Apache 2.0: Small 4 context.
- Matthew Berman on YouTube and WorldofAI on YouTube: creator coverage embedded above.
Charts: Mistral AI and Artificial Analysis, as published in the Mistral announcement. Hero: Mistral AI.
Last updated: 07/10/2026. Sourced from Mistral's announcement and documentation, the Hugging Face placeholder page and independent trackers. We have not benchmarked Mistral Large 4 ourselves; the model is in preview and prices, limits and open-weights timing may change.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






