Mistral Large 4 'Le Chonk': Europe's best open model is only eighth worldwide
Mistral calls Large 4 state-of-the-art among open models. Artificial Analysis scores it eighth, behind seven Chinese ones, at four times a peer's cost per task.

Mistral opened a public preview of Mistral Large 4 on 6 October 2026 — its largest model yet, and the one it wants Europe to rally behind. The technical numbers are big: one trillion parameters, 52 billion active, natively multimodal, a hybrid instruct-and-reasoning mixture of experts trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside the company's own European datacenters. The pitch is bigger still. Mistral calls it state-of-the-art among open models on cybersecurity, finance and law, and says it outperforms any open-weight model built in the United States or Europe. The first independent benchmark sits between that claim and a reality check: Artificial Analysis scores the preview 38.4, the most intelligent model from outside the United States and China — and only eighth among open-weight models, behind seven Chinese ones.
Reading the launch terms
What ships today is an API preview on Mistral Studio; the weights themselves arrive at the end of October. Until then Mistral says it is red-teaming the model in real-world settings with cybersecurity leaders, vetted partners and state authorities, who get the same model with reduced moderation and expanded cyber capabilities. The company trained it on 3,800 GPUs in its own European facilities, fed it more than 160 languages including every official European Union language, and funded the work with a EUR 3 billion Series D it describes as the largest equity round ever raised by a European technology company. Two footnotes matter for reading every number below. The reinforcement-learning run behind the preview is, by Mistral's own account, still in flight, so the published scores are a snapshot of a moving target. And the spec sheets disagree on size: Mistral reports 52 billion active parameters, while Artificial Analysis lists 49 billion and a roughly half-million-token context window.
Mistral's own scorecard
Every claim in this section is Mistral's, produced by Mistral, on harnesses Mistral chose. The company says the model is competitive with the strongest open-source models globally, that it beats every open-weight model from the US and Europe, and that it leads open models on cybersecurity, finance and law. On cyber it reports solving 93% of the challenges in Cybench, a set of 40 security-competition exercises, and scoring 82% on a test that reproduces a real vulnerability in open-source software and then patches it — the highest of any model. Part of that edge is structural rather than competitive: several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse the task. A defender proving a flaw is real is exactly the work safety filters can block, so the refusal gap is a genuine argument. It is also, until an outside party reruns it, a vendor's reading of one test.
Where the independent index puts it
Artificial Analysis, whose Intelligence Index v4.3.2 folds ten evaluations into a single number, is the first outside party to score the preview. It lands at 38.4. Among models built for open weights that is eighth, and the seven ahead of it are all Chinese.
| Model | Intelligence Index | Cost per task (USD) |
|---|---|---|
| MiMo-V2.6-Pro | 46.3 | $0.13 |
| GLM-5.3 (Max) | 44.8 | $2.01 |
| Kimi K3 (Max) | 43.6 | $2.00 |
| Qwen3.8 2.4T | 39.9 | $2.16 |
| Mistral Large 4 (Preview) | 38.4 | $1.13 |
That ordering still lets Mistral keep its narrowest claim. The strongest open model outside China before this was South Korea's Motif 3 at 33.6, so the preview is genuinely the best open-weight model from the United States or Europe — a line the evaluator itself endorses. The gap to the frontier is the harder fact. Claude Opus 5.5 leads at 57.6, GPT-6 Astra at 52.7 and Gemini 4 Argon at 52.6; the best open model, MiMo-V2.6-Pro, is nearly eight points clear, and the best closed model nineteen. Because the weights are not out, Artificial Analysis currently classifies the preview as proprietary, not open.
Fast, verbose, and four times the price of a peer
Cost is where the open-weight comparison bites hardest. Running the Intelligence Index costs $1.13 per task for Large 4, against $0.27 for DeepSeek V4.1 Flash — which scores higher — and $0.13 for MiMo-V2.6-Pro. Artificial Analysis calls it more than four times the cost per task of open-weight models at similar intelligence. A 50% launch discount for the first two weeks, cutting token prices to $0.68 per million input and $2.09 per million output, brings the per-task figure to about $0.57 — still above GLM-5.3-Flash at $0.25. The driver is verbosity: across the index the model generated about 200 million output tokens against a class median of 81 million, which the evaluator labels "very verbose." Speed is the counterweight. Large 4 produced 116.1 output tokens per second against a median of 86.1, and reached its first token in 1.46 seconds against 3.81. Standard pricing is $1.36 per million input tokens and $4.18 per million output, with cached input at $0.14.
The cyber claim is the one that mostly holds
On the independent Cyber Index the picture is more flattering. Large 4 scores 50, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56, and on CyberGym-E2E-AA it scores 82%, ahead of MiMo (79%) and GPT-6 Luna max (78%). Artificial Analysis says that once the weights ship, the model will rank in the top three open-weight models on the Cyber Index. Elsewhere the gains are thinner. On GDP.pdf, a document-reasoning test, it scores 19%, on par with MiMo and behind Kimi K3 at 22%. The evaluator credits part of an 18-point improvement over Mistral Large 3 to an API change that now accepts 100 images per request, up from eight. One caveat qualifies the whole section: the expanded-cyber, reduced-moderation variant handed to the authorities is not the one the public API serves.
What would settle it
Reception split along familiar European lines — sovereignty memes, a r/LocalLLM thread titled "Europe finally takes the lead" at roughly 1,200 points, and "LE CHATON FAT IS REAL" on r/singularity at about 654. That is reception, not measurement. Two open questions decide how this reads a month from now. The license is unnamed, and Mistral has shipped open weights under Apache 2.0, a revenue-capped MIT and a research-only license in the past, so the terms will decide how open "open" turns out to be; last week's Aleph Alpha Kolibri release left the same gap. And the model is still training, so 38.4 could rise before the weights land — the run, not the freeze, is the snapshot. What would move this past vendor-and-evaluator framing is concrete: a published license and architecture, per-task scores instead of an aggregate, an independent rerun at matched effort, and a second evaluation once the reinforcement-learning run finishes and the weights of this model and its rivals such as GLM-5.3 and DeepSeek V4.1 Flash can finally be run side by side. Europe's best open model is a real achievement. Being eighth is the part the marketing did not mention. For now the sovereignty claim is auditable and the capability claim is not.


