Mistral Large 4 preview: 1-trillion-parameter model scores 38 on independent index, open weights due end of October
Mistral opened an API preview of the 1.05-trillion-parameter Large 4, which Artificial Analysis scores at 38, four times Large 3 and the best from outside the US and China. It ties GPT-6 Luna at about ten times the per-token price; its case rests on open weights due this month, under a licence not yet published.
What happened
On 6 October, Paris-based Mistral released a public preview of Mistral Large 4, which the company nicknames “le Chonk”. Developers can call it now through the Mistral Studio API under the name mistral-large-4.
- Total parameters
- 1.05 trillion
- Active per token
- 52 billion
- Context (official docs)
- 1M tokens
- API list price (per 1M tokens in/out)
- $1.36 / $4.18
Source: Mistral model documentation, checked 9 Oct 2026. A 50% preview sale price of $0.68 / $2.09 also applies.
According to Mistral’s announcement:
- It sees images. It is natively multimodal: text and image input (via a 1.6-billion-parameter vision encoder), text output.
- It can reason. Ordinary answers and a step-by-step reasoning mode live in the same model; API users choose whether reasoning is on.
- It was trained in Europe, from scratch, on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacentres. The preview runs on the same machines.
- It is multilingual. The training data covered more than 160 languages, including every official language of the EU.
- Weights are coming “by the end of the month”. Until then, Mistral is red-teaming the model with cybersecurity firms, vetted partners and state authorities, who get a version with lighter moderation.
How good is it? From 9 to 38 in ten months
The clearest yardstick is Mistral’s previous flagship. Large 3, released in December 2025, had 675 billion total and 41 billion active parameters and a 256,000-token context. Large 4 is about 50% bigger in total and, per the official docs, has a context window about four times longer.
The bigger change is capability. Artificial Analysis (AA), an independent benchmarking firm, combines 10 tests, from building spreadsheets and slides to command-line programming, scientific coding and expert-level exam questions, into one Intelligence Index. On its figures as of 9 October:
AA Intelligence Index v4.3.2; most models shown at their highest reasoning setting. Source: Artificial Analysis model pages and launch analysis, checked 6 and 9 Oct 2026
Three ways to read that chart:
- Against itself: from 9 to 38 in ten months, roughly a fourfold increase. AA says it gives France back “the most intelligent model from outside the US and China”, ahead of models from South Korea and the United Arab Emirates.
- Against other open-weight models: it is level with DeepSeek V4.1 Flash (39), slightly ahead of DeepSeek V4 Pro (36) and behind Zhipu’s GLM-5.3 (45). Mistral’s claim that it “significantly outperforms any open-weight model developed in the US or Europe” is consistent with AA’s data, but it is not the strongest open-weight model in the world.
- Against the best closed models: about 20 points behind Claude Opus 5.5. The British developer Simon Willison put it this way: it is “certainly not a Fable-class model”, but it puts Mistral back to “maybe about 6 months behind the frontier”.
The gap is widest on Terminal-Bench 4.0, which asks a model to finish jobs on its own in a command line. On AA’s measurements Large 4 Preview completes about 27% of tasks, against about 60% for Claude Opus 5.5 and 59% for GPT-6 Astra.
Mistral’s own results
The scores below were published by Mistral. Many of the tests were run by third parties such as AA and vals.ai, but Mistral chose which tests to show and which rivals to compare against.
Cybersecurity is its standout. CyberGym-E2E asks a model to reproduce a real vulnerability in open-source software and then patch it. Mistral reports 82%, the highest of any model; AA’s own write-up gives the same 82%, ahead of MiMo-V2.6-Pro (79%) and GPT-6 Luna (78%). On AA’s broader Cyber Index it scores 50, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro (56).
Chart published by Mistral; measured by Artificial Analysis. Source: Mistral, “Introducing Mistral Large 4”, 6 Oct 2026
Mistral points out that Claude Opus 5.5 and GPT-6 Astra score near zero on this test because they refuse to do it. Its argument: defending software often starts with proving that a flaw is real, which is exactly what closed models’ safety filters tend to block, while attackers are busy jailbreaking those same models.
Elsewhere the results are mixed:
| Test | What it measures | Large 4 Preview | Rivals (same chart) |
|---|---|---|---|
| AutomationBench | 657 business workflows across Gmail, Sheets, Slack and other apps | 59.9% | GLM-5.3 62.2%, Kimi K3 58.3%, DeepSeek V4 Pro 56.7%, Mistral Medium 3.5 6.3% |
| Finance Agent v2 (vals.ai) | Financial analysis tasks | 54.7% | GLM-5.3 55.8%, GPT-6 Astra 53.5%, Kimi K3 53.1% |
| DeepSWE 1.1 | Changing code in real repositories | 62% | Kimi K3 68%, GLM-5.3 61%, DeepSeek V4 Pro 57% |
| AA-Briefcase | Long office tasks: spreadsheets, slides, PDFs | 1,393 | MiMo-V2.6-Pro 1,516, GLM-5.3 1,510, DeepSeek V4 Pro 1,256 |
| Dense 200 | Locating objects precisely in crowded images | 42% | GPT-6 Astra 41% |
Mistral also commissioned a blind review of coding quality from Surge AI, in which annotators could not see which model wrote what. Large 4 Preview scored 3.74 out of 5, second of five models: ahead of Kimi K3, GLM-5.3 and GLM-5.2, behind Claude Opus 5 (4.22).
Price: not a bargain through the API
| Model | AA index | Per 1M tokens, input / output | AA cost per task | Open weights |
|---|---|---|---|---|
| Mistral Large 4 Preview | 38 | $1.36 / $4.18 (sale: $0.68 / $2.09) | $1.13 (about $0.57 at sale price) | End of Oct (planned) |
| GPT-6 Luna | 38 | $0.10 / $0.50 | $0.07 | No |
| Claude Haiku 5.5 | 43 | $0.10 / $0.50 (prompts up to 100k tokens) | $0.21 | No |
| DeepSeek V4.1 Flash | 39 | $0.30 / $1.20 (peak hours) | $0.27 | Yes |
| DeepSeek V4 Pro | 36 | $1.32 / $3.96 (peak hours; half price otherwise) | $0.67 | Yes |
| Gemini 3.8 Flash | 41 | $0.75 / $3.75 ($1.50 / $7.50 from 2027) | $1.24 | No |
| GPT-6.1 Sol | 52 | $2 / $10 | $0.72 | No |
| Claude Sonnet 5.5 | 56 | $2 / $10 | $5.46 | No |
Prices are from the official pricing pages of Mistral, OpenAI, Anthropic, Google and DeepSeek, checked on 9 October 2026, at standard API rates. “AA cost per task” is AA’s average cost, at list prices, of one Intelligence Index task, so it reflects both price and how many tokens a model uses.
Two things stand out:
- Middling price, heavy usage. Large 4’s per-token price is close to DeepSeek V4 Pro’s, but AA found it generated about 200 million tokens to complete the index, more than twice the median of 81 million for comparable models. AA’s conclusion is that it costs more than four times as much per task as open-weight models of similar intelligence. Even at the 50% launch discount, which AA says runs for the first two weeks, it costs twice as much per task as DeepSeek V4.1 Flash ($0.27).
- Same score, very different bill. GPT-6 Luna, which also scores 38, costs 7 cents per task. Judged purely on capability per dollar through an API, Large 4 has no edge today.
Why it matters for Europe and for self-hosting
Mistral’s slogan for the launch is “Forged in Europe. Built for AI sovereignty.” Its case has three parts:
- Trained and served in Europe, on Mistral’s own infrastructure. The company will also offer a deployment that it runs end to end, independently of other digital service providers and under European law.
- Open weights. Companies and governments will be able to run the model in their own private cloud or server room, keeping data in-house, with no risk of a provider changing its policy at a critical moment. Mistral’s example: losing access to a model in the middle of a cyber incident is itself a security risk.
- Money. Mistral calls Large 4 the first milestone funded by its €3 billion Series D, which it describes as the largest equity round ever raised by a European technology company. The money is going mainly into more compute in its European datacentres.
Every model above it on AA’s chart is either a closed American model or Chinese. For European organisations bound by rules such as the GDPR and unwilling to send data to foreign clouds, a European model close to the second tier that they can run on their own hardware did not exist until now.
But running it yourself is a big job. At 1.05 trillion parameters, the weights alone take about 2.1 TB at 16-bit precision, and about 0.5 TB even compressed to 4 bits. With 80 GB NVIDIA H100 GPUs, that means at least 27 cards in the first case and 7 in the second, before counting the extra memory needed for long inputs. (These are YLEM’s estimates at 2 bytes and 0.5 bytes per parameter; Mistral’s documentation currently lists the memory requirement as “N/A”.) This is hardware for large companies and government datacentres, not a laptop.
Things to keep in mind
- No weights, no licence yet. As of 9 October, both the Hugging Face link (Mistral-Large-4-1T-A52B) and the licence field in Mistral’s documentation say “Coming soon”. Mistral’s licences have varied: Large 3 used the permissive Apache 2.0, which allows free commercial use; the earlier Large 2.1 used the Mistral Research License (MRL), under which commercial use requires a separate licence. Which one Large 4 gets will decide how much “open weights” means for businesses.
- The specs differ by source. Mistral’s documentation says 52 billion active parameters and a 1M-token context; AA and Simon Willison say 49 billion active, and AA lists the API’s context as 524,000 tokens.
- It is still training. Mistral says the reinforcement-learning run behind the preview is “still in flight”. The released model and weights may score differently from today’s preview.
- Mistral picked most of the comparisons. It compares itself mostly with Chinese open-weight models and rarely with closed frontier models; on agentic coding tests such as Terminal-Bench the gap to Claude and GPT frontier models is large.
- Cyber capability cuts both ways. A model that doesn’t refuse is useful to defenders and attackers alike. Mistral says Large 4 refuses malicious cyber requests more often than any other open-weight model, but once weights are public anyone can strip those safeguards out.
What it means for you
- Chat users: nothing changes yet. Mistral’s announcement only mentions the API preview and says nothing about when Large 4 comes to its Le Chat app.
- Developers: you can try it now through the API in Mistral Studio at the preview sale price ($0.68 per million input tokens, $2.09 per million output). If you just want a cheap, capable model, GPT-6 Luna and DeepSeek V4.1 Flash score about the same for less.
- European businesses, governments and security teams: this is the first model near the second tier that was trained in Europe and is due to have open weights. Two things to watch at the end of the month: whether the weights arrive on time, and whether the licence allows free commercial use.
- For a broader ranking of models, see our models page; for the Chinese open-weight approach, see DeepSeek V4.
Sources
- Mistral · Introducing Mistral Large 4
- Mistral · Large 4 model documentation and preview pricing
- Artificial Analysis · Mistral Large 4 launch analysis
- Artificial Analysis · Mistral Large 4 Preview model page
- Simon Willison · Introducing Mistral Large 4
- Mistral · Large 3 documentation
- Anthropic · Claude API pricing
- OpenAI · API pricing
- Google · Gemini API pricing
- DeepSeek · Models & pricing