Models·4 min read
By BitsMindsSource: Mistral AI

Mistral Large 4: Europe’s 1-Trillion-Parameter Open Model

Mistral has opened a public preview of Large 4, a 1.05-trillion-parameter mixture-of-experts with 52 billion active parameters, image input and a 1M-token context. It leads Mistral’s open-model coding chart, the weights are due by the end of October, and the launch price is $0.68 in and $2.09 out per million tokens.

Mistral Large 4, le Chonk The pixel-block Mistral logo cast as a thick, heavy slab with deep extruded sides, its face banded yellow to orange to red, standing on a dark warm floor under a single light, above a label reading Large 4, 1.05T parameters, 52B active. MISTRAL LARGE 4 1.05T PARAMS · 52B ACTIVE BITSMINDS.COM
Share:

Mistral has built its biggest model yet, and it is not shy about the size. Mistral Large 4, which the Paris company opened as a public preview on Tuesday, is a sparse mixture-of-experts with 1.05 trillion total parameters, of which 52 billion are active for each token. Its model page calls it “unofficially ML4, very officially: le Chonk”. It is a hybrid model that handles both quick instructions and longer reasoning, reads images as well as text, and according to the model documentation has a one-million-token context window and a 1.6-billion-parameter vision encoder.

For now it is reachable only through the API in Mistral Studio. The weights are promised by the end of October; the date reported by VentureBeat and Reuters is 27 October. Before then, Mistral is running a roughly three-week test with developers, cybersecurity firms and government authorities, and says reinforcement learning on the model is still under way, so the checkpoint that ships may score differently from the one being tested today.

Where it lands

Mistral compares Large 4 with other open-weight models, not the closed frontier. On DeepSWE v1.1, a software-engineering benchmark, it scores 61.7%, just ahead of Zhipu’s GLM-5.3 and clear of DeepSeek V4 Pro, Qwen 3.8 Max and Reflection’s Beam, which was previewed a day earlier. That is still well short of the closed models: VentureBeat notes that the public DeepSWE leaderboard has Claude Opus 5 and GPT-6 Astra at about 74%. Mistral also reports 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4.

DeepSWE v1.1: Large 4 against open rivalsTasks solved, % — higher is better · Mistral, 6 Oct 2026Mistral Large 4GLM-5.3DeepSeek V4 ProQwen 3.8 MaxReflection Beam02040608010061.761.057.051.044.0DeepSWE v1.1
Mistral’s own comparison of open-weight models. The public DeepSWE leaderboard has closed models such as Claude Opus 5 and GPT-6 Astra near 74%. Data: Mistral, via VentureBeat.

The more revealing number may be the blind test Mistral commissioned from Surge AI, in which people rated code from five models on a scale of one to five. Large 4 averaged 3.74, a little ahead of GLM-5.3 and Moonshot’s Kimi K3, and well behind Claude Opus 5 at 4.22. Outside coding, Mistral says Large 4 beats GPT-6 Astra on Vals.ai’s finance and legal evaluations, is the best open model on Harvey’s Legal Agent Benchmark, and edges Astra on Dense200 visual grounding, 42% to 41%. All of these are Mistral’s own figures, and the model has not yet appeared on independent leaderboards such as Artificial Analysis.

Blind human ratings of coding outputMean score on a 1–5 scale (axis from 0) — higher is better · Surge AI for MistralClaude Opus 5Mistral Large 4GLM-5.3Kimi K30123454.223.743.603.59Surge human eval
Human raters preferred Large 4’s code to the Chinese open models but placed it clearly behind Anthropic’s previous flagship. Data: Mistral.

A cyber model by design

Security is the clearest part of the pitch. Mistral says Large 4 solves 93% of Cybench’s 40 capture-the-flag exercises, ranks in the top five of the Artificial Analysis Cyber Index, and scores 82% on a test that asks a model to reproduce a vulnerability and then patch it, the highest result it has seen. It claims Claude Opus 5.5 and GPT-6 Astra score close to zero on that test because they refuse the task. The same model, Mistral says, also has the highest refusal rate for harmful cyber requests of any open model it tested, and vetted partners and state authorities are getting versions with lighter moderation and stronger cyber abilities during the preview.

The pricing undercuts most frontier APIs. The list price is $1.36 per million input tokens and $4.18 per million output tokens, and Mistral’s docs show a launch sale at half that: $0.68 in, $0.07 for cached input and $2.09 out. Self-hosting on private cloud or on-premises hardware will be possible once the weights are released, and Mistral is also serving the model from a European deployment it runs end to end.

Built in Europe

Large 4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral’s own European data centres, on data covering more than 160 languages, including every official EU language. Its predecessor, Large 3, had 675 billion total parameters and 41 billion active. The reinforcement-learning stage alone uses about 3,000 GPUs and produces some 33 billion tokens a day, and Mistral says it has seen no sign of the gains levelling off.

The model is the first big product of the €3 billion Series D Mistral raised in September, and the company says it will serve as the base for a family of specialised models. Chief scientist Guillaume Lample told VentureBeat that ML4 “is at the frontier of open weight models” and that larger ones are coming. The open-weight claim is the one that matters most to Mistral’s European customers, and it gets tested at the end of the month, when the weights and their licence terms are expected to be published.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Claude Haiku 5.5: Anthropic's fastest model, at a tenth of Haiku 4.5's price A terracotta and ivory stopwatch, tipped to show its depth, has the official Claude asterisk inlaid in clay as the hub of its green sweep hand. A paper tag tied to the crown reads minus 90 percent against Haiku 4.5. Beside it, the title names Claude Haiku 5.5 and quotes its API prices for prompts up to 100,000 tokens: 10 US cents per million input tokens and 50 cents per million output tokens. The stopwatch is an editorial metaphor for Anthropic's claim that Haiku 5.5 is its fastest model at standard speed, not an Anthropic product; the 90 percent figure applies to prompts up to 100,000 tokens, and longer prompts cost five times as much. 51015202530354045505560 CLAUDE HAIKU 5.5 PER-TOKEN PRICE −90% VS HAIKU 4.5 PROMPTS UP TO 100K CLAUDE Haiku 5.5 $0.10 INPUT $0.50 OUTPUT PER MILLION TOKENS BITSMINDS.COM
Models

Claude Haiku 5.5 Is Out at a Tenth of Haiku 4.5's Price

Reflection Beam: a large core, a narrow active path An original monochrome optical sculpture on a dark reflective laboratory bench. A narrow white beam passes through a hollow silver ring and branches toward a few illuminated cartridges within a large graphite core. Most cartridges remain dark; their paths recombine into a beam that reaches a solid circular receiver. The hollow and solid circles echo Reflection's mark. The labels read BEAM, 501B total parameters, 23B active per token, PREVIEW and WEIGHTS PLANNED — OCT. The illustration represents sparse inference in the model, not physical hardware or its exact expert layout. BitsMinds original editorial vector artwork for reflection-ai-beam-501b-open-weight-model. Facts verified on 6 October 2026 at https://reflection.ai/blog/introducing-beam: 501B total parameters, 23B active, weights planned later in October under Apache 2.0. The module layout is symbolic. No performance ranking or measured-cost claim is depicted. Self-contained SVG with the project Reflection wordmark. BEAM 501B TOTAL PARAMETERS 23B ACTIVE PER TOKEN PREVIEW WEIGHTS PLANNED · OCT BITSMINDS.COM
Models

Reflection’s Beam: A 501B Open Model Built to Be Frugal

Opus 5.5 vs GPT-6 Astra vs Grok 4.7 Build No Man's Sky
Models

Opus 5.5 vs GPT-6 Astra vs Grok 4.7 Build No Man's Sky