Introducing Mistral Large 4 | Mistral

Le Chonk
Today, we’re launching a public preview of Mistral Large 4. Unofficially ML4, very formally: le Chonk. ML4 pushes the frontier of open-weight efficiency. You can attempt the preview API right this moment on Mistral Studio. Weights drop finish of this month.
Frontier efficiency
ML4 is a 1 trillion-parameter natively multimodal mannequin with 49 billion lively parameters. It is our largest and most succesful mannequin thus far, and it continues to enhance quickly as we refine it.
The mannequin demonstrates distinctive efficiency throughout coding, agentic workflows, and multimodal understanding. It already achieves efficiency aggressive with the strongest open-source fashions globally, whereas considerably outperforming any open-weight mannequin developed within the US or Europe. On vital enterprise workloads, together with cybersecurity, finance and regulation, we discover it to be state-of-the-art amongst open fashions. In some domains comparable to visible grounding, it goes additional nonetheless, surpassing even frontier closed fashions.
We will launch the weights by the top of the month. Until then, we’re red-teaming the mannequin in real-world settings with cybersecurity leaders, vetted companions, and state authorities, who will entry the identical mannequin with lowered moderation and expanded cyber capabilities.
Demos
Forged in Europe. Built for AI sovereignty.
ML4 was educated from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s personal datacenters in Europe. The public preview is served on that very same development projects. It is a big milestone in our long-term capital allocation throughout development projects, analysis, and product improvement: state-of-the-art efficiency in vital verticals, delivered by way of open weights, designed to offer prospects management over their AI.
This is especially necessary in cybersecurity, the place provider-level refusals can block reputable vulnerability analysis and incident response, and the place shedding entry to a functionality mid-incident can itself develop into a vital safety danger. ML4 pairs top-tier cyber efficiency with open weights and self-deployment, giving organizations each the aptitude and the autonomy to run superior safety work beneath their very own insurance policies.
The mannequin shall be obtainable throughout a number of areas worldwide, together with a European deployment that Mistral operates end-to-end, independently of different digital service suppliers and beneath European regulation. Fun truth: a big share of ML4’s coaching information was multilingual, spanning greater than 160 languages, together with each official language of the European Union.
We’ve been working intently with main enterprises throughout the globe in finance, engineering, manufacturing, logistics, prescribed drugs, science, transport, public sector, and different mission-critical industries to coach ML4. In truth, the mannequin makes use of the identical coaching, customization, and RL surroundings we provide our prospects by way of Mistral Forge.
Try it right this moment
There continues to be extra to come back. As we work towards releasing the weights, we’ll share additional particulars on the mannequin structure, further benchmarks, and our post-training methodology.
This mannequin will even function the inspiration for a brand new technology of specialised and optimized Mistral fashions. In the meantime, we invite you to try the preview API and share your suggestions with us on social media.
Capabilities deep-dive
Cybersecurity
ML4 is without doubt one of the globe’s strongest AI fashions for cybersecurity. On the Artificial Analysis Cyber Index, an impartial analysis of how nicely AI fashions discover and repair safety flaws in actual software program, it ranks among the many prime 5 fashions globally and leads open-weight fashions developed exterior China by a large margin. On one of many index’s exams, which asks a mannequin to breed an actual vulnerability in open-source software program after which patch it, ML4 scores 82%, the very best of any mannequin. It additionally solves 93% of the challenges in Cybench, a set of 40 workouts drawn from safety competitions, one of many highest scores reported for an open-weight mannequin.
That prime rating displays a sensible benefit. Several main closed fashions, together with Claude Opus 5.5 and GPT-6 Astra, rating close to zero on the identical check as a result of they refuse to carry out the duty. Yet defending software program usually begins with proving {that a} flaw is actual, precisely the form of work security filters in closed fashions can block. This issues much more as menace actors more and more jailbreak those self same fashions to help offensive cyber exercise: defenders want programs that may match these capabilities with out being constrained by the identical refusals. ML4 can do this work, and its capabilities lengthen past what it was explicitly educated for: in inner testing, it proved helpful for analysing malware, prioritising vulnerabilities, and writing detection guidelines. For organisations that want sovereign, auditable AI for safety operations, will probably be in a position to run on personal cloud or on-premise.
ML4 in opposition to the sphere : effectively reasoning over numerous complicated challenges
Malware reverse-engineering: fixing an out-of-distribution investigation process
Agentic coding
ML4 excels throughout software program engineering, repository understanding, and sophisticated terminal workflows, scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. Its mixed Coding Agent Index rating of 49.8% locations it forward of DeepSearch V4 Pro 0813 and Qwen3.8 Max.
* Scores on DeepSWE, Terminal-Bench 4, and SWE Atlas QnA use the numbers reported by the ArtificialAnalysis coding index. Vibe Code Bench v1.1 makes use of the numbers reported by vals.ai.
We additionally ran a blind human analysis with Surge AI on coding high quality: skilled annotators rated mannequin outputs on a 1–5 scale, with mannequin identities hidden. ML4 Preview ranked second of 5 fashions (3.74), forward of Kimi K3 (3.59), GLM-5.3 (3.60) and GLM-5.2 (3.40), and behind solely Claude Opus 5 (4.22).
Agentic Workflows
ML4 runs general-purpose brokers that collect data, use instruments, and produce completed deliverables throughout complicated workflows. On AutomationBench — 657 commerce workflows throughout apps like Gmail, Google Sheets, Slack, and Salesforce — it scores 59.9%, forward of Kimi K3, MiMo-V2.6-Pro, and DeepSearch V4 Pro.
It’s simply as sturdy on the skilled deliverables that data work really produces: spreadsheets, slides, and PDFs. On AA-Briefcase, which evaluates long-horizon data work, it reaches 1,393 Elo, forward of DeepSearch V4 Pro.
Multimodal
ML4 is a step change within the capacity of our fashions to grasp photographs. It causes powerfully throughout complicated paperwork, charts, and pure photographs, and brings imaginative and prescient to the industries the place notion is vital comparable to engineering, manufacturing, and earth commentary.
The mannequin can additional mix visible grounding with agentic capabilities: from inspecting gigapixel satellite tv for pc imagery — serving to disaster-response groups act when time counts — to analyzing engineering-drawings — zooming in, inspecting, and verifying till the reply is precise. In our demos above, ML4 grounds dense pure scenes, verifies mechanical elements in technical drawings, retrieves proof from PDFs, and scans large geospatial photographs for the hardest-to-find objects.
On visible grounding notably, we discover ML4 to be one of the crucial succesful fashions we examined, as an example surpassing GPT-6-Astra on Dense 200 (42% vs 41%).
Science and Math
ML4 brings sturdy scientific capabilities, constructed by combining AI-driven strategies with our researchers’ experience in arithmetic, physics, and chemistry.
It’s extremely proficient at agentic coding for scientific duties comparable to information evaluation, modeling, and simulating bodily actuality, which lets researchers give attention to the questions somewhat than the plumbing. In benchmarks, ML4 is cutting-edge on SciCode-Verified amongst open-weight fashions. In follow, it will probably generate a full Hartree–Fock simulation in a single shot — a fancy, multi-step chemistry process constructed from a collection of superior routines.
ML4’s math is stronger too, in each formal reasoning and utilized arithmetic. In our human evaluations it causes extra exactly and with extra construction than GLM-5.3, and it will probably maintain lengthy, domain-specific applied-mathematics duties, together with work related to frontier theoretical physics.
Together, these capabilities make ML4 a powerful analysis assistant throughout the total technical workflow — from the primary query to the ultimate consequence.
%201_1i0lB0.webp?dpl=6ac4f7597a7b31848042d469)
SciCode-Verified exams the capabilities of fashions to implement complicated scientific workflows in code for domains comparable to physics, arithmetic, materials science and biology.

Internal eval on STEM duties (math and physics) of ML4 in opposition to GLM5.3
Knowledge Work
ML4 is our most succesful mannequin for the real-world duties which professionals deal with on daily basis. It can create, edit and repair complicated spreadsheets and paperwork, displaying exemplary efficiency on each authorized and monetary benchmarks.
Notably, we evaluated ML4 by way of third celebration evaluators (vals.ai) on consultant duties for each authorized and monetary duties, discovering the mannequin exceeds GPT-6-Astra in each instances. On HarveyAI’s Legal Agent benchmark, ML4 outperforms all open-source fashions.
FinWorkBench exams mannequin capabilities at creating/enhancing spreadsheets on actual life Finance and Accounting use instances.
Financial evaluation calls for precision and the power to synthesize data from a number of sources, a course of that continues to be time-consuming at many monetary establishments right this moment. In this demo, ML4 and Mistral Medium 3.5 tackle the identical multistep company finance problem, looking by way of public firm filings and monetary studies, comparable to these obtainable through EDGAR and equal European databases. An animated semantic map traces every mannequin’s journey towards an answer, highlighting each doc retrieved alongside the best way. Each monitor’s place displays the proof gathered, the outcomes of calculations, and the questions that stay unresolved. Viewers can observe how the investigations unfold and examine the distinct paths every mannequin takes earlier than arriving at its ultimate reply.
Model Safety
ML4 has saturated our benchmarks on robustness to oblique immediate injections, placing it on the frontier of OSS fashions (in comparison with GLM-5.2, GLM-5.3, Kimi-K2.6, Kimi-K3, DS-V4-Pro-0813). On Lakera’s public B3 AI Security Benchmark, ML4 resists 93.3% of assaults – we see no increased scores amongst rivals.
ML4 additionally engages extra responsibly with customers than any of our earlier fashions. We spotlight our outcomes on the KORA Benchmark, the place ML4 once more sits at our highest measured rating amongst OSS fashions (1.691, with 2 being the utmost denoted as “Exemplary”).
Of specific relevance is the mannequin’s propensity to refuse malicious requests concerning cybersecurity. Despite sturdy efficiency on Cyber benchmarks, the common refusal fee of the mannequin on cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm is increased than all OSS fashions.
Human Evaluation
We ran an inner analysis wherein skilled annotators throughout coding, computer-aided design (CAD), finance, arithmetic and physics in contrast Mistral Large 4 with GLM-5.3. ML4 was most well-liked in CAD and STEM, whereas acting on par or near GLM-5.3 in finance and coding.

Reinforcement studying at scale
Base fashions are enhancing quick, and our post-training has to maintain tempo. A recipe tuned for yesterday’s mannequin leaves functionality on the desk with right this moment’s frontier, as a result of floor reality samples that after pushed a mannequin to its limits received’t anymore. We use Reinforcement Learning (RL) as a result of it adapts because the mannequin does: we prepare on the outcomes of the mannequin’s personal makes an attempt, and we are able to elevate the issue and the breadth of the duties because it will get stronger.
Our RL library was designed to make new environments simple so as to add and prepare at scale. A shared, composable interface permits a single coaching run to mix duties starting from single-turn chat and sophisticated scientific downside fixing to security alignment, factuality, and long-horizon instrument use. These environments share scaffolds and sources comparable to code sandboxes, internet search, and exterior APIs. The similar composability extends to verification, with reward fashions, unit exams, LLM judges, and static checks mixed as wanted for every process.
At runtime, an autoscaling fleet of actors generates tens of hundreds of rollouts in parallel whereas mannequin coaching proceeds asynchronously. The technology and coaching pipeline is optimized for lengthy trajectories, supporting rollout budgets of thousands and thousands of tokens throughout a number of compactions whereas protecting staleness low. Novel strategies and optimizations throughout each phases decrease off-policy drift and allow secure RL over lengthy horizons.
At our present scale (3k GPUs), a single coaching run produces roughly 33 billion tokens per day, of which round 16 billion trainable completion tokens after filtering and masking. We can see the run progress immediately within the coaching rollouts: coaching rewards rise throughout a number of consultant environments because the coverage learns to unravel more and more complicated duties. Below are just a few examples.

The enhancements should not particular to the environments we prepare on; they switch to downstream evals, and the ultimate mannequin owes them to each post-training phases (supervised fine-tuning and RL), as proven within the charts.

What comes subsequent
This is just the start. ML4 is the primary milestone on the roadmap funded by our €3 billion Series D — the biggest fairness spherical ever raised by a European expertise firm. That capital is already being put to work: we’re considerably scaling up our compute capability in our personal European datacenters, and way more is coming on-line within the months forward.
More compute means extra coaching. The reinforcement studying run behind this preview continues to be in flight, and the mannequin is displaying no indicators of saturation — there may be substantial headroom forward. As we scale up coaching on our expanded development projects, we count on giant and fast enhancements within the weeks and months to come back.
We will launch the weights by the top of the month, together with extra particulars on the structure, further benchmarks, and our post-training methodology. And ML4 is just the inspiration: it should function the bottom for a brand new technology of specialised and optimized Mistral fashions, constructed for the industries and workloads our prospects care about most.
The tempo of progress from right here shall be quick. Stay tuned.
