Why Mistral Large 4 puts cyber defense at the center of AI

Mistral Large 4 is a public preview of Mistral’s largest model, with weights expected at the end of October. Its clearest pitch is cybersecurity: Mistral says the model can reproduce and patch vulnerabilities that some closed models refuse to handle.

Why Mistral Large 4 puts cyber defense at the center of AI

Mistral Large 4 arrives as a direct attempt to make Europe more competitive in advanced AI while sharpening one particular use case: security work. The model is now in public preview through Mistral Studio, with model weights expected at the end of October.

The company is positioning ML4 as both a technical step forward and a sovereignty play. It is a one-trillion-parameter model trained in Mistral’s own European data centers, and Mistral says it is built for organizations that need more control over where AI runs and what it is allowed to do.

A larger open-weight contender from Europe

Mistral describes Mistral Large 4, also called ML4 and nicknamed "le Chonk," as its biggest model so far. The company says it is natively multimodal, with one trillion parameters and 49 billion active parameters.

On X, Mistral says ML4 is the strongest open-weight model from the US or Europe across aggregated benchmarks. The company also claims state-of-the-art performance in cyber defense, manufacturing, and finance, along with visual grounding results that surpass even closed frontier models.

Independent benchmark data gives a more measured view. In the Artificial Analysis Intelligence Index, which combines ten benchmarks across different domains, ML4 reaches 38 points. That is a major jump from Mistral Large 3, which scored 9 points, and Mistral Medium 3.5, which scored 14.

Still, the gap to the highest-ranked closed systems remains large. Claude Opus 5.5 (Max) leads the same index with 58 points. ML4 also edges past GLM-5.2 from Z.ai, which Mistral offers on its own platform.

Security is the main argument

The most distinctive part of Mistral’s pitch is not that ML4 beats every leading model overall. It does not, according to the Artificial Analysis Intelligence Index. The stronger claim is that ML4 can support cyber defense workflows that some competitors block.

Mistral says ML4 ranks among the top five models worldwide in the Artificial Analysis Cyber Index. Among open-weight models developed outside China, the company says it leads by a wide margin.

One test highlights the tension between capability and policy. The task asks a model to reproduce a real vulnerability in open-source software and then patch it. ML4 scores 82 percent, which the source says is the highest score of any model.

But the result is not only about raw skill. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse the task entirely. That means the test captures provider safety rules as well as model performance.

Mistral argues that defensive security often requires proving that a vulnerability exists before it can be fixed. In the company’s view, closed-model safety filters can block that legitimate work, while attackers can still jailbreak models. Mistral also says losing model access during an incident can become a security risk.

ML4 is designed to run in private clouds or on-premise. That matters for organizations that cannot rely only on outside providers when investigating sensitive code, infrastructure, or incidents.

The safety balance is still an open question

Mistral is also emphasizing refusal behavior. The company says ML4 refuses malicious cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm more often than any other open model.

In Lakera’s B3 AI Security Benchmark, Mistral says ML4 blocks 93.3 percent of attacks. At the same time, the source notes that Mistral does not explain how the model consistently separates legitimate vulnerability research from preparation for an attack.

That unresolved distinction is central to the product story. Mistral wants ML4 to be useful for real security teams without becoming broadly permissive for malicious use. During the preview period, the company is red-teaming the model with security firms, vetted partners, and government agencies. Those groups receive the same version with reduced moderation and expanded cyber capabilities.

Coding, vision, and enterprise tasks

ML4 also shows strength beyond cybersecurity. In the Artificial Analysis Coding Agent Index, it scores 49.8 percent, placing it ahead of Deepseek V4 Pro and Qwen3.8 Max.

Mistral also worked with Surge AI on a blind coding evaluation. Professional annotators rated code quality without knowing which model produced it. ML4 placed second out of five models with 3.74 out of 5 points, ahead of GLM-5.3 and Kimi K3. Claude Opus 5 led with 4.22 points.

For image understanding, Mistral says ML4 can analyze documents, charts, gigapixel satellite imagery, and technical drawings, including zooming in and checking details independently. On the Dense 200 benchmark, ML4 scores 42 percent, just ahead of GPT-6 Astra at 41 percent. According to an evaluation by Vals.ai, ML4 also beats GPT-6 Astra on legal and financial tasks.

Infrastructure, pricing, and what comes next

Mistral says ML4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European data centers. The preview runs on the same infrastructure. The company plans a European variant operated entirely by Mistral, independent of other service providers and under European law.

The training data covers more than 160 languages, including all official EU languages. Mistral says training involved companies in finance, manufacturing, logistics, pharma, shipping, and the public sector. The company used the same training and RL environment offered to customers through Mistral Forge.

For reinforcement learning post-training, Mistral uses about 3,000 GPUs. A single training run generates roughly 33 billion tokens per day. Mistral says the RL run behind the preview is still ongoing and has not plateaued, and it expects significant improvements over the coming weeks.

The compute expansion is backed by the Series D round of 3 billion euros, described in the source as the largest equity round ever raised by a European tech company. Mistral has also been building its European infrastructure for months, including an $830 million loan for a data center near Paris and plans for 200 megawatts of compute capacity in Europe by the end of 2027.

Some important details are still pending. The model documentation lists a fine-grained mixture-of-experts architecture with 1.05 trillion total parameters, a 1.6 billion parameter vision encoder, and a one million token context window. Mistral plans to release architecture details, the license, and post-training methods with the weights at the end of the month.

During the preview, Mistral charges $0.68 per million input tokens, $2.09 per million output tokens, and $0.07 for cached inputs. The documentation also lists prices at double those rates: $1.36 for input, $4.18 for output, and $0.14 for cached inputs.

ML4 is also intended to become the base for a new generation of specialized Mistral models. The company’s direction appears increasingly enterprise-focused: in May, Mistral renamed its chatbot Le Chat to Vibe and rebuilt it as a work tool.

The broader message is clear. Mistral wants ML4 to compete not only on benchmark scores, but on control, deployment flexibility, and security access. For customers weighing closed-model safety filters against hands-on cyber defense needs, that may be the more important contest.