Mistral Large 4 is Mistral AI's new flagship. It has about one trillion parameters, it is trained from scratch in Europe, and its open weights are due at the end of October 2026. It isn't the smartest model available. What makes it notable is that it leads in a few areas where closed flagships either refuse or fall behind: defensive security work, locating objects in images, and legal research agents.
Mistral released it on October 6, 2026 as a public preview. Its nickname is "Le Chonk". This post covers what the model is, what it does well, where it falls short, and how to try it through AIHubMix.
At a glance
| Mistral Large 4 | |
|---|---|
| Developer | Mistral AI (France) |
| Released | October 6, 2026 |
| Status | Public preview |
| API ID | mistral-large-4 |
| Parameters | 1.05T total, about 50B active |
| Context window | 1M tokens |
| Input / output | Text and images in, text out |
| Reasoning | On by default, can be turned off |
| Price per 1M tokens | $1.36 input, $4.18 output |
| Open weights | Planned for late October |
On AIHubMix the model ID is mistral-large-4-0, which Mistral also accepts as an alias. Specs here follow Mistral's model card. Mistral has given both 49B and 52B as the active parameter count. The license for the open weights hasn't been announced.
It is a mixture-of-experts model, so only a small share of its parameters is active for each token. It has a 1.6B-parameter vision encoder, accepts up to 100 images per request, and covers more than 160 languages. Mistral trained it on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacentres, not on top of another company's base model.
What it does well
The figures below come from Mistral's launch post and from independent runs by Artificial Analysis and Vals AI. Mistral's own figures are self-reported, and all of them were measured on the preview while training was still running.
| Area | Benchmark | Large 4 | For comparison |
|---|---|---|---|
| Security | CyberGym-E2E | 82% | MiMo-V2.6-Pro 79% |
| Security | Cybench (40 CTFs) | 93% | not compared |
| Vision | DIOR-RSVG | 73% | GPT-6 Astra 68% |
| Vision | Dense200 | 42% | GPT-6 Astra 41% |
| Legal | Harvey Legal Agent | 15.8% | Kimi K3 12.9% |
| Finance | Finch | 67% | DeepSeek V4 Pro 67% |
Defensive security. CyberGym-E2E asks a model to reproduce a real vulnerability and then patch it. Large 4 has the top score in Artificial Analysis's run. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score close to zero because they refuse the task. It doesn't say yes to everything, though. Mistral reports that it refuses genuinely harmful cyber requests more often than any other open model it tested, and that it blocked 93.3% of attacks on Lakera's B3 prompt-injection benchmark.
Finding things in images. Large 4 can do visual grounding: it can mark where an object is, not just describe it. DIOR-RSVG tests this on satellite and aerial imagery, and Dense200 tests it on crowded scenes. Mistral's demos include scanning gigapixel satellite images and checking parts on engineering drawings.
Legal and finance work. On Harvey's legal agent benchmark, Vals AI ranks it sixth of 76 models and first among open models. On Finch, a finance benchmark, Mistral's chart shows it level with DeepSeek V4 Pro.
Artificial Analysis summed up the release this way: "France is back to having the most intelligent model from outside the US and China."
Where it falls short
Large 4 is not a frontier all-rounder:
- General intelligence. It scores 38 on Artificial Analysis's Intelligence Index. That ties GPT-6 Luna and sits behind the open models GLM-5.3 (45) and Kimi K3 (44).
- Terminal-heavy coding. It scores 22.7% on Terminal-Bench 4 in Vals AI's run, which puts it 21st of 45. Claude Opus 5.5 scores 65.2%.
- Maths proofs and long PDF reports. It scores 10% on ProofBench and 19% on GDP.pdf. GPT-6 Astra scores about 32% on GDP.pdf.
- Cost per task. By Artificial Analysis's count, one Intelligence Index task costs $1.13 at list price. Similarly capable open models such as GLM-5.3-Flash cost about $0.25.
How it compares with Large 3
| Large 3 | Large 4 | |
|---|---|---|
| Parameters | 675B total, 41B active | 1.05T total, about 50B active |
| Context | 256K | 1M |
| Images per request | 8 | 100 |
| Price per 1M tokens | $0.50 / $1.50 | $1.36 / $4.18 |
| Weights | Apache 2.0, available now | Due late October |
Large 4 costs about 2.7 times as much per token as Large 3. For simple, high-volume text work that Large 3 already handles well, the upgrade may not pay for itself.
Two things to know before you call it
Thinking is on by default. On Mistral's own API, reasoning_effort takes "none" or "high". Through AIHubMix, tests on October 9 found that the model thinks no matter what reasoning_effort says, "none" included, and the thinking tokens are billed as output. To turn thinking off, send "reasoning": {"effort": "none"} in the request body, as the sample below does. When thinking is on, Mistral's reasoning guide says to send it back on the next turn. Through AIHubMix it arrives in a separate reasoning_details field, and the next request accepts it as part of the assistant message.
It is a preview. Mistral's lifecycle policy allows silent updates to preview models and gives one month's notice before retiring one. Keep a small set of test prompts and rerun them from time to time.
Try it through AIHubMix
AIHubMix serves Large 4 through its OpenAI-compatible Chat Completions endpoint. If you haven't set up a key yet, the quick start takes a few minutes.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com/v1",
)
response = client.chat.completions.create(
model="mistral-large-4-0",
messages=[{
"role": "user",
"content": "Write a Sigma rule that flags PowerShell launched by Office apps, "
"and explain each condition in one line.",
}],
extra_body={"reasoning": {"effort": "none"}}, # turns thinking off
)
print(response.choices[0].message.content)
The answer is always a plain string in content. Leave out the reasoning field and the model thinks first. The thinking arrives as a list in message.model_extra["reasoning_details"], and its tokens count as output. With thinking on, give max_tokens plenty of headroom. In testing, a 300-token cap was used up by the thinking and returned an empty answer.
Who should try it
- Security teams doing vulnerability reproduction, patching, malware analysis, or detection rules that closed models decline.
- Teams working with images, such as satellite and aerial inspection, engineering drawings, and dense scenes, where boxes matter more than captions.
- Legal and finance builders who want an open model near the top of agent benchmarks.
- European organisations that need a model that is neither American nor Chinese, with an EU-hosted deployment and a path to self-hosting.
For general coding agents or cost-sensitive everyday work, compare it with other models first. The AIHubMix model list lets you run the same prompts against Large 4 and its alternatives with one key.
FAQ
What is Mistral Large 4?
It is Mistral AI's flagship model, released in public preview on October 6, 2026. It is a mixture-of-experts model with 1.05 trillion parameters, takes text and image input, and returns text.
Is Mistral Large 4 open source?
Not yet. Mistral plans to publish the weights at the end of October 2026, but hasn't announced the license. Large 3's Apache 2.0 license doesn't automatically carry over.
How much does it cost?
The list price is $1.36 per million input tokens, $0.14 per million cached input tokens, and $4.18 per million output tokens.
How long is the context window?
1M tokens, according to Mistral's model card. That is four times Large 3's 256K.
Is it better than GPT-6 or Claude?
Only in specific areas: defensive security tasks the closed models refuse, visual grounding, and Harvey's legal agent benchmark. On broad intelligence and agentic coding, the closed flagships remain well ahead.
Does it support reasoning?
Yes, and it is on by default. Through AIHubMix the thinking comes back in a separate reasoning_details field. To turn it off, send a reasoning object with effort set to none in the request body. Setting reasoning_effort to none doesn't turn it off on AIHubMix.
Can it read images?
Yes. It accepts up to 100 images per request and can mark where objects are located in them. It cannot generate images.
Sources
- Introducing Mistral Large 4 (Mistral AI)
- Mistral Large 4 model card (Mistral Docs)
- Mistral Large 4.0 on AIHubMix
- Mistral Large 4: France is home to the most intelligent model outside the US and China (Artificial Analysis)
- Mistral Large 4 benchmarks (Vals AI)



