🔍 Read the full analysis: Beyond The US And China: Where Mistral Large 4 Fits In AI on ThorstenMeyerAI.com
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral’s Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a substantial increase from the company’s previous flagship but below leading US and Chinese models. The preview is available through Mistral’s API; the company says it plans to release model weights by the end of October, with the licence not yet published.
French AI company Mistral has released Large 4 as a research public preview, and results in the Artificial Analysis Intelligence Index v4.3.2 put it at 38.4—below leading models from US and Chinese labs, but well above Mistral’s previous flagship. The result strengthens Mistral’s position as a European model developer while leaving its standing against frontier competitors, its eventual open-weights terms and its suitability for demanding agent tasks unsettled.
Artificial Analysis reports that Large 4 scored 38.4 on its Intelligence Index. In the same version of the index, the source lists leading US models at scores above 52, including Anthropic’s Claude Opus 5.5 at 57.6 and OpenAI’s GPT-6 Astra at 52.7. Several Chinese models also score above Large 4, including Z.ai’s GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. These are results on one benchmark, not a complete measure of performance in every task or deployment.
The gain over Mistral’s earlier models is marked: the supplied source gives Mistral Large 3 a score of 9 and Medium 3.5 a score of 14 on the same index version. Large 4 has one trillion total parameters, with 49 billion active, accepts text and images, produces text, and has a 512,000-token context window. Mistral says reinforcement learning is ongoing, so its performance could change.
Large 4 is currently available through Mistral’s API as a research public preview. The company has promised model weights by the end of October, but the source says the licence has not been published. Listed API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; the source reports a 50% discount for the first two weeks. Those introductory terms should not be mistaken for the standard rate.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
A European Model With a Narrow Lead
Large 4 matters because it gives buyers and developers another significant model provider outside the US and China. A major performance gain from Mistral’s preceding models suggests the company is making progress, and the preview offers a way to test its capabilities before the promised weights arrive. But the benchmark does not show that Mistral has caught the leading labs: the highest listed scores remain well above 38.4, and Chinese models also sit ahead of it.
That distinction is relevant for organisations weighing regional supply, provider diversity or future access to model weights against performance and cost. The source estimates that Large 4 costs $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models score 41.8 and 39.5, respectively, on the cited index. This comparison comes from the supplied source’s benchmark-cost figures; actual spending will vary by workload, token use, and pricing arrangements.
The same source reports that Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. That is a raw figure with the comparison group described by the source, not a guarantee of token use in ordinary applications. If similar verbosity appears in a buyer’s own tasks, it could affect cost and latency. Organisations should measure their specific workloads rather than infer performance from the index alone.
As an affiliate, we earn on qualifying purchases.
How the Benchmark Frames Large 4
The Artificial Analysis Intelligence Index combines evaluations of agentic knowledge work, real-world work tasks, software workflows and coding, according to the supplied source. It is therefore especially relevant to claims about multi-step work, but it remains a benchmark assembled from selected tasks. Its score should not be read as a universal ranking for every model use, such as image understanding, retrieval, or a company’s internal workflows.
The index comparison also puts the “outside the US and China” description in perspective. The supplied source says Large 4 is the strongest model in that geographic framing, while arguing that few other labs in the region are competing at this level. That is the source’s interpretation of the competitive field, not a benchmark score in itself. Large 4’s listed position among models with weights available is also provisional: Mistral has not yet released its weights, and the source says its licence remains unpublished.
Claims about reliability need separate treatment from index results. The supplied source author says hands-on testing found confident false statements, but provides no testing protocol, sample size or independent verification. That observation is not an Artificial Analysis finding. The source also cites AA-Omniscience figures for other models, but those measures concern a different evaluation and should not be conflated with the Intelligence Index score.
“Reinforcement learning is still running, so scores may move.”
— Mistral, according to the supplied source
As an affiliate, we earn on qualifying purchases.
Weights, Licence and Reliability
Several details remain open. Mistral has not yet provided the weights or their licence, so developers cannot assess the eventual use conditions from the supplied information. The source does not specify an exact date for the release beyond “the end of October,” nor does it clarify whether that timetable could change.
The score could shift while reinforcement learning continues, and the source gives no later benchmark update. It also does not provide enough information to independently evaluate the author’s report of hallucinations: the number and type of test interactions, prompts and comparison method are unspecified. That account should be treated as a reported observation, not a controlled finding.
Price comparisons are also tied to the source’s stated prices and index-task estimates. Costs can vary with input-output mix, caching, discounts and how a task is implemented. The reported benchmark figures alone do not establish what a particular business will pay or whether Large 4 will be more or less effective on its own work.
As an affiliate, we earn on qualifying purchases.
The October Weights Release
The next milestone is Mistral’s promised model-weight release by the end of October. Developers will then be able to examine the release terms and test whether their intended deployment is permitted, subject to the licence Mistral publishes. Until then, Large 4 remains a research preview accessed through the company’s API.
Potential buyers can compare it with alternatives using their own tasks, including the accuracy, token consumption, latency and failure handling required for their applications. Further Artificial Analysis results may also clarify whether ongoing reinforcement learning changes the score. The supplied source does not confirm a separate release date for updated benchmark results.
text and image AI processing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral released Large 4 as a research public preview through its API. The model is multimodal for text and image input, produces text, and has a 512,000-token context window.
How did Large 4 score?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the supplied source. Several leading US and Chinese models scored higher on that index.
Are Large 4’s weights available now?
No. The source says Mistral has promised weights by the end of October. The model is currently available as an API preview, and the licence for the planned weights has not been published.
Does the benchmark prove Large 4 is unsuitable for agents?
No. The index includes agentic tasks, but one benchmark cannot determine performance in every deployment. The supplied source author raises concerns about cost, verbosity and observed hallucinations; those concerns should be checked against the intended workload, and the hallucination report lacks a disclosed test protocol.
How much does the API cost?
The source lists standard prices of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It also reports a 50% discount for the first two weeks; buyers should confirm current pricing and applicable terms with Mistral.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
