AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: A Standout Outside The US And China, With Agent Trade-Offs on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral has released Large 4 as a research preview, scoring 38.4 on the Artificial Analysis Intelligence Index. That marks a sharp rise from the company’s previous models, but the source’s comparison places it below current US and Chinese flagships and raises questions about cost, verbosity and reliability in agent workflows.

Mistral released Large 4 as a research preview, and the model scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, according to the supplied source. The result marks a significant improvement over Mistral’s previous models, but the same comparison ranks several current US and Chinese models higher, leaving Large 4’s value for demanding agent work an open question.

Large 4 is a one-trillion-parameter model with 49 billion active parameters. It accepts text and images, produces text, and has a 512,000-token context window. Mistral is offering it through its API as a research preview; the source says the company plans to release model weights at the end of October. The model’s licence has not yet been published.

In the cited Artificial Analysis comparison, Large 4 scored 38.4. The source reports that Mistral Large 3 scored 9 and Medium 3.5 scored 14 on the same index version. Those results suggest a substantial step forward for Mistral, while the table places Large 4 below the listed US frontier models and several Chinese models, including GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash. Rankings apply to the specific index version and model set cited; they do not establish performance on every use case.

The source lists API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. It says Mistral is offering a 50% discount for the first two weeks. Artificial Analysis reportedly estimates a cost of $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those models scored 41.8 and 39.5 respectively in the same comparison, according to the source.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4, a multimodal research-preview model whose Artificial Analysis score shows substantial progress but leaves it behind leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Progress Meets Agent Costs

Large 4’s score matters because the index cited by the source includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. The reported result therefore speaks partly to work involving multiple steps and tool use, rather than only short-form question answering. It does not, by itself, predict how the model will perform in a particular company’s workflow.

For buyers, the comparison makes price per completed task as relevant as the headline token rate. The source reports that Large 4 used 200 million output tokens across the index tasks, compared with a median of 81 million for comparable models. If those figures are representative of a buyer’s workload, lengthy answers could add cost and latency. Actual results will depend on prompts, task mix, system design and any discounts.

The release also offers a measure of European model development. The source describes Large 4 as the highest-scoring model from outside the United States and China in the cited comparison. That is a narrow distinction, not evidence that it matches the leading systems overall: the supplied table shows a 38.4 score against 57.6 for the listed leader. The model’s progress may matter to organisations seeking another provider, but its position and practical fit require more than a geographic label.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Large 4 Compares

The supplied source bases its rankings on Artificial Analysis Intelligence Index v4.3.2, which it identifies as the current version. It reports that the top listed US models scored from 51.8 to 57.6, while the highest-scoring listed Chinese models ranged from 39.5 to 44.8. Large 4’s score of 38.4 is below those entries, though ahead of some older models in the table.

The source characterizes the release as a large jump from Mistral’s earlier scores, while cautioning that comparisons depend on which competitors and versions are included. It also says reinforcement learning is still underway, so Large 4’s results may change. The model is not yet an open-weights release: until the planned weight release, customers are accessing it through Mistral’s API, and the licence remains unpublished.

“Reinforcement learning is still running.”

— Mistral, as reported in the supplied source

Amazon

large format light table

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Reliability and Final Scores

Several important details remain unsettled. The source says Mistral plans to publish Large 4’s weights at the end of October, but the licence has not been published. Until the weights and licence are available, prospective users cannot assess the release’s terms for local deployment from the information provided.

The reported benchmark score may also change because Mistral says reinforcement learning is ongoing. The source includes a tester’s observation of confident hallucinations, but this is an individual hands-on assessment, not a result from the Artificial Analysis index. The article supplies no testing protocol or sample size for that observation, so it should not be treated as a general error rate.

It is also unclear whether the reported task costs and output-token totals will hold across customer workloads. The figures come from the cited benchmark comparison, and the initial discount is time-limited. No independent enterprise deployment results, long-run reliability data or workload-specific cost estimates are supplied.

Amazon

multimodal AI model

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The October Weights Release

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. The licence will be important for organisations considering deployment, but the source gives no confirmed terms. Mistral’s ongoing reinforcement learning may also lead to revised model performance or new benchmark results.

For now, buyers can compare the research-preview API’s documented prices and benchmark figures with their own task requirements. A grounded evaluation would need to measure accuracy, tool-use success, latency and total cost on representative workflows, especially for long-running agents. The supplied source does not report when further independent evaluations or customer results will be available.

Amazon

AI model cost analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is Mistral’s one-trillion-parameter model with 49 billion active parameters, available as a research preview through the company’s API. It accepts text and images, returns text, and has a 512,000-token context window, according to the supplied source.

How did Large 4 score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The source’s comparison places several listed US and Chinese models higher, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5.

Is Large 4 open source or open weights now?

No. The source says it is currently a proprietary API research preview, with weights planned for release at the end of October. It also says the licence has not been published.

What are the reported API prices?

The listed standard prices are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source reports a 50% discount for the first two weeks; buyers should confirm current pricing with Mistral.

What should buyers know about agent use?

The cited analysis raises questions about cost, output volume and reliability for multi-step tasks. Its benchmark figures are not a substitute for testing Large 4 on a buyer’s own workflow, and the hallucination concern in the supplied source is a tester’s observation rather than a benchmark rate.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

China’s demographic decline is not the disaster many fear

Recent data suggests China’s population decline is less severe and more manageable than widespread fears, with implications for its economy and policy.

Web Search API

Cloudflare’s beta Web Search API connects AI applications to live web results through AI Gateway, with three providers available at launch.

Mobilisiert, Nicht Ausgegeben: Was Von Europas €200-Milliarden-KI-Offensive üBrig Bleibt

Die EU kündigt eine KI-Investitionsoffensive von €200 Mrd. an, doch nur ein Bruchteil ist echtes öffentliches Geld, der Rest ist unsicher und noch nicht investiert.

The Governance Quandaries Of Self-Watching Smart Cities

Exploring the governance issues, vendor lock-in, and societal risks of digital twins in smart cities amid ongoing debates and emerging models.