AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Is Mistral Large 4 Right For Your AI Agents? Consider Its Limits on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get office and shipping supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a substantial jump from its predecessor but below leading US and Chinese models. Its benchmark cost, high output volume and reported hallucination concerns may limit its fit for long-running AI agents; the model is still in preview, and Mistral says reinforcement learning is ongoing.

Mistral released Large 4 as a research public preview, and independent benchmark data places it at 38.4 on Artificial Analysis Intelligence Index v4.3.2. The score marks a large improvement over Mistral’s prior models, but remains below leading US and Chinese systems, raising questions about whether its price and performance suit businesses running multi-step AI agents.

Artificial Analysis’ current index ranks Large 4 below the leading US models and several Chinese models. The source’s table gives Anthropic’s Claude Opus 5.5 a score of 57.6, while the strongest listed Chinese model, GLM-5.3, scores 44.8. Large 4 is close to OpenAI’s smaller GPT-6 Luna, listed at about 38, and ahead of some earlier Chinese systems. These rankings are specific to Index v4.3.2; they do not establish how every model performs on every real-world task.

The model has 1 trillion total parameters, with 49 billion active, accepts text and images, produces text, and has a 512,000-token context window, according to the source. Mistral’s API offers it as a research public preview. The source reports standard pricing of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; a 50% discount is available for the first two weeks. Mistral has promised to release the weights at the end of October, but the licence was unpublished in the material provided.

The improvement over earlier Mistral models is substantial on the same index: Large 3 scored 9 and Medium 3.5 scored 14, compared with Large 4’s 38.4. Artificial Analysis data cited by the source also shows Large 4 used 200 million output tokens across index tasks, versus a median of 81 million for comparable models. The source calculates a cost of $1.13 per index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, both of which scored higher.

At a glance
analysisWhen: Released yesterday, according to the so…
The developmentMistral has released Large 4 as a research public preview, with benchmark data showing a marked improvement over earlier Mistral models but continued gaps in capability and cost efficiency for agent workloads.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Workflows Face Cost and Error Risks

The score matters to buyers because the index includes tasks designed to reflect agentic work, including knowledge work, software workflows and coding. A model used for an agent must often make several decisions in sequence, call tools and carry information between steps. A weakness at any step can affect what follows, so a benchmark gap may matter more in a long task than in a short exchange. The index alone, however, cannot predict the outcome of a particular company’s workflow.

Output volume also affects the economics of agents. If a model uses more tokens to complete comparable tasks, buyers may face higher costs and longer waits, even when the advertised per-token price looks manageable. The source’s task-cost figures suggest lower-cost alternatives scored higher on this benchmark, while a higher-scoring frontier model may cost more. Actual expenses will depend on the prompts, task mix, caching and provider pricing a buyer uses.

The source author also reports seeing confident false statements in hands-on use. That is an attributed observation, not a result from the Artificial Analysis index. In an agent system, an unsupported claim can become an input to later actions; teams would need to test factual reliability and add safeguards before relying on the model for consequential work.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Large Jump, Not a Frontier Lead

The comparison presented in the source makes two points at once. Large 4’s score rose sharply from Mistral’s previous entries, suggesting the company has narrowed some of its performance gap. Yet it remains below the top systems in the same index. Describing it as the most intelligent model from outside the United States and China may be accurate within that geographic framing, but it does not mean it leads the broader field or ranks above the Chinese models listed.

Availability also affects how customers can judge it. Large 4 is currently described as a proprietary API preview, with weights promised later and licence terms not yet published. The index comparison is a snapshot: Mistral says reinforcement learning is still in progress and scores could change. Buyers assessing deployment should distinguish the current preview’s measured performance from any future version or eventual open-weight release.

“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”

— ThorstenMeyerAI.com author

Amazon

AI agent performance monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results May Still Change

Large 4 remains a research public preview, and the source says Mistral’s reinforcement learning work is ongoing. It is not yet clear how much the benchmark score, output efficiency or reliability might change. The promised end-of-October weights and their licence are also not available in the source material, leaving open how developers will be able to inspect, modify or deploy them.

The source does not provide a full breakdown of Large 4’s performance across individual benchmark tasks, nor independently documented testing of the author’s hallucination observations. Benchmark costs may also vary with real workloads and provider conditions. The figures support comparisons on the cited index, but do not settle whether Large 4 is suitable for a particular agent, industry or risk level.

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch for Weights and Retesting

The next milestones are Mistral’s planned end-of-October weight release and publication of its licence terms, as reported by the source. Updated Artificial Analysis results could offer a clearer comparison if Mistral’s ongoing training changes the model’s performance. Until then, organisations considering the preview can test it against their own agent tasks, tracking completion rates, factual errors, token use, latency and total task cost against alternatives.

For higher-impact workflows, those trials should also measure how errors propagate across tool calls and whether verification or human review can catch them. The current data points to a significant improvement over earlier Mistral models, but it does not by itself establish that Large 4 is the best operational or economic choice.

Amazon

AI model cost management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source material. That score is a benchmark result, not a guarantee of performance on a specific task.

Is Mistral Large 4 available to run with its weights?

The source describes it as a research public preview on Mistral’s API. Mistral has promised weights for the end of October, but the source says the licence had not been published.

Why might its cost be a concern for AI agents?

The source reports a cost of $1.13 per Artificial Analysis index task and unusually high output-token use compared with the cited model median. Long-running agents may repeat model calls, making token use and task-level cost relevant; actual costs depend on the workload and pricing.

Does the benchmark prove that Large 4 hallucinates?

No. The source author reports seeing confident false statements during hands-on testing, but identifies that as a personal observation rather than an Artificial Analysis finding. Buyers would need to test reliability on their own tasks.

Could Large 4’s benchmark score change?

Yes. The source says Mistral reported that reinforcement learning is ongoing and that scores may move. The preview’s present results should not be treated as final.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Advancing Digital Financial Literacy In Finland

Finland is expanding initiatives to improve digital financial literacy, aiming to enhance consumer understanding and resilience in digital finance.

How ChatGPT’s Rank Monitoring Boosts Brand Visibility In AI Search

New ChatGPT rank monitor helps brands track AI-driven mentions and share-of-voice, transforming AI search visibility strategies amid rising AI adoption.

Latest Developments From NTSB On Miami B-767 Runway Crash

The NTSB has issued a new investigative update on the Miami B-767 runway excursion, clarifying key findings and ongoing uncertainties. Read the latest details here.

Stock Yards Bancorp Surges In Global Coverage

Stock Yards Bancorp experiences a significant increase in global media mentions, raising investor interest and market attention.