Mistral Large 4 Benchmarks: Official and Independent Results

Mistral Large 4 scores 61.7% on DeepSWE v1.1, 93% on Cybench and 15.8% on Harvey’s Legal Agent Benchmark in Mistral AI’s launch results, while independent indexes place it mid-pack overall: 38 on the Artificial Analysis Intelligence Index (#64 of 225) and 48.05% on the Vals Index (#32 of 44). Its clearest lead is in cybersecurity, legal agents and visual grounding; GPT-6 Astra and Claude Opus 5.5 still score higher on general intelligence indexes.

Updated 2026-10-06

Mistral Large 4 at a glance

Developer
Mistral AI (Paris, France)
Released
October 6, 2026 (public preview)
Parameters
1.05T total, 49B active per token (Mixture of Experts)
Context window
524,288 tokens via the API
Max output
262,144 tokens
Input / output
Text and images in, text out
API price (sale)
$0.68 input / $2.09 output per 1M tokens (list $1.36 / $4.18)
API model name
mistral-large-4
Reasoning
reasoning_effort: "high" or "none"
Open weights
Scheduled for the end of October 2026

Sources: Mistral AI announcement and pricing page, OpenRouter model listing, Hugging Face model page. Checked Oct 6, 2026.

The short version

Mistral Large 4 is strongest at security work, legal and finance agents, and pointing at objects in images; it is average on broad intelligence indexes. That pattern holds across both Mistral’s own charts and the independent leaderboards.

  • Security: 82% on CyberGym-E2E and 93% on Cybench, ahead of every open-weight model Mistral compared against.
  • Legal: 15.8% on Harvey’s Legal Agent Benchmark, ranked #6 of 75 models by Vals AI, above GPT-6 Astra (5.4%).
  • Coding agents: 61.7% on DeepSWE v1.1, above DeepSeek V4 Pro 0813 (57) and Qwen3.8 Max (51), below Kimi K3 (68).
  • Vision: 42.0 on Dense200 grounding, slightly above GPT-6 Astra (41.5).
  • Overall: Artificial Analysis gives it 38, behind Claude Opus 5.5 (58) and GPT-6 Astra (53) but ahead of DeepSeek V4 Pro (36.0).

Coding and agent benchmarks (Mistral-reported)

On coding agents Mistral Large 4 beats DeepSeek V4 Pro and Qwen3.8 Max but trails GLM-5.3 and Kimi K3 on some tests. Mistral ran these with Artificial Analysis harnesses; each rival used its own preferred agent tool (for example Claude Code for Qwen3.8 Max and Codex for DeepSeek), so treat small gaps as close calls.

Mistral reports a combined Coding Agent Index of 49.8%. In the Surge AI blind human evaluation of coding answers, raters scored it 3.74 out of 5, ahead of Kimi K3 (3.59) and GLM-5.3 (3.60) but behind Claude Opus 5 (4.22).

BenchmarkMistral Large 4DeepSeek V4 Pro 0813Qwen3.8 MaxGLM-5.3Kimi K3
DeepSWE v1.161.7%57516168
Terminal-Bench 4.028.3%10174021
SWE-Atlas-QnA59.4%66625962
AutomationBench (657 workflows)59.9%56.757.2 (2.4T A95B)62.258.3
Surge AI human eval, coding (1–5)3.74——3.603.59
Source: Mistral AI Mistral Large 4 announcement (vendor-reported, preview model), checked Oct 6, 2026.

Cybersecurity and safety benchmarks

Cybersecurity is where Mistral Large 4 stands out most. On CyberGym-E2E, which asks a model to reproduce and then patch real vulnerabilities, it scores 82% while Mistral reports that Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse the tasks. On the Artificial Analysis Cyber Index it completes 50 tasks, against 33 for GPT-6 Astra and 29 for Claude Opus 5.5.

It also refuses clearly harmful requests: an average of 95.3 across JailbreakBench, StrongREJECT and AgentHarm, and 93.3% attack resistance on Lakera’s B3 agent security test. Mistral gives cyber leaders, vetted partners and state authorities a reduced-moderation build for red-teaming before the weights ship.

BenchmarkMistral Large 4Best rival shownOther rivals
CyberGym-E2E (reproduce + patch)82%MiMo-V2.6-Pro 79Grok 4.7 74, Kimi K3 58, GLM-5.3 29
Cybench (40 CTF tasks)93%Kimi K3 90DeepSeek V4 Pro 0813 88, GLM-5.3 85
AA Cyber Index (successes)50GLM-5.3-Flash 50Kimi K3 41, GPT-6 Astra 33, Claude Opus 5.5 29
Harmful cyber request refusal (avg)95.3Kimi K3 93.7GLM-5.3 87, DeepSeek V4 Pro 0813 77.7
B3 Agent Security (Lakera)93.3%GLM-5.3 93.3Kimi K3 88.1, DeepSeek V4 Pro 0813 85.2
Source: Mistral AI Mistral Large 4 announcement charts (vendor-reported), checked Oct 6, 2026.

Independent results: Artificial Analysis

Artificial Analysis scores Mistral Large 4 at 38 on its Intelligence Index v4.3.2, placing it #64 of 225 models and 8th among open-weight models. It ranks ahead of DeepSeek V4 Pro (36.0) and GLM-5.2 (33.7), and behind GLM-5.3 (45) and MiMo-V2.6-Pro (46).

It is fast but talkative. Output speed is 116.1 tokens per second against a median of 87, yet it used 200M output tokens to finish the index against a median of 81M, so long answers add to cost and wait time. Artificial Analysis currently lists it as a proprietary model because the weights are not public yet.

MetricMistral Large 4Reference point
Intelligence Index v4.3.238 (#64 of 225)Claude Opus 5.5 58, GPT-6 Astra 53, Gemini 4 Argon 53, GPT-6.1 Sol 52
Open-weight rank8thDeepSeek V4 Pro 36.0, GLM-5.2 33.7
Output speed116.1 tokens/s (#51)Median 87 tokens/s
Time to first token1.46 s—
Output tokens for the full index200MMedian 81M
Cost per index task (list price)$1.13—
Source: Artificial Analysis model page and trendingtopics.eu summary of AA data, checked Oct 6, 2026.

Independent results: Vals AI

Vals AI places Mistral Large 4 at 48.05% on the Vals Index v2.1, #32 of 44 models, almost level with Qwen 3.8 Max (48.27%) and DeepSeek V4 Pro 0813 (47.63%). Gemini 4 Argon leads that index at 68.90%, followed by Claude Opus 5.5 at 66.97%.

Its best Vals placements are legal and finance work; its weakest are formal proofs and puzzle-style science reasoning.

Vals benchmarkScoreRank
Vals Index v2.148.05%#32 of 44
Harvey’s Legal Agent Benchmark15.83%#6 of 75
BioMysteryBench67.04%#17 of 24
Vibe Code Bench 1-10014.26%#16 of 21
Terminal-Bench 4.022.73%#20 of 44
Finance Agent v254.68%#22 of 75
Tax Agent Bench63.29%#25 of 66
Vibe Code Bench v1.178.40%#25 of 109
IOI45.28%#28 of 41
Code Migration30.56%#38 of 74
ProofBench v1.110.00%#42 of 47
MysteryMechanism12.16%#24 of 24
Source: Vals AI Mistral Large 4 model page and Vals Index, checked Oct 6, 2026.

How to read these numbers

Vendor scores and independent scores answer different questions, and the gaps between them are worth knowing before you choose a model.

  • Terminal-Bench 4.0 differs by harness: 28.3% in Mistral’s Artificial Analysis run versus 22.73% on Vals AI.
  • Mistral calls its SciCode-Verified result state of the art among open-weight models, yet its own chart shows MiMo-V2.6-Pro (91.9), GLM-5.3 (92.5) and Qwen3.8 2.4T A95B (93.8) above its 91.8.
  • The internal expert win rate against GLM-5.3 is 68% for STEM on the domain chart and 69% on the breakdown chart; CAD is 62%, finance 50% and code 48%.
  • All Mistral numbers are for the preview build released on Oct 6, 2026; scores can change once the final weights ship.
  • Reasoning effort changes cost more than quality for some users: on Hacker News, Simon Willison reported little difference between reasoning_effort "high" and "none" in his test.

Try the model on your own prompts

Benchmarks show averages; your prompts are what matter. You can send Mistral Large 4 your own coding, legal or finance questions in the chat on our homepage, free for 3 messages a day without an account. Token costs for this level of performance are broken down on /api-pricing, and the /vs-deepseek-v4 and /vs-mistral-large-3 guides go head to head.

Frequently asked questions

More about Mistral Large 4