The short version
Mistral Large 4 is strongest at security work, legal and finance agents, and pointing at objects in images; it is average on broad intelligence indexes. That pattern holds across both Mistral’s own charts and the independent leaderboards.
- Security: 82% on CyberGym-E2E and 93% on Cybench, ahead of every open-weight model Mistral compared against.
- Legal: 15.8% on Harvey’s Legal Agent Benchmark, ranked #6 of 75 models by Vals AI, above GPT-6 Astra (5.4%).
- Coding agents: 61.7% on DeepSWE v1.1, above DeepSeek V4 Pro 0813 (57) and Qwen3.8 Max (51), below Kimi K3 (68).
- Vision: 42.0 on Dense200 grounding, slightly above GPT-6 Astra (41.5).
- Overall: Artificial Analysis gives it 38, behind Claude Opus 5.5 (58) and GPT-6 Astra (53) but ahead of DeepSeek V4 Pro (36.0).