vs Small 4

Mistral Large 4 vs Mistral Small 4

Same 8 prompts, same settings, both models through the Mistral API on October 9, 2026. Large 4 did the better job on 6 of 8 tasks and tied on 2; Small 4 finished in 29.8 seconds for the whole set against 55.1 for Large 4, and cost 6 times less. Below is every task, who did it better and why.

Updated 2026-10-09 · By suifeng

Popular prompts

Open chat

Opens full-screen chat. Your chats are saved to your account.

Mistral Large 4 at a glance

Developer
Mistral AI (Paris, France)
Released
October 6, 2026 (public preview)
Parameters
1.05T total, 49B active per token (Mixture of Experts)
Context window
524,288 tokens via the API
Max output
262,144 tokens
Input / output
Text and images in, text out
API price (sale)
$0.68 input / $2.09 output per 1M tokens (list $1.36 / $4.18)
API model name
mistral-large-4
Reasoning
reasoning_effort: "high" or "none"
Open weights
Scheduled for the end of October 2026

Sources: Mistral AI announcement and pricing page, OpenRouter model listing, Hugging Face model page. Checked Oct 6, 2026.

Why we ran this

Our free chat gives every new account its first 3 replies on Mistral Large 4 and then switches to Mistral Small 4; Pro and credit packs stay on Large 4. We wanted to know, with real tasks rather than benchmark tables, when the bigger model earns its price: at list price Large 4 costs $1.36 per million input tokens and $4.18 per million output tokens, Small 4 $0.15 and $0.60, about nine times more on input and seven times more on output.

Method: eight prompts that look like actual work (not riddles), sent once each to mistral-large-4 and mistral-small-2603 through the chat completions API with reasoning_effort set to none, temperature 0.2 and a 1,500-token output cap. Time is wall-clock from one machine in Asia, so your absolute numbers will differ; the ratio between the two models is the point. One run per prompt, so treat small differences as noise.

The scoreboard

Large 4 wins where completeness and instruction-following matter; Small 4 matches it on well-defined conversions and is two to three times faster.

TaskMistral Large 4Mistral Small 4Better answerWhy
Contract clause risks13.3 s (first token 1.4 s), 607 tokens3.7 s (first token 0.6 s), 469 tokensLarge 4Five risks, each tied to the exact words; Small 4 found three and quoted whole sentences
Two-quarter comparison6.8 s (first token 0.8 s), 373 tokens3.1 s (first token 0.6 s), 375 tokensTieBoth got gross profit ($2.56M / $2.30M) and operating profit ($0.46M / $0.60M) right; Large 4 asked sharper questions
Python code review10.5 s (first token 0.8 s), 1042 tokens4.6 s (first token 0.6 s), 785 tokensLarge 4Both found the injection, the off-by-one and the leaked connection; Large 4 fixed them with context managers and named columns
Prices from a 2,800-token page3.2 s (first token 1.0 s), 234 tokens3.8 s (first token 0.7 s), 488 tokensLarge 4Large 4 listed the four models the prompt asked for; Small 4 listed Large 4 price tiers instead. Both priced the workload right
Messy text to JSON3.0 s (first token 1.0 s), 195 tokens1.9 s (first token 0.6 s), 201 tokensTieIdentical numbers down to the cent; Small 4 was faster
Translation with an idiom5.4 s (first token 1.9 s), 255 tokens3.3 s (first token 0.7 s), 279 tokensLarge 4Large 4 translated the idiom in all four languages; Small 4 left “heads up” in English in the Japanese version
Scheduling puzzle10.9 s (first token 0.9 s), 830 tokens8.7 s (first token 0.6 s), 1500 tokens , cut offLarge 4Large 4 found the single valid order A-D-B-C-E with a case analysis; Small 4 hit the 1,500-token cap halfway through
Rewrite marketing copy2.0 s (first token 1.2 s), 24 tokens0.8 s (first token 0.6 s), 38 tokensLarge 4Large 4: 22 words, no hype; Small 4 used “Boost” and “innovate daily”
Our own run, October 9, 2026, Mistral API, reasoning_effort none, temperature 0.2, max 1,500 output tokens, one run per prompt.

Where Large 4 was clearly better

The pattern across the six wins: Large 4 does what the prompt literally asked and keeps going until the job is complete; Small 4 answers the general shape of the question.

  • Contract clause: asked for every risk with the exact words that create it, Large 4 listed five risks (including the combined effect of the three termination phrases) and quoted the precise fragment for each, then wrote a full replacement clause with a 30-day notice and pro-rated payment. Small 4 listed three risks, quoted whole sentences, and left “[X days]” in its rewrite.
  • Code review: both caught the SQL injection, the loop that skips the first row and the unclosed connection. Large 4’s fix used context managers, sqlite3.Row with named columns and “is not None” so an empty status string still filters; Small 4 kept positional indexes and closed the connection with an “if conn in locals()” check.
  • Extraction from a long page: the prompt pasted our 2,800-token API pricing page and asked for every Mistral model with per-million prices. Large 4 returned the four models (Large 4, Large 3, Medium 3.5, Small 4). Small 4 returned ten rows of Large 4 price tiers (batch, priority, regional) and no other model. Both then priced 50M input and 5M output tokens correctly: $88.90 list, $44.45 on sale.
  • Translation with an idiom: Large 4 rendered “heads up” and “swamped” idiomatically in French, German, Spanish and Japanese. Small 4 was fine in the three European languages but left “heads up” in English inside the Japanese sentence and added a note about it.
  • Scheduling puzzle: five talks, five constraints. Large 4 worked through the four placements of the A-D block and found the single valid order, A-D-B-C-E, in 830 tokens. Small 4 explained the constraints at length and ran into the 1,500-token cap before finishing its second case.
  • Copy rewrite under constraints (60 words, no hype words): Large 4 returned 22 plain words. Small 4 returned 30 words that opened with “Boost team productivity” and ended with “innovate daily”.

Where Small 4 was just as good

Two tasks were ties, and on both Small 4 was faster. The invoice-to-JSON conversion came back identical from both models down to the cent (subtotal 2,580, tax 490.20, total 3,070.20), in 1.9 seconds from Small 4 against 3.0 from Large 4. The two-quarter finance comparison got the same gross profit and operating profit from both; Large 4’s three CFO questions were more specific (asking whether the margin drop is mix shift or pricing pressure), Small 4’s were adequate.

Speed: Small 4 started answering after 0.6 to 0.7 seconds on every task; Large 4 took 0.8 to 1.9 seconds to the first token and about two to three times longer to finish. For short, well-defined jobs the wait is the only thing you give up.

Have a question about this?

Ask Mistral Large 4 directly. Free account, 15 credits every day.

Ask now

What each answer cost

Token counts are the API’s own usage figures; dollars are at Mistral’s list prices (Large 4 was on a limited-time 50% sale when we ran this, which halves its column). On our site every answer in this test cost one credit on either model, because one credit covers up to about $0.0067 of model cost at the free plan’s rate; the price gap only shows up on long replies and long documents.

TaskLarge 4 tokens in / outLarge 4 costSmall 4 tokens in / outSmall 4 costRatio
Contract clause risks84 / 607$0.002797 / 469$0.00039x
Two-quarter comparison97 / 373$0.0017119 / 375$0.00027x
Python code review136 / 1042$0.0045149 / 785$0.00059x
Prices from a 2,800-token page2823 / 234$0.00483149 / 488$0.00086x
Messy text to JSON122 / 195$0.0010150 / 201$0.00017x
Translation with an idiom69 / 255$0.001281 / 279$0.00026x
Scheduling puzzle71 / 830$0.003683 / 1500$0.00094x
Rewrite marketing copy78 / 24$0.000294 / 38$0.00006x
All 8 tasks3480 / 3560$0.01963922 / 4135$0.00316x
Usage reported by the Mistral API on October 9, 2026; prices from Mistral’s pricing page the same day (Large 4 $1.36 / $4.18, Small 4 $0.15 / $0.60 per million tokens, list).

Which one to use

Use Small 4 when the output format is fixed and the input is short: converting text to JSON or a table, extracting fields you name explicitly, quick arithmetic on numbers you supply, short translations between major European languages. Use Large 4 when you will act on the answer without checking it: legal or contract review, code you will ship, anything that asks for “every” item in a long document, constraint puzzles, Japanese or other languages where idiom matters, and writing under strict rules.

On our chat the free plan runs your first 3 replies on Large 4 so you can see the difference on your own task, then switches to Small 4; a credit pack or Pro keeps every reply on Large 4. The pricing page lists both.

How to reproduce it

Both models are called the same way; only the model name changes. This is the request we used, minus the API key.

# MISTRAL_KEY holds your Mistral Studio API key (set it in your shell)
curl https://api.mistral.ai/v1/chat/completions \
  -H "Authorization: Bearer $MISTRAL_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-large-4",
    "messages": [{"role": "user", "content": "I am a freelance designer. List every risk in this clause for me, quoting the exact words that create each risk, and rewrite the clause so it is fair to both sides: ..."}],
    "reasoning_effort": "none",
    "temperature": 0.2,
    "max_tokens": 1500
  }'
# then the same call with "model": "mistral-small-2603"

FAQ

vs Small 4 questions

Keep reading

More about Mistral Large 4

Ask Mistral Large 4 something real

Paste a bug, a contract clause or a messy spreadsheet note. A free account gives you 15 credits every day.

Start chatting