How to Run Mistral Large 4 Locally

You cannot run Mistral Large 4 locally yet: Mistral AI plans to publish the weights at the end of October 2026, and even then the 1.05-trillion-parameter model needs about 1.05 TB of GPU memory at FP8 or about 525 GB at 4-bit for the weights alone, which means a multi-GPU server rather than a desktop PC.

Updated 2026-10-06

Mistral Large 4 at a glance

Developer
Mistral AI (Paris, France)
Released
October 6, 2026 (public preview)
Parameters
1.05T total, 49B active per token (Mixture of Experts)
Context window
524,288 tokens via the API
Max output
262,144 tokens
Input / output
Text and images in, text out
API price (sale)
$0.68 input / $2.09 output per 1M tokens (list $1.36 / $4.18)
API model name
mistral-large-4
Reasoning
reasoning_effort: "high" or "none"
Open weights
Scheduled for the end of October 2026

Sources: Mistral AI announcement and pricing page, OpenRouter model listing, Hugging Face model page. Checked Oct 6, 2026.

Can you run Mistral Large 4 locally today?

No. As of October 6, 2026 the Hugging Face repo mistralai/Mistral-Large-4.0-1T05-A52B is listed as an upcoming release with no files. Mistral's launch post promises the weights by the end of the month; the Hugging Face page shows an ETA of October 31, and several outlets report October 27.

Mistral has not announced the license, the file precision (BF16, FP8 or a 4-bit format) or a recommended hardware list yet. Everything below is arithmetic from the published parameter count, so you can budget hardware now and confirm the exact file size on release day.

If you only want to use the model, you do not need any of this: you can chat with Mistral Large 4 on our homepage right now, 3 messages a day without an account and 15 credits a day with a free account.

Weight memory by precision

Weight memory is simple: parameters × bytes per parameter. Mistral Large 4 has 1.05 trillion total parameters (1.05e12), and in a Mixture-of-Experts model every expert must sit in memory even though only 49B parameters are used per token.

Real quantized files are slightly larger than the pure bit count because they also store scaling factors. For example, a format averaging 4.5 bits per weight gives 1.05e12 × 4.5 / 8 = 590.6 GB instead of 525 GB. The 1.6B-parameter vision encoder adds little: 1.6e9 × 2 bytes = 3.2 GB at BF16.

PrecisionBytes per parameterFormulaWeights only
BF16 / FP1621.05e12 × 2≈ 2,100 GB (2.1 TB)
FP8 / INT811.05e12 × 1≈ 1,050 GB
6-bit0.751.05e12 × 0.75≈ 788 GB
4-bit (NVFP4, Q4-class)0.51.05e12 × 0.5≈ 525 GB
3-bit0.3751.05e12 × 0.375≈ 394 GB
2-bit0.251.05e12 × 0.25≈ 263 GB
Source: our arithmetic from the 1.05T parameter count in Mistral Docs and the Hugging Face repo name; weights only, checked Oct 6, 2026.

Which GPU setups fit Mistral Large 4

An 8× B200 node is the most comfortable single-node target for FP8; 8× H200 fits FP8 weights with only about 78 GB to spare; 8× H100 only works at 4-bit. The table compares total memory with each weight size and shows what is left for KV cache and runtime overhead.

Two nodes of 8× H200 are needed to hold BF16 weights. A 512 GB unified-memory workstation cannot hold the 4-bit weights and would need a 3-bit or smaller quantization.

SetupTotal memoryBF16 (2,100 GB)FP8 (1,050 GB)4-bit (525 GB)
1× 24 GB consumer GPU24 GBNoNoNo (weights are about 22× its memory)
4× H200 141 GB564 GBNoNoYes, about 39 GB left (very tight)
8× H100 80 GB640 GBNoNoYes, about 115 GB left
8× H200 141 GB1,128 GBNoYes, about 78 GB left (tight)Yes, about 603 GB left
8× B200 192 GB1,536 GBNoYes, about 486 GB leftYes, about 1,011 GB left
16× H200 (two nodes)2,256 GBYes, about 156 GB leftYes, about 1,206 GB leftYes
512 GB unified memory512 GBNoNoNo (3-bit at 394 GB leaves about 118 GB)
Source: our arithmetic (GPU count × memory per GPU minus weight size), weights only, checked Oct 6, 2026.

KV cache: the memory you need on top of the weights

Plan memory beyond the weights for the KV cache, which stores keys and values for every token in every active conversation. It grows linearly with context length and with the number of users you serve at once.

The formula is: KV cache bytes = 2 (keys and values) × layers × KV heads × head dimension × bytes per value × tokens × concurrent sequences. Mistral has not published Mistral Large 4's layer count or attention layout yet; they will be in the model config when the weights ship, and then the exact per-token cost follows from this formula.

Example arithmetic for a cache costing 100 KB per token: one conversation filling the full 524,288-token context needs 524,288 × 100 KB ≈ 52.4 GB, while 8 users at 32,768 tokens each need 8 × 32,768 × 100 KB ≈ 26.2 GB. On 8× H200 at FP8 that leaves little of the 78 GB headroom for activations, so long contexts there call for an FP8 KV cache or shorter limits.

  • Halve the cache by storing it in FP8 instead of BF16 where your inference engine supports it.
  • Cap the context length you actually need; 32,768 tokens is 1/16 of the 524,288-token API window.
  • Limit concurrent sequences on tight setups instead of letting the server preallocate for many users.

Software support: vLLM, SGLang, llama.cpp and Ollama

As of October 6, 2026 no Mistral Large 4 support has been announced in vLLM, SGLang, llama.cpp or Hugging Face Transformers; a search of their GitHub issues and pull requests found no model-specific work. Ollama and LM Studio build on llama.cpp-style GGUF files, so they depend on that support landing first.

The previous generation shows the likely path. Mistral Large 3 shipped with FP8 weights, an NVFP4 variant and day-one vLLM support (vLLM 1.12.0 or newer with mistral_common 1.8.6 or newer), and its model card recommends a single node of B200s or H200s for FP8 and a single node of H100s or A100s for NVFP4.

ModelTotal paramsActive per tokenFP8 weights4-bit weights
Mistral Large 3675B41B≈ 675 GB≈ 338 GB
Mistral Large 41.05T49B≈ 1,050 GB≈ 525 GB
DeepSeek V4 Pro (0813)1.6T49B≈ 1,600 GB≈ 800 GB
Source: Hugging Face model cards (Mistral Large 3, Mistral Large 4 repo), OpenRouter model listing for DeepSeek V4 Pro; memory is our arithmetic, checked Oct 6, 2026.

Why speed depends on the 49B active parameters

Total parameters set how much memory you need; active parameters set how much work each token costs. Mistral Large 4 activates 49B parameters per token (52B counting embeddings and output layers), so at FP8 each generated token reads roughly 49e9 × 1 byte = 49 GB of weights.

That is why a 1.05T MoE model can generate faster than its size suggests: Artificial Analysis measured 116.1 output tokens per second on the hosted API. Your own speed will be set by GPU memory bandwidth and how the experts are split across cards.

Prepare now, download on release day

Free up disk space first: the FP8 download alone is about 1.05 TB, and a BF16 copy would be about 2.1 TB. These commands check space and fetch the repo once Mistral publishes the files.

  • Use the model today in the chat on our homepage, no setup needed.
  • Build against the hosted API now (see /api and /api-pricing) and switch the endpoint to your own server later.
  • Follow /huggingface and /updates for the weights date, license and file formats.
# Weight size in GB for a given precision (bytes per parameter)
python3 -c "print(1.05e12 * 1 / 1e9, 'GB at FP8'); print(1.05e12 * 0.5 / 1e9, 'GB at 4-bit')"

# Make sure the target disk has room (about 1.05 TB for FP8)
df -h /models

# Download after the weights are published on Hugging Face
hf download mistralai/Mistral-Large-4.0-1T05-A52B --local-dir /models/mistral-large-4

Frequently asked questions

More about Mistral Large 4