You cannot run Mistral Large 4 locally yet: Mistral AI plans to publish the weights at the end of October 2026, and even then the 1.05-trillion-parameter model needs about 1.05 TB of GPU memory at FP8 or about 525 GB at 4-bit for the weights alone, which means a multi-GPU server rather than a desktop PC.
Updated 2026-10-06
Mistral Large 4 at a glance
Developer
Mistral AI (Paris, France)
Released
October 6, 2026 (public preview)
Parameters
1.05T total, 49B active per token (Mixture of Experts)
Sources: Mistral AI announcement and pricing page, OpenRouter model listing, Hugging Face model page. Checked Oct 6, 2026.
Can you run Mistral Large 4 locally today?
No. As of October 6, 2026 the Hugging Face repo mistralai/Mistral-Large-4.0-1T05-A52B is listed as an upcoming release with no files. Mistral's launch post promises the weights by the end of the month; the Hugging Face page shows an ETA of October 31, and several outlets report October 27.
Mistral has not announced the license, the file precision (BF16, FP8 or a 4-bit format) or a recommended hardware list yet. Everything below is arithmetic from the published parameter count, so you can budget hardware now and confirm the exact file size on release day.
If you only want to use the model, you do not need any of this: you can chat with Mistral Large 4 on our homepage right now, 3 messages a day without an account and 15 credits a day with a free account.
Weight memory by precision
Weight memory is simple: parameters × bytes per parameter. Mistral Large 4 has 1.05 trillion total parameters (1.05e12), and in a Mixture-of-Experts model every expert must sit in memory even though only 49B parameters are used per token.
Real quantized files are slightly larger than the pure bit count because they also store scaling factors. For example, a format averaging 4.5 bits per weight gives 1.05e12 × 4.5 / 8 = 590.6 GB instead of 525 GB. The 1.6B-parameter vision encoder adds little: 1.6e9 × 2 bytes = 3.2 GB at BF16.
Precision
Bytes per parameter
Formula
Weights only
BF16 / FP16
2
1.05e12 × 2
≈ 2,100 GB (2.1 TB)
FP8 / INT8
1
1.05e12 × 1
≈ 1,050 GB
6-bit
0.75
1.05e12 × 0.75
≈ 788 GB
4-bit (NVFP4, Q4-class)
0.5
1.05e12 × 0.5
≈ 525 GB
3-bit
0.375
1.05e12 × 0.375
≈ 394 GB
2-bit
0.25
1.05e12 × 0.25
≈ 263 GB
Source: our arithmetic from the 1.05T parameter count in Mistral Docs and the Hugging Face repo name; weights only, checked Oct 6, 2026.
Which GPU setups fit Mistral Large 4
An 8× B200 node is the most comfortable single-node target for FP8; 8× H200 fits FP8 weights with only about 78 GB to spare; 8× H100 only works at 4-bit. The table compares total memory with each weight size and shows what is left for KV cache and runtime overhead.
Two nodes of 8× H200 are needed to hold BF16 weights. A 512 GB unified-memory workstation cannot hold the 4-bit weights and would need a 3-bit or smaller quantization.
Setup
Total memory
BF16 (2,100 GB)
FP8 (1,050 GB)
4-bit (525 GB)
1× 24 GB consumer GPU
24 GB
No
No
No (weights are about 22× its memory)
4× H200 141 GB
564 GB
No
No
Yes, about 39 GB left (very tight)
8× H100 80 GB
640 GB
No
No
Yes, about 115 GB left
8× H200 141 GB
1,128 GB
No
Yes, about 78 GB left (tight)
Yes, about 603 GB left
8× B200 192 GB
1,536 GB
No
Yes, about 486 GB left
Yes, about 1,011 GB left
16× H200 (two nodes)
2,256 GB
Yes, about 156 GB left
Yes, about 1,206 GB left
Yes
512 GB unified memory
512 GB
No
No
No (3-bit at 394 GB leaves about 118 GB)
Source: our arithmetic (GPU count × memory per GPU minus weight size), weights only, checked Oct 6, 2026.
KV cache: the memory you need on top of the weights
Plan memory beyond the weights for the KV cache, which stores keys and values for every token in every active conversation. It grows linearly with context length and with the number of users you serve at once.
The formula is: KV cache bytes = 2 (keys and values) × layers × KV heads × head dimension × bytes per value × tokens × concurrent sequences. Mistral has not published Mistral Large 4's layer count or attention layout yet; they will be in the model config when the weights ship, and then the exact per-token cost follows from this formula.
Example arithmetic for a cache costing 100 KB per token: one conversation filling the full 524,288-token context needs 524,288 × 100 KB ≈ 52.4 GB, while 8 users at 32,768 tokens each need 8 × 32,768 × 100 KB ≈ 26.2 GB. On 8× H200 at FP8 that leaves little of the 78 GB headroom for activations, so long contexts there call for an FP8 KV cache or shorter limits.
Halve the cache by storing it in FP8 instead of BF16 where your inference engine supports it.
Cap the context length you actually need; 32,768 tokens is 1/16 of the 524,288-token API window.
Limit concurrent sequences on tight setups instead of letting the server preallocate for many users.
Software support: vLLM, SGLang, llama.cpp and Ollama
As of October 6, 2026 no Mistral Large 4 support has been announced in vLLM, SGLang, llama.cpp or Hugging Face Transformers; a search of their GitHub issues and pull requests found no model-specific work. Ollama and LM Studio build on llama.cpp-style GGUF files, so they depend on that support landing first.
The previous generation shows the likely path. Mistral Large 3 shipped with FP8 weights, an NVFP4 variant and day-one vLLM support (vLLM 1.12.0 or newer with mistral_common 1.8.6 or newer), and its model card recommends a single node of B200s or H200s for FP8 and a single node of H100s or A100s for NVFP4.
Model
Total params
Active per token
FP8 weights
4-bit weights
Mistral Large 3
675B
41B
≈ 675 GB
≈ 338 GB
Mistral Large 4
1.05T
49B
≈ 1,050 GB
≈ 525 GB
DeepSeek V4 Pro (0813)
1.6T
49B
≈ 1,600 GB
≈ 800 GB
Source: Hugging Face model cards (Mistral Large 3, Mistral Large 4 repo), OpenRouter model listing for DeepSeek V4 Pro; memory is our arithmetic, checked Oct 6, 2026.
Why speed depends on the 49B active parameters
Total parameters set how much memory you need; active parameters set how much work each token costs. Mistral Large 4 activates 49B parameters per token (52B counting embeddings and output layers), so at FP8 each generated token reads roughly 49e9 × 1 byte = 49 GB of weights.
That is why a 1.05T MoE model can generate faster than its size suggests: Artificial Analysis measured 116.1 output tokens per second on the hosted API. Your own speed will be set by GPU memory bandwidth and how the experts are split across cards.
Prepare now, download on release day
Free up disk space first: the FP8 download alone is about 1.05 TB, and a BF16 copy would be about 2.1 TB. These commands check space and fetch the repo once Mistral publishes the files.
Use the model today in the chat on our homepage, no setup needed.
Build against the hosted API now (see /api and /api-pricing) and switch the endpoint to your own server later.
Follow /huggingface and /updates for the weights date, license and file formats.
# Weight size in GB for a given precision (bytes per parameter)
python3 -c "print(1.05e12 * 1 / 1e9, 'GB at FP8'); print(1.05e12 * 0.5 / 1e9, 'GB at 4-bit')"
# Make sure the target disk has room (about 1.05 TB for FP8)
df -h /models
# Download after the weights are published on Hugging Face
hf download mistralai/Mistral-Large-4.0-1T05-A52B --local-dir /models/mistral-large-4
Frequently asked questions
No. Even the 4-bit weights are about 525 GB, roughly 22 times the memory of a 24 GB consumer GPU. Mistral Large 4 needs a multi-GPU server; on a normal computer, use the chat on our homepage instead.
For weights alone: about 2,100 GB at BF16, 1,050 GB at FP8 and 525 GB at 4-bit (1.05e12 parameters × bytes per parameter). Add memory for the KV cache and runtime on top.
Only at 4-bit. 8× H100 80 GB gives 640 GB, which holds the 525 GB of 4-bit weights with about 115 GB left. FP8 weights (1,050 GB) do not fit.
Mistral says by the end of October 2026. The Hugging Face repo mistralai/Mistral-Large-4.0-1T05-A52B shows an ETA of October 31, and some press reports say October 27.
Not yet. As of October 6, 2026 none of vLLM, SGLang, llama.cpp or Transformers has announced Mistral Large 4 support. Mistral Large 3 had vLLM support on release day, so vLLM is the most likely first engine.
Not at 4-bit: those weights are about 525 GB. A 3-bit quantization (about 394 GB) would fit with roughly 118 GB left for the KV cache and the operating system.
Mistral Large 4 has 1.05T parameters versus 675B, so it needs about 1.56× the memory: about 1,050 GB at FP8 versus 675 GB. Large 3 FP8 fits on one 8× H200 node with room to spare; Large 4 FP8 fits with only about 78 GB left.