To call Mistral Large 4, send a POST request to https://api.mistral.ai/v1/chat/completions with your Mistral API key as a Bearer token and "model": "mistral-large-4" (alias mistral-large-4-0). The API serves a 524,288-token context with up to 262,144 output tokens, takes text and images, and switches reasoning on or off with reasoning_effort.
Updated 2026-10-06
Mistral Large 4 at a glance
Developer
Mistral AI (Paris, France)
Released
October 6, 2026 (public preview)
Parameters
1.05T total, 49B active per token (Mixture of Experts)
Sources: Mistral AI announcement and pricing page, OpenRouter model listing, Hugging Face model page. Checked Oct 6, 2026.
Mistral Large 4 API quick facts
Mistral Large 4 has been in public preview on Mistral Studio since Oct 6, 2026 (version label v26.10). The same model is available through OpenRouter and Vercel AI Gateway under their own IDs, all served by Mistral.
Item
Value
Endpoint
POST https://api.mistral.ai/v1/chat/completions
Auth header
Authorization: Bearer followed by your Mistral API key
Model ID (Mistral)
mistral-large-4 (alias mistral-large-4-0)
Model ID (OpenRouter)
mistralai/mistral-large-4-0
Model ID (Vercel AI Gateway)
mistral/mistral-large-4
Context served by the API
524,288 tokens (Mistral docs list 1M)
Max output
262,144 tokens
Input / output
Text and images in, text out
Reasoning switch
reasoning_effort: "high" or "none"
Price per 1M tokens
$0.68 in / $2.09 out on sale (list $1.36 / $4.18)
Source: Mistral Docs model card and reasoning guide, OpenRouter endpoints API, Vercel AI Gateway, checked Oct 6, 2026.
Step 1: get a key and make your first call
Create an API key in Mistral Studio (console.mistral.ai), store it in an environment variable, and send a chat request. The example below is the same request shape Mistral shows on the model card, with reasoning switched on.
# MISTRAL_KEY holds your Mistral Studio API key (set it in your shell)
curl https://api.mistral.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MISTRAL_KEY" \
-d '{
"model": "mistral-large-4",
"messages": [{"role": "user", "content": "Explain mixture-of-experts in two sentences."}],
"reasoning_effort": "high"
}'
Step 2: call it from Python or TypeScript
Mistral’s official SDKs expose the same parameters. In Python the model card imports the client from mistralai.client; in TypeScript the package is @mistralai/mistralai. Both read the key from the environment, so it never sits in your source code.
import os
from mistralai.client import Mistral
client = Mistral(api_key=os.environ["MISTRAL_KEY"])
response = client.chat.complete(
model="mistral-large-4",
messages=[{"role": "user", "content": "Write a SQL query that finds duplicate emails."}],
reasoning_effort="none",
)
print(response.choices[0].message.content)
# TypeScript equivalent:
# import { Mistral } from '@mistralai/mistralai';
# const client = new Mistral({ apiKey: process.env.MISTRAL_KEY });
# const res = await client.chat.complete({ model: 'mistral-large-4', messages, reasoning_effort: 'high' });
reasoning_effort: "high" vs "none"
Mistral Large 4 is one hybrid model: the same model ID answers instantly or thinks first, depending on reasoning_effort. With "high", message.content comes back as a list of chunks, a thinking chunk followed by a text chunk. With "none", it comes back as a plain string with minimal thinking. OpenRouter lists high as the default effort and reasoning as optional.
For multi-turn conversations with reasoning on, Mistral’s docs say to send the full previous assistant message back in the history, including its thinking chunks. If your code only keeps the final text, store the whole message object instead.
Because thinking is generated text, "high" makes answers longer. Artificial Analysis measured 200M output tokens for its full test set against a median of 81M, so use "none" for extraction, rewriting and short lookups and keep "high" for code and multi-step problems.
Send an image
Mistral Large 4 reads images through a 1.6B-parameter vision encoder and answers in text; it cannot generate images. Pass the image as a content part next to your text. Mistral reports strong grounding results (42.0 on Dense200), so asking for object positions or chart values is a good use.
curl https://api.mistral.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MISTRAL_KEY" \
-d '{
"model": "mistral-large-4",
"reasoning_effort": "none",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Read the values in this bar chart and return them as a table."},
{"type": "image_url", "image_url": "https://example.com/chart.png"}
]
}]
}'
Function calling and JSON output
Mistral Large 4 supports function calling and structured outputs on both /v1/chat/completions and /v1/conversations. Define tools with a JSON schema and let the model pick one; OpenRouter also accepts tool_choice values none, auto, required or a specific function. For strict JSON replies, add response_format.
curl https://api.mistral.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MISTRAL_KEY" \
-d '{
"model": "mistral-large-4",
"messages": [{"role": "user", "content": "What is the weather in Paris?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}],
"tool_choice": "auto"
}'
Every feature and where to call it
Mistral’s model card lists the features below for Mistral Large 4. The batch endpoint is also the cheapest route, at 50% off standard token prices.
Feature
Endpoint or parameter
Chat and reasoning
/v1/chat/completions with reasoning_effort
Stateful conversations
/v1/conversations
Function calling
tools and tool_choice on chat or conversations
Structured outputs / JSON
response_format
Image input
image_url content parts in messages
Document QnA
document_url content parts in messages
Prefix
final assistant message with "prefix": true to fix how the reply starts
Batch jobs (−50%)
/v1/batch
Agents and built-in tools
/v1/agents
Sampling controls (OpenRouter)
temperature, top_p, stop, seed, frequency and presence penalty, logprobs
Source: Mistral Docs Mistral Large 4 model card and OpenRouter endpoints API, checked Oct 6, 2026.
Context window: 524,288 tokens in practice
Plan for 524,288 input tokens. Mistral’s model card lists a 1M-token context, but OpenRouter, Vercel AI Gateway and Artificial Analysis all list 524,288 (512K), Vals AI lists 512K, and a models.dev pull request reports that the live API enforces 524,288. Trim or chunk longer inputs to stay inside that limit.
Output can reach 262,144 tokens per reply. Filling the whole 524,288-token context once costs about $0.36 at the current sale price.
Use it through OpenRouter or Vercel AI Gateway
If you already route traffic through OpenRouter, use the ID mistralai/mistral-large-4-0 on its OpenAI-compatible endpoint; the price is the same $0.68 / $2.09 per 1M tokens because Mistral is the only provider. On Vercel AI Gateway the ID is mistral/mistral-large-4, and the AI SDK call is streamText with that model string.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENROUTER_KEY" \
-d '{
"model": "mistralai/mistral-large-4-0",
"messages": [{"role": "user", "content": "Summarize the GDPR in five bullet points."}]
}'
# Vercel AI SDK (TypeScript):
# import { streamText } from 'ai';
# const result = streamText({ model: 'mistral/mistral-large-4', prompt: 'Hello' });
Where Mistral Large 4 is not available yet
As of Oct 6, 2026 you can reach Mistral Large 4 through Mistral Studio, OpenRouter and Vercel AI Gateway. Mistral has not announced it on Azure, Amazon Bedrock, Google Vertex AI or IBM watsonx, and self-hosting has to wait for the open weights, which Mistral plans for the end of October (see the Hugging Face page). To skip keys and code entirely, chat with Mistral Large 4 on our homepage.
Frequently asked questions
Use "mistral-large-4"; Mistral’s model card also lists the alias "mistral-large-4-0". On OpenRouter the ID is mistralai/mistral-large-4-0 and on Vercel AI Gateway it is mistral/mistral-large-4.
Set reasoning_effort to "high" to get thinking chunks plus the answer, or to "none" for a plain-string answer with minimal thinking. It is the same model either way.
Mistral’s docs list 1M, but the API serves 524,288 tokens, the figure shown by OpenRouter, Vercel AI Gateway and Artificial Analysis. Build for 524,288 input tokens and up to 262,144 output tokens.
Yes, it accepts images as input through a 1.6B-parameter vision encoder and returns text. It does not generate images.
Yes. Function calling and structured outputs work on /v1/chat/completions and /v1/conversations, and OpenRouter also supports response_format and tool_choice.
Not yet. As of Oct 6, 2026 Mistral has announced it only on Mistral Studio, with OpenRouter and Vercel AI Gateway also serving it through Mistral.
Currently $0.68 per 1M input tokens and $2.09 per 1M output tokens on a temporary 50% sale, against a list price of $1.36 / $4.18. The API pricing page has a calculator for monthly estimates.