How to Use the Mistral Large 4 API

To call Mistral Large 4, send a POST request to https://api.mistral.ai/v1/chat/completions with your Mistral API key as a Bearer token and "model": "mistral-large-4" (alias mistral-large-4-0). The API serves a 524,288-token context with up to 262,144 output tokens, takes text and images, and switches reasoning on or off with reasoning_effort.

Updated 2026-10-06

Mistral Large 4 at a glance

Developer
Mistral AI (Paris, France)
Released
October 6, 2026 (public preview)
Parameters
1.05T total, 49B active per token (Mixture of Experts)
Context window
524,288 tokens via the API
Max output
262,144 tokens
Input / output
Text and images in, text out
API price (sale)
$0.68 input / $2.09 output per 1M tokens (list $1.36 / $4.18)
API model name
mistral-large-4
Reasoning
reasoning_effort: "high" or "none"
Open weights
Scheduled for the end of October 2026

Sources: Mistral AI announcement and pricing page, OpenRouter model listing, Hugging Face model page. Checked Oct 6, 2026.

Mistral Large 4 API quick facts

Mistral Large 4 has been in public preview on Mistral Studio since Oct 6, 2026 (version label v26.10). The same model is available through OpenRouter and Vercel AI Gateway under their own IDs, all served by Mistral.

ItemValue
EndpointPOST https://api.mistral.ai/v1/chat/completions
Auth headerAuthorization: Bearer followed by your Mistral API key
Model ID (Mistral)mistral-large-4 (alias mistral-large-4-0)
Model ID (OpenRouter)mistralai/mistral-large-4-0
Model ID (Vercel AI Gateway)mistral/mistral-large-4
Context served by the API524,288 tokens (Mistral docs list 1M)
Max output262,144 tokens
Input / outputText and images in, text out
Reasoning switchreasoning_effort: "high" or "none"
Price per 1M tokens$0.68 in / $2.09 out on sale (list $1.36 / $4.18)
Source: Mistral Docs model card and reasoning guide, OpenRouter endpoints API, Vercel AI Gateway, checked Oct 6, 2026.

Step 1: get a key and make your first call

Create an API key in Mistral Studio (console.mistral.ai), store it in an environment variable, and send a chat request. The example below is the same request shape Mistral shows on the model card, with reasoning switched on.

# MISTRAL_KEY holds your Mistral Studio API key (set it in your shell)
curl https://api.mistral.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MISTRAL_KEY" \
  -d '{
    "model": "mistral-large-4",
    "messages": [{"role": "user", "content": "Explain mixture-of-experts in two sentences."}],
    "reasoning_effort": "high"
  }'

Step 2: call it from Python or TypeScript

Mistral’s official SDKs expose the same parameters. In Python the model card imports the client from mistralai.client; in TypeScript the package is @mistralai/mistralai. Both read the key from the environment, so it never sits in your source code.

import os
from mistralai.client import Mistral

client = Mistral(api_key=os.environ["MISTRAL_KEY"])

response = client.chat.complete(
    model="mistral-large-4",
    messages=[{"role": "user", "content": "Write a SQL query that finds duplicate emails."}],
    reasoning_effort="none",
)
print(response.choices[0].message.content)

# TypeScript equivalent:
# import { Mistral } from '@mistralai/mistralai';
# const client = new Mistral({ apiKey: process.env.MISTRAL_KEY });
# const res = await client.chat.complete({ model: 'mistral-large-4', messages, reasoning_effort: 'high' });

reasoning_effort: "high" vs "none"

Mistral Large 4 is one hybrid model: the same model ID answers instantly or thinks first, depending on reasoning_effort. With "high", message.content comes back as a list of chunks, a thinking chunk followed by a text chunk. With "none", it comes back as a plain string with minimal thinking. OpenRouter lists high as the default effort and reasoning as optional.

For multi-turn conversations with reasoning on, Mistral’s docs say to send the full previous assistant message back in the history, including its thinking chunks. If your code only keeps the final text, store the whole message object instead.

Because thinking is generated text, "high" makes answers longer. Artificial Analysis measured 200M output tokens for its full test set against a median of 81M, so use "none" for extraction, rewriting and short lookups and keep "high" for code and multi-step problems.

Send an image

Mistral Large 4 reads images through a 1.6B-parameter vision encoder and answers in text; it cannot generate images. Pass the image as a content part next to your text. Mistral reports strong grounding results (42.0 on Dense200), so asking for object positions or chart values is a good use.

curl https://api.mistral.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MISTRAL_KEY" \
  -d '{
    "model": "mistral-large-4",
    "reasoning_effort": "none",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Read the values in this bar chart and return them as a table."},
        {"type": "image_url", "image_url": "https://example.com/chart.png"}
      ]
    }]
  }'

Function calling and JSON output

Mistral Large 4 supports function calling and structured outputs on both /v1/chat/completions and /v1/conversations. Define tools with a JSON schema and let the model pick one; OpenRouter also accepts tool_choice values none, auto, required or a specific function. For strict JSON replies, add response_format.

curl https://api.mistral.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MISTRAL_KEY" \
  -d '{
    "model": "mistral-large-4",
    "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
          "type": "object",
          "properties": {"city": {"type": "string"}},
          "required": ["city"]
        }
      }
    }],
    "tool_choice": "auto"
  }'

Every feature and where to call it

Mistral’s model card lists the features below for Mistral Large 4. The batch endpoint is also the cheapest route, at 50% off standard token prices.

FeatureEndpoint or parameter
Chat and reasoning/v1/chat/completions with reasoning_effort
Stateful conversations/v1/conversations
Function callingtools and tool_choice on chat or conversations
Structured outputs / JSONresponse_format
Image inputimage_url content parts in messages
Document QnAdocument_url content parts in messages
Prefixfinal assistant message with "prefix": true to fix how the reply starts
Batch jobs (−50%)/v1/batch
Agents and built-in tools/v1/agents
Sampling controls (OpenRouter)temperature, top_p, stop, seed, frequency and presence penalty, logprobs
Source: Mistral Docs Mistral Large 4 model card and OpenRouter endpoints API, checked Oct 6, 2026.

Context window: 524,288 tokens in practice

Plan for 524,288 input tokens. Mistral’s model card lists a 1M-token context, but OpenRouter, Vercel AI Gateway and Artificial Analysis all list 524,288 (512K), Vals AI lists 512K, and a models.dev pull request reports that the live API enforces 524,288. Trim or chunk longer inputs to stay inside that limit.

Output can reach 262,144 tokens per reply. Filling the whole 524,288-token context once costs about $0.36 at the current sale price.

Use it through OpenRouter or Vercel AI Gateway

If you already route traffic through OpenRouter, use the ID mistralai/mistral-large-4-0 on its OpenAI-compatible endpoint; the price is the same $0.68 / $2.09 per 1M tokens because Mistral is the only provider. On Vercel AI Gateway the ID is mistral/mistral-large-4, and the AI SDK call is streamText with that model string.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENROUTER_KEY" \
  -d '{
    "model": "mistralai/mistral-large-4-0",
    "messages": [{"role": "user", "content": "Summarize the GDPR in five bullet points."}]
  }'

# Vercel AI SDK (TypeScript):
# import { streamText } from 'ai';
# const result = streamText({ model: 'mistral/mistral-large-4', prompt: 'Hello' });

Where Mistral Large 4 is not available yet

As of Oct 6, 2026 you can reach Mistral Large 4 through Mistral Studio, OpenRouter and Vercel AI Gateway. Mistral has not announced it on Azure, Amazon Bedrock, Google Vertex AI or IBM watsonx, and self-hosting has to wait for the open weights, which Mistral plans for the end of October (see the Hugging Face page). To skip keys and code entirely, chat with Mistral Large 4 on our homepage.

Frequently asked questions

More about Mistral Large 4