Home/Models/Mistral Large 4
524K contextMultimodalReasoningResponses API

Mistral Large 4 on QuickSilver Pro

Mistral Large 4 is Mistral AI's frontier multimodal model, released in preview on October 6, 2026: text and image input, a 524K-token context, and a build aimed at coding, reasoning and agentic work. QuickSilver Pro serves Mistral's own endpoint at $0.68 input / $2.09 output per million tokens — the same as the model's list price on OpenRouter, with no markup.

$0.68 input · $2.09 output per 1M tokens
ByRaullen Chai·Updated

At a glance

Context
524K tokens
Input / 1M
$0.68
Output / 1M
$2.09
Thinks by default
Yes

Mistral's flagship for coding and agents — text and image input, 524K context, and reasoning you can switch off per request.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.68$2.09—
OpenRouter list price (mistralai/mistral-large-4-0)$0.68$2.09same

When to use

Use Mistral Large 4 for agent loops, multi-file coding and long-document work where you want a current European flagship: the 524K-token context holds a large repository or a stack of reports, it reads screenshots and diagrams alongside text, and it supports function calling (including a forced tool choice) and JSON Schema structured output. It thinks before it answers by default; for quick, direct replies pass `reasoning: {"enabled": false}`. Cached input costs $0.07 per million tokens.

When to use something else

Reasoning is on by default, is thorough, and can spend thousands of output tokens before the answer starts, so turn it off for routine chat and set a generous `max_tokens` when you keep it on. For high-volume simple extraction a flash-tier model such as DeepSeek V4 Flash ($0.086/$0.173) costs far less. A single response is capped at 131,072 output tokens. This is a preview release served from Mistral's own endpoint, which does not offer zero data retention, and behavior may change before general availability.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-large-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-compatible. One-line migration via base_url.

FAQ

$0.68 per million input tokens, $2.09 per million output tokens, and $0.07 per million cached-input tokens — the same as the model's list price on OpenRouter, with no markup. Reasoning tokens are billed as output tokens.

Yes. By default it thinks before it answers; the reasoning trace comes back in a separate field, and its tokens count toward `max_tokens` and are billed at the output rate. Pass `reasoning: {"enabled": false}` in the request to turn thinking off for that call, which then bills no reasoning tokens.

Yes. It accepts text and images as input and returns text. It supports streaming, function calling (including a forced tool choice), and JSON Schema structured output. Send images as URLs or inline as base64 data URLs.

Set base_url=https://api.quicksilverpro.io/v1, use your QSP key, and set model="mistral-large-4". Both Chat Completions and Responses API clients use the same public model ID.

Try Mistral Large 4 with double credits — up to $50 in bonus credits

Get API Key