Home/Models/DeepSeek V4.1 Flash
1M contextReasoningResponses API

DeepSeek V4.1 Flash on QuickSilver Pro

DeepSeek V4.1 Flash is DeepSeek's V4.1 Flash: a new model architecture that is faster and stronger than V4 Flash, with a 1M-token context and a native Responses API for Codex. QuickSilver Pro serves it at FP8 precision for $0.125 / $0.55 per million tokens — less than half of DeepSeek's own peak-hour list rate.

$0.125 input · $0.55 output per 1M tokens
ByRaullen Chai·Updated

At a glance

Context
1M tokens
Input / 1M
$0.125
Output / 1M
$0.55
Thinks by default
Yes

The newest DeepSeek Flash architecture for Codex, agentic coding, and long-context automation.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.125$0.55lowest-cost
DeepSeek list price (DeepSeek API)$0.30$1.2054% lower
OpenAI (GPT-4o-mini)$0.15$0.608% lower

When to use

Reach for V4.1 Flash when you want the latest DeepSeek Flash architecture — faster responses and stronger coding/tool use than V4 Flash — for Codex, agentic coding, multi-step tool use, and long-context automation. The official build supports both OpenAI Chat Completions and Responses API clients under a stable QSP model ID.

When to use something else

For a settled, lowest-cost workhorse, V4 Flash ($0.086/$0.173) is cheaper. For genuinely hard reasoning, escalate to V4 Pro. V4.1 Flash is text-only here — if you need image input, 37 of our 71 models accept it.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-compatible. One-line migration via base_url.

FAQ

V4.1 Flash is DeepSeek's newer Flash architecture — faster and stronger on coding and tool use. V4 Flash ($0.086/$0.173) remains the settled, lowest-cost option. Both think by default and expose the OpenAI Chat Completions and Responses APIs.

V4.1 Flash reasons by default. For direct non-thinking chat through Chat Completions, pass `reasoning: { enabled: false }`; Responses API clients can set `reasoning.effort` to control depth.

Yes. Use base_url=https://api.quicksilverpro.io/v1 with model="deepseek-v4.1-flash". QSP exposes both `/v1/chat/completions` and `/v1/responses`; Codex should use `wire_api = "responses"`.

Try DeepSeek V4.1 Flash with double credits — up to $50 in bonus credits

Get API Key