DeepSeek V4.1 Flash on QuickSilver Pro
DeepSeek V4.1 Flash is DeepSeek's V4.1 Flash: a new model architecture that is faster and stronger than V4 Flash, with a 1M-token context and a native Responses API for Codex. QuickSilver Pro serves it at FP8 precision for $0.125 / $0.55 per million tokens — less than half of DeepSeek's own peak-hour list rate.
At a glance
The newest DeepSeek Flash architecture for Codex, agentic coding, and long-context automation.
Pricing comparison ($/1M tokens)
| Provider | Input | Output | vs QSP |
|---|---|---|---|
| QuickSilver Pro | $0.125 | $0.55 | lowest-cost |
| DeepSeek list price (DeepSeek API) | $0.30 | $1.20 | 54% lower |
| OpenAI (GPT-4o-mini) | $0.15 | $0.60 | 8% lower |
When to use
Reach for V4.1 Flash when you want the latest DeepSeek Flash architecture — faster responses and stronger coding/tool use than V4 Flash — for Codex, agentic coding, multi-step tool use, and long-context automation. The official build supports both OpenAI Chat Completions and Responses API clients under a stable QSP model ID.
When to use something else
For a settled, lowest-cost workhorse, V4 Flash ($0.086/$0.173) is cheaper. For genuinely hard reasoning, escalate to V4 Pro. V4.1 Flash is text-only here — if you need image input, 37 of our 71 models accept it.
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-compatible. One-line migration via base_url.
FAQ
V4.1 Flash is DeepSeek's newer Flash architecture — faster and stronger on coding and tool use. V4 Flash ($0.086/$0.173) remains the settled, lowest-cost option. Both think by default and expose the OpenAI Chat Completions and Responses APIs.
V4.1 Flash reasons by default. For direct non-thinking chat through Chat Completions, pass `reasoning: { enabled: false }`; Responses API clients can set `reasoning.effort` to control depth.
Yes. Use base_url=https://api.quicksilverpro.io/v1 with model="deepseek-v4.1-flash". QSP exposes both `/v1/chat/completions` and `/v1/responses`; Codex should use `wire_api = "responses"`.