Home/Models/Qwen3.8 Flash Next
1M contextResponses API

Qwen3.8 Flash Next on QuickSilver Pro

Qwen3.8 Flash Next is Alibaba's open-weight preview of the Qwen4 architecture — a 6B-active mixture-of-experts model with the family's full 1M-token context window and 131K max output. On QuickSilver Pro it's $0.12 input / $0.376 output per million tokens, ~20% below Alibaba's own list price of $0.15 / $0.47.

$0.12 input · $0.376 output per 1M tokens
ByRaullen Chai·Updated

At a glance

Context
1M tokens
Input / 1M
$0.12
Output / 1M
$0.376
Thinks by default
No

The cheapest way to run a next-generation Qwen on repo-scale context — a 6B-active MoE with a 1M-token window.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.12$0.376lowest-cost
Alibaba (qwen/qwen3.8-flash)$0.15$0.4720% lower
OpenAI (GPT-4o mini)$0.15$0.6037% lower

When to use

Reach for Qwen3.8 Flash Next on high-volume agent and coding workloads where cost-per-call and working-set size dominate: long-context refactors, document pipelines, and cache-heavy agents that re-read the same context. Its 6B-active routing keeps latency and price low while the Qwen4-preview stack targets agentic tasks, and the full 1M-token window is served on the route that actually enforces it.

When to use something else

For the hardest single-shot reasoning, step up to Qwen3.8 Max or GLM 5.3. As a preview-lineage release its behavior can shift with upstream revisions — pin your evals if you need a frozen target. Output is billed at $0.376 per million, so output-heavy generation should compare against DeepSeek V4 Flash and Qwen3.8 27B before committing.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-flash-next",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-compatible. One-line migration via base_url.

FAQ

Qwen3.8 Flash Next is an early open-weight release built on Alibaba's next-generation Qwen4 attention stack, tuned for long-context agentic workloads at a 6B-active MoE footprint. It ships ahead of the full Qwen4 line, so treat it as a fast-moving preview rather than a frozen production target.

Yes — the full 1,000,000-token window with up to 131K output tokens per call, and QuickSilver Pro serves the route that actually enforces those limits. Long-context calls bill the same flat $0.12 per million input tokens.

Yes — Qwen3.8 Flash Next is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="qwen3.8-flash-next". Streaming, tool calling, and usage.cost accounting all work.

Alibaba lists Qwen3.8 Flash Next at $0.15 input / $0.47 output per million tokens; QuickSilver Pro is $0.12 / $0.376, ~20% below on both legs. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `qwen/` provider prefix from the model ID.

Try Qwen3.8 Flash Next with double credits — up to $50 in bonus credits

Get API Key