Qwen3.8 Flash Next on QuickSilver Pro
Qwen3.8 Flash Next is Alibaba's open-weight preview of the Qwen4 architecture — a 6B-active mixture-of-experts model with the family's full 1M-token context window and 131K max output. On QuickSilver Pro it's $0.12 input / $0.376 output per million tokens, ~20% below Alibaba's own list price of $0.15 / $0.47.
At a glance
The cheapest way to run a next-generation Qwen on repo-scale context — a 6B-active MoE with a 1M-token window.
Pricing comparison ($/1M tokens)
| Provider | Input | Output | vs QSP |
|---|---|---|---|
| QuickSilver Pro | $0.12 | $0.376 | lowest-cost |
| Alibaba (qwen/qwen3.8-flash) | $0.15 | $0.47 | 20% lower |
| OpenAI (GPT-4o mini) | $0.15 | $0.60 | 37% lower |
When to use
Reach for Qwen3.8 Flash Next on high-volume agent and coding workloads where cost-per-call and working-set size dominate: long-context refactors, document pipelines, and cache-heavy agents that re-read the same context. Its 6B-active routing keeps latency and price low while the Qwen4-preview stack targets agentic tasks, and the full 1M-token window is served on the route that actually enforces it.
When to use something else
For the hardest single-shot reasoning, step up to Qwen3.8 Max or GLM 5.3. As a preview-lineage release its behavior can shift with upstream revisions — pin your evals if you need a frozen target. Output is billed at $0.376 per million, so output-heavy generation should compare against DeepSeek V4 Flash and Qwen3.8 27B before committing.
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-flash-next",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-compatible. One-line migration via base_url.
FAQ
Qwen3.8 Flash Next is an early open-weight release built on Alibaba's next-generation Qwen4 attention stack, tuned for long-context agentic workloads at a 6B-active MoE footprint. It ships ahead of the full Qwen4 line, so treat it as a fast-moving preview rather than a frozen production target.
Yes — the full 1,000,000-token window with up to 131K output tokens per call, and QuickSilver Pro serves the route that actually enforces those limits. Long-context calls bill the same flat $0.12 per million input tokens.
Yes — Qwen3.8 Flash Next is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="qwen3.8-flash-next". Streaming, tool calling, and usage.cost accounting all work.
Alibaba lists Qwen3.8 Flash Next at $0.15 input / $0.47 output per million tokens; QuickSilver Pro is $0.12 / $0.376, ~20% below on both legs. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `qwen/` provider prefix from the model ID.