Qwen3.8 27B on QuickSilver Pro
Qwen3.8 27B is the fast open member of Qwen's newest 3.8 generation — a 27B with the family's full 1M-token context window and 131K max output. On QuickSilver Pro it's $0.34 input / $2.04 output per million tokens, ~20% below Alibaba's own list price of $0.425 / $2.55. It replies directly — no invisible thinking tokens on your bill.
At a glance
Coding agents that need a huge working set cheaply — a fast 27B with a 1M-token context window.
Pricing comparison ($/1M tokens)
| Provider | Input | Output | vs QSP |
|---|---|---|---|
| QuickSilver Pro | $0.34 | $2.04 | lowest-cost |
| Alibaba (qwen/qwen3.8-27b) | $0.425 | $2.55 | 20% lower |
| OpenAI (GPT-4o mini) | $0.15 | $0.60 | 240% more expensive |
When to use
Reach for Qwen3.8 27B when the working set, not raw difficulty, is the constraint: repo-scale coding agents, long-document analysis, and multi-file refactors that want the full 1M-token window without flagship prices. It benchmarks with models several times its price class, and real coding-agent traffic is already routing to it in production. Direct replies keep short turns cheap and latency low.
When to use something else
For the hardest reasoning at this window size, Qwen3.8 Max and GLM 5.3 are the stronger picks. If you never need the giant window, GLM 5.3 Flash ($0.06/$0.20) undercuts it on price with the same family-fresh 2026 training. Output is billed at $2.04 per million — output-heavy generation pipelines should compare against DeepSeek V4 Flash before committing.
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-compatible. One-line migration via base_url.
FAQ
No — on QuickSilver Pro Qwen3.8 27B always replies directly, so you never bill invisible reasoning tokens. This route serves the direct-answer profile only; if a task needs deliberate step-by-step reasoning at a similar price, use GLM 5.3 Flash, where deep reasoning is always on.
Yes — the full 1,000,000-token window with up to 131K output tokens per call, and QuickSilver Pro serves the route that actually enforces those limits. Long-context calls bill the same flat $0.34 per million input tokens.
Yes — Qwen3.8 27B is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="qwen3.8-27b". Streaming, tool calling, and usage.cost accounting all work.
Alibaba lists Qwen3.8 27B at $0.425 input / $2.55 output per million tokens; QuickSilver Pro is $0.34 / $2.04, ~20% below on both legs. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `qwen/` provider prefix from the model ID.