Home/Models/Qwen3.8 27B
1M contextResponses API

Qwen3.8 27B on QuickSilver Pro

Qwen3.8 27B is the fast open member of Qwen's newest 3.8 generation — a 27B with the family's full 1M-token context window and 131K max output. On QuickSilver Pro it's $0.34 input / $2.04 output per million tokens, ~20% below Alibaba's own list price of $0.425 / $2.55. It replies directly — no invisible thinking tokens on your bill.

$0.34 input · $2.04 output per 1M tokens
ByRaullen Chai·Updated

At a glance

Context
1M tokens
Input / 1M
$0.34
Output / 1M
$2.04
Thinks by default
No

Coding agents that need a huge working set cheaply — a fast 27B with a 1M-token context window.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.34$2.04lowest-cost
Alibaba (qwen/qwen3.8-27b)$0.425$2.5520% lower
OpenAI (GPT-4o mini)$0.15$0.60240% more expensive

When to use

Reach for Qwen3.8 27B when the working set, not raw difficulty, is the constraint: repo-scale coding agents, long-document analysis, and multi-file refactors that want the full 1M-token window without flagship prices. It benchmarks with models several times its price class, and real coding-agent traffic is already routing to it in production. Direct replies keep short turns cheap and latency low.

When to use something else

For the hardest reasoning at this window size, Qwen3.8 Max and GLM 5.3 are the stronger picks. If you never need the giant window, GLM 5.3 Flash ($0.06/$0.20) undercuts it on price with the same family-fresh 2026 training. Output is billed at $2.04 per million — output-heavy generation pipelines should compare against DeepSeek V4 Flash before committing.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-compatible. One-line migration via base_url.

FAQ

No — on QuickSilver Pro Qwen3.8 27B always replies directly, so you never bill invisible reasoning tokens. This route serves the direct-answer profile only; if a task needs deliberate step-by-step reasoning at a similar price, use GLM 5.3 Flash, where deep reasoning is always on.

Yes — the full 1,000,000-token window with up to 131K output tokens per call, and QuickSilver Pro serves the route that actually enforces those limits. Long-context calls bill the same flat $0.34 per million input tokens.

Yes — Qwen3.8 27B is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="qwen3.8-27b". Streaming, tool calling, and usage.cost accounting all work.

Alibaba lists Qwen3.8 27B at $0.425 input / $2.55 output per million tokens; QuickSilver Pro is $0.34 / $2.04, ~20% below on both legs. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `qwen/` provider prefix from the model ID.

Try Qwen3.8 27B with double credits — up to $50 in bonus credits

Get API Key