QuickSilver Pro vs OpenRouter
QuickSilver Pro sells Claude at half its previous price or less — Claude Opus 5.5 is 60% and Claude Sonnet 5.5 50% below OpenRouter's public per-token rates — and lists many open-weight models like GLM-5.3, Qwen 3.7 Plus and MiniMax M3 ~20% below. GPT-6.1 Sol and Grok 4.7 are at list price. Same OpenAI-compatible API, two-line migration. OpenRouter still wins on the long tail of community models.
At a glance
| Feature | QuickSilver Pro | openrouter |
|---|---|---|
| Models in catalog | 71 — GPT-6.1 Sol, Claude 5.5, Gemini 3.8, Grok 4.7, Qwen3.8, GLM-5.3, DeepSeek, Kimi K3, FLUX images | 300+ |
| Pricing on shared models | Claude 50–67% below; many open models ~20% below; GPT-6.1 Sol and Grok 4.7 at list | Baseline |
| OpenAI-compatible surface | Yes | Yes |
| Streaming / tools / json_schema | Yes — json_schema varies by model | Yes |
| usage.cost on responses | Yes (synthetic) | Yes |
| Per-key monthly spend limits | Yes | Yes |
| Closed frontier models (Claude, GPT-6.1, Gemini, Grok) | Yes — every Claude model 50–67% below (Opus 5.5 60%, Sonnet 5.5 50%), Gemini 15% below, GPT-6.1 and Grok at list | Yes |
| First top-up bonus | First top-up matched 100%, up to $50 (credited on your next top-up) | Limited free models |
| Minimum top-up | $5 | $10 |
Pricing (per million tokens, USD)
Competitor list prices, last checked 2026-09-30.
| Model | QSP input | QSP output | openrouter input | openrouter output | vs. list |
|---|---|---|---|---|---|
| Claude Opus 5.5 | $1.60 | $8.00 | $4.00 | $20.00 | 60% |
| Claude Sonnet 5.5 | $1.00 | $5.00 | $2.00 | $10.00 | 50% |
| GLM-5.3 | $1.12 | $3.52 | $1.40 | $4.40 | 20% |
| Gemini 3.8 Flash | $0.6375 | $3.1875 | $0.75 | $3.75 | 15% |
| Kimi K3 | $2.55 | $12.75 | $3.00 | $15.00 | ~15% |
| MiniMax M3 | $0.24 | $0.96 | $0.30 | $1.20 | 20% |
| Qwen 3.7 Plus | $0.256 | $1.024 | $0.32 | $1.28 | 20% |
| GPT-6.1 Sol | $2.00 | $10.00 | $2.00 | $10.00 | at list |
Migration - two lines
from openai import OpenAI
client = OpenAI(
base_url="https://api.quicksilverpro.io/v1",
api_key=os.environ["QSP_KEY"],
)
r = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[{"role": "user", "content": "Hi"}],
)FAQ
On many shared models, yes: Claude Opus 5.5 is 60% and Claude Sonnet 5.5 50% below OpenRouter's public per-token rates; GLM-5.3, Qwen 3.7 Plus, MiniMax M3 and MiMo V2.6 Pro ~20% below; Kimi K3 and Gemini 3.8 Flash ~15%. GPT-6.1 Sol and Grok 4.7 are at list price. Prices move — the table above has exact numbers as of October 2026.
Two lines in your OpenAI SDK setup: change base_url from openrouter.ai/api/v1 to api.quicksilverpro.io/v1 and swap the API key. Model ID mappings: deepseek/deepseek-v4-flash-0731 -> deepseek-v4-flash, deepseek/deepseek-v4-pro -> deepseek-v4-pro, qwen/qwen3.7-max -> qwen3.7-max, qwen/qwen3.7-plus -> qwen3.7-plus, qwen/qwen3.7-flash -> qwen3.7-flash, qwen/qwen3.6-plus -> qwen3.6-plus, qwen/qwen3.6-35b-a3b -> qwen3.6-35b, moonshotai/kimi-k2.6 -> kimi-k2.6, moonshotai/kimi-k2.7-code -> kimi-k2.7-code, moonshotai/kimi-k3 -> kimi-k3, z-ai/glm-5.3 -> glm-5.3, z-ai/glm-5.3-flash -> glm-5.3-flash, openai/gpt-oss-120b -> gpt-oss-120b, qwen/qwen3.8-27b -> qwen3.8-27b, qwen/qwen3.8-flash -> qwen3.8-flash-next, qwen/qwen3.8-omni-flash -> qwen3.8-omni-flash and z-ai/glm-5.2 -> glm-5.2.
If your workload needs Llama, Mistral, or the long tail of community models. The closed frontier is here too — GPT-6.1 Sol, Claude Opus 5.5 and Sonnet 5.5, Gemini 3.8, Grok 4.7 — on one key, with Claude Opus 5.5 60% and Sonnet 5.5 50% below OpenRouter, and Gemini priced 15% under Google's own list.
Yes for the shared models. Streaming, tool / function calling, and standard usage accounting all work through the official OpenAI SDK, and each response returns a synthetic usage.cost computed from the public per-token rate. json_schema strict mode is model-dependent: the Claude models do not support it, so a schema there is advisory and the enforced path is a tool with `strict: true`.
DeepSeek V4 Flash + Pro, Qwen 3.6/3.7, and Kimi K2.6 all default to chain-of-thought reasoning on OpenRouter — so a one-token "Hi" can return hundreds of reasoning tokens. For DeepSeek V4 and Kimi K2.6 we pass requests through unchanged: set `reasoning: { enabled: false }` to get non-thinking low-cost chat without the thinking overhead. The Qwen 3.6/3.7 models are served in non-thinking mode on QuickSilver Pro, so replies are direct; of these only Qwen 3.7 Max can opt back in, with `enable_thinking: true`.