Pricing

Frontier and open models. One API. Priced below list.

Everything about price, in one place — no scrolling required.

Pay with

  • Card
  • Apple Pay
  • Google Pay
  • Alipay
  • WeChat Pay
  • Crypto (USDC)

Options at checkout depend on your country. Minimum top-up $5.

Sort
Filter

71 models. Sorted by Popular

Model
Context
Latency
Intelligence
InputUSD / 1M tokens
OutputUSD / 1M tokens
vs. list
GPT-6.1 SolNew
gpt-6.1-sol
GPT-6.1 Sol: near-Astra agentic coding and computer use at Sol pricing
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$2.00$4.00 >200K
$10.00$15.00 >200K
claude-opus-5-5
most capable Claude: agentic coding, knowledge work, computer use
ReasoningToolsVision
USD / 1M tokens
1M
—
58
$1.60$4.00
$8.00$20.00
−60%
claude-fable-5-1
deep reasoning, long-horizon agentic (5.1)
ReasoningToolsVision
USD / 1M tokens
1M
—
57
$4.00$10.00
$20.00$50.00
−60%
gpt-6-astra
long-horizon agentic coding, deep research, document work, image input
ReasoningToolsVision
USD / 1M tokens
1M
—
55
$10.00$20.00 >200K
$50.00$75.00 >200K
claude-sonnet-5-5
newest Sonnet: faster, fewer tokens per task, strong coding
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$1.00$2.00
$5.00$10.00
−50%
claude-haiku-5-5
newest Haiku: small, fast, 1M context, built for subagents and high-volume work
ToolsVision
USD / 1M tokens
1M
—
—
$0.05$0.25 >200K$0.10
$0.25$1.25 >200K$0.50
−50%
gemini-3.8-flash
most capable Flash for agents; 1M context
ReasoningToolsVision
USD / 1M tokens
1M
—
47
$0.6375$0.75
$3.1875$3.75
−15%
grok-4.7
long-running agentic coding, knowledge work, self-verification, image input
ReasoningToolsVision
USD / 1M tokens
500K
—
—
$2.00$4.00 >200K
$6.00$12.00 >200K
qwen3.8-max
Qwen 3.8 flagship, autonomous coding
ReasoningToolsVision
USD / 1M tokens
1M
—
47
$2.00
$6.00
glm-5.3
complex software engineering, long-horizon agents
ReasoningTools
USD / 1M tokens
1M
—
49
$1.12$1.40
$3.52$4.40
−20%
deepseek-v4.1-flash
newest DeepSeek Flash architecture, faster than V4 Flash
ReasoningTools
USD / 1M tokens
1M
—
—
$0.125$0.30
$0.55$1.20
−54–58%
kimi-k3
Multimodal reasoning, agentic coding
ReasoningToolsVision
USD / 1M tokens
1M
—
50
$2.55$3.00
$12.75$15.00
−15%
GPT-6 SolNew
gpt-6-sol
GPT-6 mid-flagship: strong reasoning, agentic coding, a rung below Astra
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$2.00$4.00 >200K
$10.00$15.00 >200K
GPT-6 LunaNew
gpt-6-luna
cheapest GPT-6 tier: high-volume chat, classification, lightweight agents
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$0.10$0.20 >200K
$0.50$0.75 >200K
deepseek-v4-pro
premium reasoning
ReasoningTools
USD / 1M tokens
1M
—
42
$0.70$1.32
$2.10$3.96
−47%
muse-spark-1.3
Flagship coding & agentic reasoning, 1M context
ReasoningToolsVision
USD / 1M tokens
1M
—
53
$1.00$1.25
$3.40$4.25
−20%
minimax-m3
long-horizon agentic coding, tool use
Tools
USD / 1M tokens
1M
—
36
$0.24$0.30
$0.96$1.20
−20%
Qwen3.8 Max PrimeNew
qwen3.8-max-prime
Qwen top-tier flagship: deepest reasoning, autonomous coding
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$4.00
$12.00
GLM 5.3 PrimeNew
glm-5.3-prime
top reasoning flagship, complex software engineering
ReasoningTools
USD / 1M tokens
1M
—
—
$2.24$2.80
$7.04$8.80
−20%
mimo-v2.6-pro
flagship open-weight coding, agentic workflows, image input
ToolsVision
USD / 1M tokens
1M
—
46
$0.348$0.435
$0.696$0.87
−20%
glm-5.3-flash
high-volume agents and coding at workhorse pricing
ReasoningTools
USD / 1M tokens
1M
—
46
$0.06$0.075
$0.20$0.25
−20%
deepseek-v4-flash
fast chat & coding, thinking on by default
ReasoningTools
USD / 1M tokens
1M
—
41
$0.086
$0.173
gpt-oss-120b
high-volume production agents, OpenAI open-weight MoE
ReasoningTools
USD / 1M tokens
131K
—
16
$0.041$0.15
$0.187$0.60
−69–73%
qwen3.8-27b
1M-context coding agents at small-model prices
Tools
USD / 1M tokens
1M
—
41
$0.34$0.425
$2.04$2.55
−20%
qwen3.8-flash-next
Cheapest next-gen Qwen: 6B-active Qwen4 preview on a 1M-token window
Tools
USD / 1M tokens
1M
—
—
$0.12$0.15
$0.376$0.47
−20%
qwen3.8-omni-flash
Qwen's first agentic omni model: image, audio + video understanding with tool use, at flash pricing
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$0.15
$0.47
glm-5.2
long-horizon agents, project-level coding
Tools
USD / 1M tokens
1M
—
—
$1.12$1.40
$3.52$4.40
−20%
nemotron-3-ultra
NVIDIA's 550B open flagship: deep reasoning and agent orchestration
Tools
USD / 1M tokens
262K
—
—
$0.50
$2.20
nemotron-3.5-lightning
fast, tool-heavy agents and high-volume automation
Tools
USD / 1M tokens
262K
—
—
$0.066
$0.176
qwen3.7-max
Qwen 3.7 flagship, agent / coding
ReasoningTools
USD / 1M tokens
1M
—
—
$1.25$1.475
$3.75$4.425
−15%
qwen3.7-flash
fast multimodal agents, visual coding, search
ToolsVision
USD / 1M tokens
1M
—
—
$0.024$0.03
$0.104$0.13
−20%
kimi-k2.6
Opus-class agentic / planning
ReasoningToolsVision
USD / 1M tokens
256K
—
—
$0.5472
$2.728
muse-spark-1.2
Coding-focused reasoning, agentic dev
ReasoningToolsVision
USD / 1M tokens
1M
—
47
$1.00$1.25
$3.40$4.25
−20%
muse-glimmer-30b
Budget agents, everyday coding, open weights
ReasoningToolsVision
USD / 1M tokens
131K
—
24
$0.28$0.35
$1.20$1.50
−20%
claude-opus-5
demanding reasoning, end-to-end coding, visual analysis, long-horizon agents
ReasoningToolsVision
USD / 1M tokens
1M
—
54
$2.00$5.00
$10.00$25.00
−60%
GPT-5.6 Sol
gpt-5.6-sol
complex reasoning, agentic coding, long-horizon tasks
Tools
USD / 1M tokens
1M
—
51
$1.60$2.00
$8.00$10.00
−20%
hy4-preview
coding & agentic workflows, reasoning (preview)
ReasoningTools
USD / 1M tokens
1M
—
—
$0.6672$0.834
$2.0008$2.501
−20%
hy3
general-purpose coding, agentic workflows
Tools
USD / 1M tokens
262K
—
—
$0.091
$0.363
mistral-large-4
Mistral's frontier multimodal flagship: coding, agents and 524K context
ReasoningToolsVision
USD / 1M tokens
524K
—
—
$0.68
$2.09
mimo-v2.6-flash
fast low-cost coding, agents, image input
ToolsVision
USD / 1M tokens
1M
—
—
$0.112$0.14
$0.224$0.28
−20%
mimo-v2.5
cost-efficient everyday coding, agentic workflows
Tools
USD / 1M tokens
1M
—
33
$0.112
$0.224
qwen3.6-plus
thinks-by-default flagship
Tools
USD / 1M tokens
1M
—
—
$0.26$0.325
$1.56$1.95
−20%
qwen3.7-plus
Qwen 3.7 agent flagship, long-horizon coding
ToolsVision
USD / 1M tokens
1M
—
—
$0.256$0.32
$1.024$1.28
−20%
qwen3.6-35b
long-context RAG, 35B MoE
Tools
USD / 1M tokens
262K
—
—
$0.112$0.14
$0.80$1.00
−20%
kimi-k2.7-code
Long-horizon agentic coding
ReasoningTools
USD / 1M tokens
256K
—
—
$0.584$0.73
$2.80$3.50
−20%
GPT-5.6 Terra
gpt-5.6-terra
everyday coding, reasoning, balanced agentic
Tools
USD / 1M tokens
1M
—
47
$1.60$2.00
$9.60$12.00
−20%
GPT-5.6 Luna
gpt-5.6-luna
high-volume chat, classification, lightweight agentic
Tools
USD / 1M tokens
1M
—
43
$0.16$0.20
$0.96$1.20
−20%
Grok 4.5
grok-4.5
coding, knowledge work, STEM
ReasoningTools
USD / 1M tokens
500K
—
45
$1.60$2.00
$4.80$6.00
−20%
grok-4.6
frontier coding, knowledge work, STEM, visual analysis
ReasoningToolsVision
USD / 1M tokens
500K
—
51
$2.00$4.00 >200K
$6.00$12.00 >200K
claude-fable-5
deep reasoning, long-horizon agentic
ReasoningToolsVision
USD / 1M tokens
1M
—
53
$4.00$10.00
$20.00$50.00
−60%
claude-opus-4-8
top-tier reasoning, coding, agentic
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$2.00$5.00
$10.00$25.00
−60%
claude-sonnet-5
newest Sonnet, stronger reasoning & coding, below list price
ReasoningToolsVision
USD / 1M tokens
1M
—
45
$1.00$3.00
$5.00$15.00
−67%
claude-haiku-4-5
fast, low-cost, high-volume tasks
ToolsVision
USD / 1M tokens
200K
—
—
$0.40$1.00
$2.00$5.00
−60%
gemini-3.7-flash
multimodal Flash for agents
ReasoningToolsVision
USD / 1M tokens
1M
—
45
$0.6375$0.75
$3.1875$3.75
−15%
gemini-3.6-flash
current general-purpose Flash GA
ReasoningToolsVision
USD / 1M tokens
1M
—
40
$0.6375$0.75
$3.1875$3.75
−15%
gemini-3.5-flash-lite
current low-cost, high-volume workloads
ToolsVision
USD / 1M tokens
1M
—
28
$0.255$0.30
$2.125$2.50
−15%
gemini-3.5-flash
next-gen Flash GA
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$1.275$1.50
$7.65$9.00
−15%
gemini-3.1-pro-preview
flagship reasoning
ReasoningToolsVision
USD / 1M tokens
1M
—
37
$1.70$2.00
$10.20$12.00
−15%
jev-1.13
Typed decisions, not chat: routing, classification and scoring with calibrated probabilities via /v1/systemone
System One
USD / 1M tokens
32K
—
—
$0.042
Free
flux.2-pro
flagship image generation
Image
—
—
—
—
$0.027 per image$0.031
flux.1-schnell
fast, high-volume image drafts
Image
—
—
—
—
$0.003 per image
sdxl-turbo
cheapest, fastest image previews
Image
—
—
—
—
$0.003 per image
flux.2-klein
balanced open FLUX.2 image generation
Image
—
—
—
—
$0.02 per image
Gemini 3 Pro Image
gemini-3-pro-image
GA pro-grade image generation
VisionImage
USD / 1M tokens
1M
—
—
$1.70$2.00
$10.20$12.00
$0.11424 per image$0.134
−15%
GPT Image 2New
gpt-image-2
OpenAI image generation at a fixed ~1.5MP output
Image
—
—
—
—
$0.01 per image
gemini-3-flash-preview
legacy integrations migrating to 3.6 Flash
ReasoningToolsVision
USD / 1M tokens
1M
—
—
$0.425
$2.55
—
gemini-3.1-flash-lite
legacy integrations migrating to 3.5 Flash-Lite
ToolsVision
USD / 1M tokens
1M
—
—
$0.2125
$1.275
—
qwen-image-max
Alibaba's Qwen-Image flagship, sharp text rendering
Image
—
—
—
—
$0.10 per image
—
seedream-5.0-pro
ByteDance's flagship photorealistic image model
Image
—
—
—
—
$0.07 per image
—
seedream-4
fast, high-quality ByteDance image generation
Image
—
—
—
—
$0.055 per image
—
bria-fibo-1.5
trained on fully licensed, commercially safe data
Image
—
—
—
—
$0.055 per image
—

Prices are exact per-token rates, not rounded — some carry more decimals than others. Each struck-through figure is that row's reference list price and links to the page that publishes it, so every line can be checked. "At list" means we sell at the reference price; "mixed" means we are below it on one side and above on the other. Every row checkable, every model verifiable →

Building an agent? Fund a key programmatically with USDC — no account, no card. x402 docs →

python
1# Two lines. That is the whole migration.
2from openai import OpenAI
3 
4client = OpenAI(
5 base_url="https://api.quicksilverpro.io/v1",
6 api_key="your-api-key",
7)
FAQ

Common questions

QuickSilver Pro is an OpenAI-compatible inference API with 71 models, one endpoint and one API key. Browse models →

Yes. To turn reasoning off for direct chat, send reasoning.enabled=false in the request body; in Python use extra_body.

Up to 20% below the standard published per-token list rate on most of the catalog, and 50–67% below on Claude. DeepSeek V4.1 Flash: $0.125 / $0.55. DeepSeek V4 Pro: $0.70 / $2.10. Kimi K3: $2.55 / $12.75. GLM 5.3: $1.12 / $3.52. Claude Opus 5.5: $1.60 / $8.00. GPT-6.1 Sol: $2.00 / $10.00. Grok 4.7: $2.00 / $6.00. Closed frontier models run through the same endpoint and the same key as the open-weight ones — Claude, GPT-6.1, Gemini and Grok included.

Yes. Change base_url to https://api.quicksilverpro.io/v1 in the official openai Python / Node / Swift SDKs. Streaming, tool calling, and usage.cost accounting all work out of the box. json_schema strict mode is model-dependent: the Claude models do not support it, so a schema there is advisory and the enforced path is a tool with strict: true.

New accounts get $0.05 free credit with no card, usable in browser chat and through the API with your account key on 5 starter models (GPT-6 Luna, Claude Haiku 5.5, DeepSeek V4.1 Flash, Qwen3.8 Flash Next and MiMo-V2.6-Flash); top up from $5 to unlock the rest. Browser chat also has a separate allowance of 5 free messages. And your first credit purchase is matched 100%, up to $50 in bonus credits. The bonus is added to your balance the next time you top up (any amount from $5): pay $20, then top up $5 later, and $20 of bonus arrives with it. A larger first payment still receives the $50 maximum. One per customer and payment card; standard pay-as-you-go after that.

Get your API key

Create an account, get your API key in 30 seconds.

Get API Key

Model launches & updates. A few emails a month.