Models

Model catalog

One OpenAI-compatible API for open-source and frontier LLMs, plus Google's Gemini family for multimodal chat, reasoning, and image generation. Prices are listed per 1M tokens in USD, from the current catalog. The model ID is what you pass in the request body.

At a glance

Model IDContextInputOutputNotes
deepseek-v4-flash1M$0.086$0.173Official 0731 agent model. 1M context with Chat Completions and Responses for Codex.
deepseek-v4.1-flash1M$0.125$0.55official DeepSeek V4.1 Flash — new architecture, faster, 1M context, Responses API
deepseek-v4-pro1M$0.70$2.10Premium reasoning with 1M-token context. Maps to o3-mini at ~6× lower output cost.
qwen3.8-max1M$2.00$6.00Alibaba's Qwen 3.8 flagship, released 2026-08-03: a 2.4T-parameter MoE that Alibaba pitches at autonomous coding and long-horizon agents. Thinks by default — the gateway suppresses thinking by default; pass enable_thinking=true to opt into the reasoning trace.
qwen3.8-max-prime1M$4.00$12.00Qwen 3.8 top-tier flagship, deepest reasoning & autonomous coding, 1M context
qwen3.8-omni-flash1M$0.15$0.47Qwen's first agentic omni model: native image, audio and video understanding with tool use and reasoning, 1M context, at Alibaba's first-party price
qwen3.7-max1M$1.25$3.75Qwen 3.7 flagship. Agent-centric, coding & productivity tasks. Thinks by default — the gateway suppresses thinking by default; pass enable_thinking=true to opt into the reasoning trace.
qwen3.7-plus1M$0.256$1.024Alibaba's hosted multimodal agent flagship. Long-running coding/agent loops, 1M context. Served in non-thinking mode: direct replies, no reasoning trace.
qwen3.7-flash1M$0.024$0.104Fast multimodal agent model for vision, visual coding, search, tool use, and high-volume automation. Served in non-thinking mode.
qwen3.6-plus1M$0.26$1.561T-parameter MoE flagship. 1M context. Top Qwen on OpenRouter by token volume. Served in non-thinking mode: direct replies, no reasoning trace.
qwen3.6-35b262K$0.112$0.8035B/3B-active MoE. Long-context RAG and summarization with strong reasoning.
qwen3.8-27b1M$0.34$2.04Qwen's fast open 27B reasoning model with a 1M-token context window
qwen3.8-flash-next1M$0.12$0.376Qwen's open Qwen4-architecture preview: 6B-active MoE, fast and cheap, with a 1M-token context window
kimi-k2.6256K$0.5472$2.728Opus-class agentic / planning. Best fit when your eval picks Claude Opus.
kimi-k2.7-code256K$0.584$2.80Always-thinking K2 tuned for long-horizon agentic coding. Fixed sampling parameters; Moonshot reports ~30% less overthinking than K2.6.
kimi-k31M$2.55$12.75Moonshot's 2.8T open-weight multimodal reasoning flagship. Complex coding, knowledge work, and long-horizon agentic workflows.
muse-spark-1.31M$1.00$3.40Coding & agentic reasoning model, 1M context
muse-spark-1.21M$1.00$3.40Coding-focused reasoning model, 1M context
muse-glimmer-30b131K$0.28$1.20Compact agentic multimodal model, distilled from Muse Spark
glm-5.31M$1.12$3.52Z.ai latest reasoning flagship, 1M context, complex software engineering & long-horizon agents
glm-5.3-prime1M$2.24$7.04Z.ai GLM 5.3 Prime, top reasoning flagship, 1M context, complex software engineering
glm-5.3-flash1M$0.06$0.20Z.ai fast open-weights reasoning model, 1M context at workhorse pricing
glm-5.21M$1.12$3.52Z.ai's large-scale reasoning flagship. 1M context, built for long-horizon agent workflows and project-level software engineering. Served in non-thinking mode: direct replies, no reasoning trace.
nemotron-3-ultra262K$0.50$2.20550B-A55B open MoE for frontier reasoning and agent orchestration
nemotron-3.5-lightning262K$0.066$0.17630B-A3B open MoE for fast, tool-heavy agents
gpt-oss-120b131K$0.041$0.187OpenAI's open-weight 117B MoE, high-volume agents & reasoning, 131K context
gpt-6-astra1M$10.00 ($20.00 >200K)$50.00 ($75.00 >200K)OpenAI's GPT-6 flagship. Long-horizon agentic coding, deep research and document work; accepts image input.
gpt-6-luna1M$0.10 ($0.20 >200K)$0.50 ($0.75 >200K)OpenAI's cost-efficient GPT-6, high-volume chat & lightweight agentic work, 1M context
gpt-6.1-sol1M$2.00 ($4.00 >200K)$10.00 ($15.00 >200K)OpenAI's GPT-6.1 Sol, near-Astra agentic coding at Sol pricing, 1M context
gpt-6-sol1M$2.00 ($4.00 >200K)$10.00 ($15.00 >200K)OpenAI's GPT-6 mid-flagship, strong reasoning & agentic coding, 1M context
gpt-5.6-luna1M$0.16$0.96OpenAI's fast, cost-efficient GPT-5.6 tier. High-volume, latency-sensitive chat and classification. Reasoning model; pass reasoning.enabled=false to suppress the trace.
gpt-5.6-terra1M$1.60$9.60OpenAI's balanced GPT-5.6 mid-tier, between Luna and Sol. Everyday coding, reasoning and agentic work.
gpt-5.6-sol1M$1.60$8.00OpenAI's flagship GPT-5.6. Complex reasoning, coding and agentic workflows, especially on hard multi-step tasks.
grok-4.5500K$1.60$4.80xAI's frontier model for coding, knowledge work and STEM. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort.
grok-4.6500K$2.00 ($4.00 >200K)$6.00 ($12.00 >200K)xAI's frontier model for coding, knowledge work, STEM and image analysis; accepts image input. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort.
grok-4.7500K$2.00 ($4.00 >200K)$6.00 ($12.00 >200K)xAI's flagship for coding, agentic tasks and knowledge work, succeeding Grok 4.6. Built for long-running software engineering and self-verification; accepts image input. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort.
minimax-m31M$0.24$0.96MiniMax's open-weight frontier model. Text-only input, tuned for long-horizon agentic coding.
mimo-v2.6-pro1M$0.348$0.696Xiaomi's flagship open-weight MiMo-V2.6-Pro (1.02T MoE, 42B active) — top open-weights model on the Artificial Analysis Intelligence Index. Coding and agentic workflows; accepts image input. Runs in non-thinking mode (reasoning off).
mimo-v2.6-flash1M$0.112$0.224Xiaomi's fast, low-cost open-weight MiMo-V2.6-Flash (309B MoE, 15B active). Everyday coding and agents; accepts image input. Runs in non-thinking mode (reasoning off).
mimo-v2.51M$0.112$0.224Xiaomi's open-weight MiMo-V2.5. Cost-efficient everyday coding and agentic workflows.
hy4-preview1M$0.6672$2.0008Tencent's Hy4 (preview). Successor to Hy3 for coding and agentic workflows with reasoning.
hy3262K$0.091$0.363Tencent's Hy3. General-purpose coding and agentic workflows. Served in non-thinking mode.
mistral-large-4524K$0.68$2.09Mistral's frontier multimodal model for coding and agents, 524K context
jev-1.1332K$0.042FreeTypeSafe's System One model on POST /v1/systemone — not a chat model. Answers typed choice / score / noul questions about a JSON state with calibrated probabilities. Bills input only.
claude-opus-5-51M$1.60$8.00Anthropic's most capable Claude and the first of the Claude 5.5 family. Leads Opus 5 and Fable 5.1 on agentic coding, knowledge work and computer use, at a lower per-token price than Opus 5. Accepts image input. Anthropic does not accept a forced tool_choice on this model; QuickSilver Pro sends one as auto plus an instruction to call that tool, so it steers rather than guarantees the call.
claude-opus-51M$2.00$10.00Anthropic's Opus 5 for demanding reasoning, coding and long-horizon agentic work. Accepts image input.
claude-fable-5-11M$4.00$20.00Anthropic's Mythos-class model, version 5.1, for deep reasoning and long-horizon agentic work.
claude-fable-51M$4.00$20.00Anthropic's Mythos-class model for deep reasoning and long-horizon agentic work.
claude-opus-4-81M$2.00$10.00Anthropic flagship. Top-tier reasoning, coding and agentic workloads.
claude-sonnet-5-51M$1.00$5.00Anthropic's newest Sonnet, faster and fewer tokens per task, 1M context
claude-sonnet-51M$1.00$5.00Anthropic's newest Sonnet. Stronger reasoning and coding at a lower price than the previous mid-tier.
claude-haiku-5-51M$0.05 ($0.25 >200K)$0.25 ($1.25 >200K)Anthropic's newest small, fast model for subagents and high-volume work, 1M context
claude-haiku-4-5200K$0.40$2.00Fast, low-cost Anthropic tier for high-volume work. Does not emit a reasoning trace.
gemini-3.7-flash1M$0.6375$3.1875fast multimodal Flash for agents; promotional pricing through 2026-12-31
gemini-3.8-flash1M$0.6375$3.1875most capable Flash for agents; promotional pricing through 2026-12-31
gemini-3.6-flash1M$0.6375$3.1875Current general-purpose Flash GA. Recommended for new Flash integrations.
gemini-3.5-flash-lite1M$0.255$2.125Current low-cost Flash-Lite GA. Recommended for high-volume workloads.
gemini-3.5-flash1M$1.275$7.65Next-gen Flash GA from Google. 1M context, thinks by default. Sits between 3 Flash Preview and 3.1 Pro on capability and price.
gemini-3.1-pro-preview1M$1.70$10.20Google's flagship reasoning model. 1M context, thinks deeply. Preview API; semantics may shift before GA.
gemini-3-pro-image1M$1.70$10.20GA pro-grade image generation. Replaces the retired preview model ID.
gemini-3-flash-preview1M$0.425$2.55Legacy compatibility only. Migrate new and existing workloads to gemini-3.6-flash.
gemini-3.1-flash-lite1M$0.2125$1.275Legacy compatibility only. Migrate new and existing workloads to gemini-3.5-flash-lite.
flux.2-pro-—$0.027 per imageFlagship image generation via /v1/images/generations. Billed per generated image.
flux.1-schnell-—$0.003 per imageultra-fast open image generation, billed per image
sdxl-turbo-—$0.003 per imagefast open image generation (SDXL Turbo), billed per image
flux.2-klein-—$0.02 per imageopen FLUX.2 image generation, billed per image
qwen-image-max-—$0.1 per imageAlibaba's Qwen-Image flagship: high-fidelity generation with strong typography, billed per image
seedream-5.0-pro-—$0.07 per imageByteDance's Seedream 5.0 Pro: flagship photorealistic image generation, billed per image
seedream-4-—$0.055 per imageByteDance Seedream 4: fast, high-quality image generation, billed per image
bria-fibo-1.5-—$0.055 per imageBria FIBO 1.5: image generation trained on fully licensed data (commercially safe), billed per image
gpt-image-2-—$0.01 per imageOpenAI's GPT Image 2 at a fixed ~1.5MP standard output, billed per image

Which model should I use?

  • Code-first default - deepseek-v4.1-flash. A practical starting point for code, chat, and production agents. 1M context; input $0.125, output $0.55.
  • Hard reasoning - deepseek-v4-pro. Use for math, theorem proving, and difficult multi-step problems. 1M context; input $0.70, output $2.10.
  • Long-document RAG - qwen3.8-flash-next. A fast, cost-efficient choice for retrieval and long-document summarization. 1M context; input $0.12, output $0.376.
  • Agentic planning - glm-5.3. Use for project-level reasoning and agentic planning. 1M context; input $1.12, output $3.52.
  • Multimodal workloads - gemini-3.8-flash. Use when prompts include images as well as text. 1M context; input $0.6375, output $3.1875.
  • Image generation - gemini-3-pro-image. Use for pro-grade generated image output. 1M context; input $1.70, output $10.20.
  • High-volume short turns - gemini-3.5-flash-lite. Use for cost-sensitive routing, classification, and extraction. 1M context; input $0.255, output $2.125.
  • Routing & classification decisions - jev-1.13. A System One model, not chat: typed answers with calibrated probabilities on POST /v1/systemone. 32K context; input $0.042, output free.

Thinking vs. non-thinking

DeepSeek V4.1 Flash and V4 Pro emit a reasoning trace before the final answer. Even a short prompt can produce extra output tokens. For inexpensive non-thinking chat on these models, pass reasoning.enabled=false.

See quickstart for full code →

Some always-reasoning models, such as Grok 4.7, reject reasoning.enabled=false. Control depth with reasoning_effort instead, or use DeepSeek V4.1 Flash with reasoning disabled.

Claude models: structured output and tool choice

Claude models do not enforce response_format: {type: "json_schema"} — the request succeeds and the reply comes back as ordinary prose. When you need schema-shaped output from a Claude model, declare the schema as a tool with strict: true and read the tool-call arguments; that path is enforced.

On claude-opus-5-5, claude-sonnet-5-5 and claude-fable-5-1 Anthropic does not accept a forced tool_choice. QuickSilver Pro accepts one anyway and sends it upstream as "auto" plus an instruction to call that tool, so the model is steered to the tool rather than strictly forced — handle a reply that arrives without tool_calls. The other Claude models accept a forced tool choice natively.

System One models (typed decisions)

System One models — Jev 1.13 today — are not chat models. You send a JSON state plus typed questions (choice, score or noul) to POST /v1/systemone and get calibrated probabilities back from a single forward pass, with no generated text to parse. Use them for routing, classification, triage, scoring and guardrail decisions. They bill input tokens only — output is free — and a Chat Completions call to one is rejected.

System One docs →

Per-model deep dives