QuickSilver Pro system status
Services
Model availability
deepseek-v4-flashdeepseek-v4.1-flashdeepseek-v4-proqwen3.8-maxqwen3.8-max-primeqwen3.8-omni-flashqwen3.7-maxqwen3.7-plusqwen3.7-flashqwen3.6-plusqwen3.6-35bqwen3.8-27bqwen3.8-flash-nextkimi-k2.6kimi-k2.7-codekimi-k3muse-spark-1.3muse-spark-1.2muse-glimmer-30bglm-5.3glm-5.3-primeglm-5.3-flashglm-5.2nemotron-3-ultranemotron-3.5-lightninggpt-oss-120bgpt-6-astragpt-6-lunagpt-6.1-solgpt-6-solgpt-5.6-lunagpt-5.6-terragpt-5.6-solgrok-4.5grok-4.6grok-4.7minimax-m3mimo-v2.6-promimo-v2.6-flashmimo-v2.5hy4-previewhy3mistral-large-4jev-1.13claude-opus-5-5claude-opus-5claude-fable-5-1claude-fable-5claude-opus-4-8claude-sonnet-5-5claude-sonnet-5claude-haiku-5-5claude-haiku-4-5gemini-3.7-flashgemini-3.8-flashgemini-3.6-flashgemini-3.5-flash-litegemini-3.5-flashgemini-3.1-pro-previewgemini-3-pro-imagegemini-3-flash-previewgemini-3.1-flash-liteflux.2-proflux.1-schnellsdxl-turboflux.2-kleinqwen-image-maxseedream-5.0-proseedream-4bria-fibo-1.5gpt-image-2Roadmap - how we become a real inference company
Now - launched on a curated catalog
LiveCustomers save 20% today on a curated catalog at low list prices. The narrow operational surface is what keeps the gap honest, and the same surface scales straight into Phase 2.
Q2 2026 - our own inference stack on H100/H200
PlannedSelf-hosted serving on dedicated GPUs using SGLang + continuous batching, EAGLE-3 speculative decoding, FP8 quantization via DeepGEMM, and SageAttention / ThunderMLA custom kernels. At that point system_fingerprint becomes stable (it changes only when we rev the stack), and repeatable-seed workflows start working properly. Target: 30-50% below current prices on the DeepSeek V4 wave.
H2 2026 - colocated data center + AIDC partnerships
FutureMove from rented (Vast.ai) to self-owned or colocated racks. Partner with AI-datacenter operators where that makes sense. The goal is the lowest-cost reliable inference for open-source models on the planet - full stack, our engineering.
About this page
Backend and website rows run reachability checks from your browser. The API row summarizes the listed model checks reported by our backend. Model rows reflect a real 1-token probe sent server-side every 3 minutes from our backend. Historical bars show the results of recent probes stored in this browser's localStorage; cleared if you switch devices.
Public uptime tracking began 2026-04-16. For a contractual SLA and third-party-monitored history, contact us.