Home/Use cases/qwen3 for long-context
Use case

Qwen3.8 Flash Next for long context

Qwen3.8 Flash Next is a fast, cost-efficient choice for retrieval and long-document workflows. Its 1M context window lets you include the retrieved material in one prompt, at $0.12 input / $0.376 output per 1M tokens.

$0.12 / $0.376 per 1M tokens

Why it's a fit for RAG

Use the 1M catalog context limit to size your retrieved material. Input costs $0.12 per 1M tokens. Evaluate recall on your own documents, especially near the context limit.

RAG pipeline pattern

Retrieve the most relevant passages, include them in the prompt, and ask for an answer grounded in those passages. If the corpus exceeds 1M, summarize by section before answering.

Quickstart code

python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.quicksilverpro.io/v1",
    api_key="sk-qsp-...",
)

document = open("annual-report.txt").read()

resp = client.chat.completions.create(
    model="qwen3.8-flash-next",
    messages=[
        {"role": "system", "content": "Answer using only the provided document."},
        {"role": "user", "content": f"Document:
{document}

Question: What was free cash flow in Q3?"},
    ],
    max_tokens=500,
)
print(resp.choices[0].message.content)

FAQ

1M is the catalog limit. For critical retrieval, combine vector search with prompts that put relevant passages first.

This model runs in non-thinking mode. Do not budget for a hidden reasoning trace.

Its catalog capabilities are text input, text output, streaming and tool calling, exposed through Chat Completions and Responses. Benchmark multi-tool workflows before production.

Start with your own key

Get API Key