Qwen3.8 Flash Next for long context
Qwen3.8 Flash Next is a fast, cost-efficient choice for retrieval and long-document workflows. Its 1M context window lets you include the retrieved material in one prompt, at $0.12 input / $0.376 output per 1M tokens.
Why it's a fit for RAG
Use the 1M catalog context limit to size your retrieved material. Input costs $0.12 per 1M tokens. Evaluate recall on your own documents, especially near the context limit.
RAG pipeline pattern
Retrieve the most relevant passages, include them in the prompt, and ask for an answer grounded in those passages. If the corpus exceeds 1M, summarize by section before answering.
Quickstart code
from openai import OpenAI
client = OpenAI(
base_url="https://api.quicksilverpro.io/v1",
api_key="sk-qsp-...",
)
document = open("annual-report.txt").read()
resp = client.chat.completions.create(
model="qwen3.8-flash-next",
messages=[
{"role": "system", "content": "Answer using only the provided document."},
{"role": "user", "content": f"Document:
{document}
Question: What was free cash flow in Q3?"},
],
max_tokens=500,
)
print(resp.choices[0].message.content)FAQ
1M is the catalog limit. For critical retrieval, combine vector search with prompts that put relevant passages first.
This model runs in non-thinking mode. Do not budget for a hidden reasoning trace.
Its catalog capabilities are text input, text output, streaming and tool calling, exposed through Chat Completions and Responses. Benchmark multi-tool workflows before production.