Function calling
Pass tools to the chat-completions endpoint and the model can request that you call them. The wire format matches OpenAI — same tool_calls in the message, same role: tool replies on the way back.
Best models for tool use
- DeepSeek V4.1 Flash — production default for tool-calling agents. Reliable JSON args, low latency; thinks by default so you may want
reasoning.enabled=falseto keep tool selection fast. - GLM 5.3 — agentic / planning workloads where tool chaining needs project-level reasoning. Reasoning is always on; use reasoning_effort to control depth.
- For pure tool calling, prefer a non-thinking setup (pass
reasoning.enabled=false) — an always-on chain-of-thought trace means you pay for tokens you don't need.
Python — single tool
import os, json
from openai import OpenAI
client = OpenAI(
base_url="https://api.quicksilverpro.io/v1",
api_key=os.environ["QSP_KEY"],
)
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
resp = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=messages,
tools=tools,
)
msg = resp.choices[0].message
if msg.tool_calls:
call = msg.tool_calls[0]
args = json.loads(call.function.arguments)
# ... call your tool ...
result = {"city": args["city"], "temp_c": 22, "conditions": "clear"}
messages.append(msg)
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
final = client.chat.completions.create(model="deepseek-v4.1-flash", messages=messages, tools=tools)
print(final.choices[0].message.content)Parallel tool calls
When the model wants to call multiple tools in one turn, it returns multiple entries in tool_calls. Resolve them in any order; reply with one role: tool message per call, each referencing the matching tool_call_id.
Streaming tool calls
With stream=true, tool calls arrive as deltas just like content. The function name is streamed once; arguments accumulate across chunks as JSON fragments. You typically buffer until finish_reason=tool_calls before executing.
Choosing a tool
tool_choice defaults to "auto": the model decides whether to call a tool. "none" forbids tool calls, and "required" demands at least one call — with parallel calls enabled it can still return several, so read every entry in tool_calls. Naming one function restricts the choice to that function.
Some models only decide for themselves. On claude-opus-5-5, claude-sonnet-5-5 and claude-fable-5-1 Anthropic does not accept a forced choice; QuickSilver Pro sends it as "auto" plus an instruction to call that tool, which steers the model but does not guarantee the call. Forcing a specific function is rejected by glm-5.3 and glm-5.3-flash. On those models, passing a single tool and asking for it in the prompt raises the odds but guarantees nothing — the reply can still come back as plain text with no tool_calls, so handle that case. When a call must happen, use a model that accepts "required" or a named function.
Strict mode
For tighter schemas, set strict: true inside the function object. The model is constrained to produce arguments that match your JSON schema exactly — including refusing unknown fields. See Structured output for the equivalent on plain replies. On Claude models a strict tool is the only schema-enforced path: response_format is not honoured there.