文档目录
LLM API Quickstart
Call CAPI language models through a supported synchronous or streaming protocol.
3 分钟阅读
此页正文目前仅提供英文版本,我们正在陆续补充中文翻译。界面与导航已完成中文化。
Language model requests are synchronous: you send a request and get the answer back in the same connection. CAPI supports OpenAI Chat Completions, the OpenAI Responses request and response format, and Anthropic Messages. Requests are relayed to the configured channel, so feature availability depends on the selected model and upstream.
Pick a protocol
You only need one, and it should match the client you already use:
| Protocol | Base path | Use when |
|---|---|---|
| OpenAI Chat Completions | /v1/chat/completions | You already use the OpenAI SDK or a compatible client. |
| OpenAI Responses | /v1/responses | Your client uses the OpenAI Responses API, including structured input or streaming. |
| Anthropic Messages | /v1/messages | Your code targets Claude's native schema. |
Use a model that is enabled for your workspace and compatible with the selected route. Responses requests are forwarded in the Responses schema; built-in tools and other model-specific features work only when the configured upstream supports them. CAPI does not currently expose Gemini's native generateContent route.
OpenAI-compatible request
curl https://capi.minapp.xin/api/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6",
"messages": [
{ "role": "system", "content": "You are a concise assistant." },
{ "role": "user", "content": "Explain task queues in two sentences." }
]
}'Because the schema matches OpenAI, the official SDK works unchanged:
from openai import OpenAI
client = OpenAI(
base_url="https://capi.minapp.xin/api/v1",
api_key="YOUR_API_TOKEN",
)
response = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "Explain task queues."}],
)
print(response.choices[0].message.content)Note the model: claude-opus-5 through the OpenAI-compatible route. Model choice and protocol are independent.
Streaming
Set stream: true to receive server-sent events. CAPI forwards provider deltas without buffering, so time-to-first-token matches the upstream provider.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://capi.minapp.xin/api/v1",
apiKey: process.env.CAPI_API_KEY,
});
const stream = await client.chat.completions.create({
model: "gpt-5.6",
messages: [{ role: "user", content: "Write a haiku about tides." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Tool calling
Chat Completions forwards supported request fields to the configured upstream. Tool calling depends on the selected channel and model:
{
"model": "gpt-5.6",
"messages": [{ "role": "user", "content": "What's the weather in Lisbon?" }],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Look up current weather for a city.",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
}
]
}Choosing a model
Pricing and context windows differ by an order of magnitude across the catalog, so route deliberately:
- Reasoning and long context —
claude-opus-5,gpt-5.6-sol,gemini-3.1-pro - Balanced production default —
gpt-5.6,claude-sonnet-5,gemini-3.1-flash - High-volume, low-cost —
deepseek-v4-flash,glm-5-air,gpt-5-mini
The full list, with per-token pricing, is in the model catalog.
Usage accounting
Every response reports token usage, and the settled cost is available on the same request in the dashboard:
{
"usage": {
"prompt_tokens": 42,
"completion_tokens": 118,
"total_tokens": 160
}
}Successful responses include x-request-id and x-capi-request-id headers for support and log correlation. Anthropic Messages responses also include Anthropic's request-id header. Relay errors include the same request ID in the response headers and error message.
Next steps
- Chat Completions reference
- Anthropic Messages reference
- Authentication — key scoping and rate limits.