Documentation menu
Documentation/LLM API/LLM API Quickstart

LLM API Quickstart

Call CAPI language models through a supported synchronous or streaming protocol.

2 min read

Language models are synchronous: you send a request and get the answer back in the same connection. CAPI exposes each provider's native shape, so existing clients work by changing the base URL.

Pick a protocol

You only need one, and it should match the client you already use:

ProtocolBase pathUse when
OpenAI Chat Completions/v1/chat/completionsYou already use the OpenAI SDK or a compatible client.
OpenAI Responses/v1/responsesYou want reasoning items, built-in tools, and server-side state.
Anthropic Messages/v1/messagesYour code targets Claude's native schema.
Gemini generateContent/v1beta/models/{model}:generateContentYou use the Google GenAI SDKs.

Any model in the catalog can be addressed through any of these routes where the provider supports it.

OpenAI-compatible request

SHELL
curl https://capi.ai/api/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6",
    "messages": [
      { "role": "system", "content": "You are a concise assistant." },
      { "role": "user", "content": "Explain task queues in two sentences." }
    ]
  }'

Because the schema matches OpenAI, the official SDK works unchanged:

PYTHON
from openai import OpenAI

client = OpenAI(
    base_url="https://capi.ai/api/v1",
    api_key="YOUR_API_TOKEN",
)

response = client.chat.completions.create(
    model="claude-opus-5",
    messages=[{"role": "user", "content": "Explain task queues."}],
)

print(response.choices[0].message.content)

Note the model: claude-opus-5 through the OpenAI-compatible route. Model choice and protocol are independent.

Streaming

Set stream: true to receive server-sent events. CAPI forwards provider deltas without buffering, so time-to-first-token matches the upstream provider.

JAVASCRIPT
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://capi.ai/api/v1",
  apiKey: process.env.CAPI_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "gpt-5.6",
  messages: [{ role: "user", content: "Write a haiku about tides." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Tool calling

Tool and function definitions pass through to the provider. The response carries tool_calls, and you send results back as role: "tool" messages, exactly as with OpenAI:

JSON
{
  "model": "gpt-5.6",
  "messages": [{ "role": "user", "content": "What's the weather in Lisbon?" }],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Look up current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": { "city": { "type": "string" } },
          "required": ["city"]
        }
      }
    }
  ]
}

Choosing a model

Pricing and context windows differ by an order of magnitude across the catalog, so route deliberately:

  • Reasoning and long contextclaude-opus-5, gpt-5.6-sol, gemini-3.1-pro
  • Balanced production defaultgpt-5.6, claude-sonnet-5, gemini-3.1-flash
  • High-volume, low-costdeepseek-v4-flash, glm-5-air, gpt-5-mini

The full list, with per-token pricing, is in the model catalog.

Usage accounting

Every response reports token usage, and the settled cost is available on the same request in the dashboard:

JSON
{
  "usage": {
    "prompt_tokens": 42,
    "completion_tokens": 118,
    "total_tokens": 160
  }
}

Next steps