Build
Endpoints
The three request formats Kael speaks, the routes for each, and what comes back.
Routes
Kael speaks three request formats. Pick the one your SDK or framework already uses; the answer and the price are the same in all three.
| Route | Format |
|---|---|
POST/v1/chat/completions | OpenAI Chat Completions |
POST/v1/responses | OpenAI Responses |
POST/v1/messages | Anthropic Messages |
GET/v1/models | Lists the models you can call |
- OpenAI SDKs use the base URL
https://api.quancis.space/v1and add/chat/completionsor/responsesthemselves. - Anthropic SDKs use
https://api.quancis.spacewith no/v1, because they add/v1/messagesthemselves. - One API key works on every route. See Authentication.
Model name
Send kael-beta as model. Any other value returns a 400 with the code model_not_found. The older names quan-pro and quan-4 still resolve, but they are deprecated.
GET /v1/models
{
"object": "list",
"data": [
{"id": "kael-beta", "object": "model", "created": 1755000000, "owned_by": "quancis"}
]
}The same request in each format
POST /v1/chat/completions
{
"model": "kael-beta",
"messages": [
{"role": "system", "content": "You are a careful code reviewer."},
{"role": "user", "content": "Review this function for bugs: ..."}
],
"reasoning_effort": "high"
}| Item | Chat Completions | Responses | Messages |
|---|---|---|---|
| System prompt | a system or developer message | instructions | top-level system |
| Thinking level | reasoning_effort | reasoning.effort | reasoning_effort |
| Answer text | choices[0].message.content | the message item in output | the text block in content |
| Tokens billed | usage.prompt_tokens, usage.completion_tokens | usage.input_tokens, usage.output_tokens | usage.input_tokens, usage.output_tokens |
| Cached tokens | usage.prompt_cache_hit_tokens | usage.input_tokens_details.cached_tokens | Not reported |
| Why it stopped | finish_reason: stop, length, tool_calls | status: completed, incomplete | stop_reason: end_turn, max_tokens, tool_use |
Response shapes
The Chat Completions response is shown in the Quickstart. The other two look like this. A Responses output list can hold more than one kind of item, so read the one whose type is message.
{
"id": "...",
"status": "completed",
"model": "kael-beta",
"output": [
{
"type": "message",
"role": "assistant",
"content": [
{"type": "output_text", "text": "The range goes one step too far: ..."}
]
}
],
"usage": {
"input_tokens": 38,
"output_tokens": 214,
"total_tokens": 252,
"input_tokens_details": {"cached_tokens": 0}
}
}The model field always says kael-beta. The usage you see is what you are billed. See Pricing and billing.