Build
Streaming
Stream answers over server-sent events, including Kael's short thinking titles.
Turn it on
Set stream to true on any of the three formats. The response is then a stream of server-sent events, and the OpenAI and Anthropic SDKs handle it for you.
What to expect from the stream
- The answer starts after the draft is done. Draft text is never shown to you, so the first answer token arrives later than with a single model call.
- Until then, with thinking on, you get thinking titles. These are short lines such as "Reading the stack trace". They are written by a separate summarizer. Raw reasoning is never sent, so do not parse the titles or depend on their wording.
- Keep-alives arrive about once a second while Kael is quiet. Ignore them (see below).
By format
| Item | Chat Completions | Responses | Messages |
|---|---|---|---|
| Answer text | choices[0].delta.content | Content delta events | text_delta events |
| Thinking titles | choices[0].delta.reasoning_content | response.reasoning_text.delta | thinking blocks (thinking_delta) |
| Keep-alive | a chunk with content: "" | response.ping | the standard ping event |
| End of stream | data: [DONE] | response.completed (or response.incomplete); no [DONE] | message_stop |
| Usage | final chunk, if you set stream_options.include_usage | in the response.completed event | in the final message_delta event |
On the Responses format, Kael sends content deltas and a final completed event, not every item lifecycle event. If your client expects each lifecycle event, treat the completed event as the source of truth.
Example
Python
stream = client.chat.completions.create(
model="kael-beta",
messages=[{"role": "user", "content": "Explain this stack trace: ..."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if not chunk.choices: # the final usage chunk has no choices
print("\n", chunk.usage)
continue
delta = chunk.choices[0].delta
title = getattr(delta, "reasoning_content", None)
if title:
print("[thinking]", title)
if delta.content: # empty strings are keep-alives
print(delta.content, end="", flush=True)A Chat Completions stream looks like this (shortened):
Server-sent events
data: {"choices":[{"index":0,"delta":{"reasoning_content":"Reading the stack trace"}}]}
data: {"choices":[{"index":0,"delta":{"content":""}}]}
data: {"choices":[{"index":0,"delta":{"content":"The error comes from"}}]}
data: [DONE]Failures
- Errors before the stream starts use a normal HTTP status. See Errors.
- Errors after the stream starts cannot change the status, which is already
200. They arrive as adata:event holding anerrorobject. Treat that as a failed request. - If the connection drops, send the request again. A retry is a new request.
- Without streaming, the connection stays open until the answer is ready, which can take minutes on hard requests. Set client timeouts above 8 minutes, or stream.