Quancisuancis

Build

Streaming

Stream answers over server-sent events, including Kael's short thinking titles.

Turn it on

Set stream to true on any of the three formats. The response is then a stream of server-sent events, and the OpenAI and Anthropic SDKs handle it for you.

What to expect from the stream

  • The answer starts after the draft is done. Draft text is never shown to you, so the first answer token arrives later than with a single model call.
  • Until then, with thinking on, you get thinking titles. These are short lines such as "Reading the stack trace". They are written by a separate summarizer. Raw reasoning is never sent, so do not parse the titles or depend on their wording.
  • Keep-alives arrive about once a second while Kael is quiet. Ignore them (see below).

By format

ItemChat CompletionsResponsesMessages
Answer textchoices[0].delta.contentContent delta eventstext_delta events
Thinking titleschoices[0].delta.reasoning_contentresponse.reasoning_text.deltathinking blocks (thinking_delta)
Keep-alivea chunk with content: ""response.pingthe standard ping event
End of streamdata: [DONE]response.completed (or response.incomplete); no [DONE]message_stop
Usagefinal chunk, if you set stream_options.include_usagein the response.completed eventin the final message_delta event

On the Responses format, Kael sends content deltas and a final completed event, not every item lifecycle event. If your client expects each lifecycle event, treat the completed event as the source of truth.

Example

Python
stream = client.chat.completions.create(
    model="kael-beta",
    messages=[{"role": "user", "content": "Explain this stack trace: ..."}],
    stream=True,
    stream_options={"include_usage": True},
)

for chunk in stream:
    if not chunk.choices:  # the final usage chunk has no choices
        print("\n", chunk.usage)
        continue

    delta = chunk.choices[0].delta
    title = getattr(delta, "reasoning_content", None)
    if title:
        print("[thinking]", title)
    if delta.content:  # empty strings are keep-alives
        print(delta.content, end="", flush=True)

A Chat Completions stream looks like this (shortened):

Server-sent events
data: {"choices":[{"index":0,"delta":{"reasoning_content":"Reading the stack trace"}}]}

data: {"choices":[{"index":0,"delta":{"content":""}}]}

data: {"choices":[{"index":0,"delta":{"content":"The error comes from"}}]}

data: [DONE]

Failures

  • Errors before the stream starts use a normal HTTP status. See Errors.
  • Errors after the stream starts cannot change the status, which is already 200. They arrive as a data: event holding an error object. Treat that as a failed request.
  • If the connection drops, send the request again. A retry is a new request.
  • Without streaming, the connection stays open until the answer is ready, which can take minutes on hard requests. Set client timeouts above 8 minutes, or stream.