Quancisuancis

Platform

Limits

Context, output, rate, concurrency, and time limits.

Limits

LimitValueNotes
Context window1M tokensHow much conversation Kael can read in one request.
Response length128,000 tokensFixed. max_tokens does not change it. Ask for shorter answers in your prompt.
Requests per minute150 per accountCounted across all your keys and the Playground.
Requests at once30 per accountRequests that are still running.
Requests per minute from one network300 per IP addressChecked before your key. Offices and shared networks count together.
Time per request8 minutesA longer request is cut off with a timeout error.
InputText and imagesNo PDFs, audio, or video.
OutputTextKael does not create images or files.

These are the limits today. Kael is in beta, so they can change as we learn from real traffic.

When you hit one

Going over a rate or concurrency limit returns 429 with the code rate_limit_exceeded or concurrency_limit_exceeded. Wait, then retry with exponential backoff. See Errors.

Long requests can run for minutes. Stream them, and set your client timeout above 8 minutes.

What we do not publish yet

There is no uptime commitment or SLA yet, and no worst-case latency figures for long prompts. See Not available yet.