Platform
Limits
Context, output, rate, concurrency, and time limits.
Limits
| Limit | Value | Notes |
|---|---|---|
| Context window | 1M tokens | How much conversation Kael can read in one request. |
| Response length | 128,000 tokens | Fixed. max_tokens does not change it. Ask for shorter answers in your prompt. |
| Requests per minute | 150 per account | Counted across all your keys and the Playground. |
| Requests at once | 30 per account | Requests that are still running. |
| Requests per minute from one network | 300 per IP address | Checked before your key. Offices and shared networks count together. |
| Time per request | 8 minutes | A longer request is cut off with a timeout error. |
| Input | Text and images | No PDFs, audio, or video. |
| Output | Text | Kael does not create images or files. |
These are the limits today. Kael is in beta, so they can change as we learn from real traffic.
When you hit one
Going over a rate or concurrency limit returns 429 with the code rate_limit_exceeded or concurrency_limit_exceeded. Wait, then retry with exponential backoff. See Errors.
Long requests can run for minutes. Stream them, and set your client timeout above 8 minutes.
What we do not publish yet
There is no uptime commitment or SLA yet, and no worst-case latency figures for long prompts. See Not available yet.