Platform
Prompt caching
Cached input costs less. How it is applied and how to see it in your usage.
How it works
Caching is automatic. There is nothing to turn on: no headers, no flags, no parameters. When the beginning of a request matches one Kael has seen recently, those input tokens are served from cache and billed at the cached rate.
What a hit costs
| Input tokens | Off, low, high, max | z-low, z-high |
|---|---|---|
| Not cached | $2.00 | $6.00 |
| Served from cache (80% less) | $0.40 | $1.20 |
A request only gets the cached rate on the tokens that hit, so its overall saving is somewhere between none and 80%, depending on how much of the input was cached.
See your hits
| Where | Field |
|---|---|
| Chat Completions | usage.prompt_cache_hit_tokens |
| Responses | usage.input_tokens_details.cached_tokens |
| Messages | Not reported in the response |
| Dashboard | The Cache column on each request |
| Usage page | Cached and uncached input, per day |
prompt_tokens already includes the cached tokens, so the tokens you paid full price for are prompt_tokens minus the cached count.
Getting more hits
- Caching matches the start of a request, so keep the beginning stable from one request to the next and add what changes at the end.
- In a conversation, append new turns rather than rewriting earlier ones.
- Check the cached count in your own responses to see what you actually get.
Answers are not cached
Kael does not cache answers. Sending the same request twice generates two fresh answers, and you are billed for both.