Quancisuancis

Platform

Prompt caching

Cached input costs less. How it is applied and how to see it in your usage.

How it works

Caching is automatic. There is nothing to turn on: no headers, no flags, no parameters. When the beginning of a request matches one Kael has seen recently, those input tokens are served from cache and billed at the cached rate.

What a hit costs

Input tokensOff, low, high, maxz-low, z-high
Not cached$2.00$6.00
Served from cache (80% less)$0.40$1.20

A request only gets the cached rate on the tokens that hit, so its overall saving is somewhere between none and 80%, depending on how much of the input was cached.

See your hits

WhereField
Chat Completionsusage.prompt_cache_hit_tokens
Responsesusage.input_tokens_details.cached_tokens
MessagesNot reported in the response
DashboardThe Cache column on each request
Usage pageCached and uncached input, per day

prompt_tokens already includes the cached tokens, so the tokens you paid full price for are prompt_tokens minus the cached count.

Getting more hits

  • Caching matches the start of a request, so keep the beginning stable from one request to the next and add what changes at the end.
  • In a conversation, append new turns rather than rewriting earlier ones.
  • Check the cached count in your own responses to see what you actually get.

Answers are not cached

Kael does not cache answers. Sending the same request twice generates two fresh answers, and you are billed for both.