In document Q&A, coding agents, and multi-turn conversations, content such as product documentation, codebases, tool definitions, and system prompts is typically sent repeatedly. Passing it again with every request not only lengthens the context but also incurs input charges over and over.
Context Caching reuses these repeated request prefixes. Cached content is billed at a lower price: for kimi-k3, the cache-hit price is only one-tenth of the cache-miss price, which lowers the cost of high-frequency calls.
Cache Write is now billed as a separate item, and two cache TTL (time to live) options are offered: 5m and 1h. You can choose a TTL based on the interval between calls, keeping cache costs transparent and making savings easy to observe.
Best Practices: Best Practices for context caching - Kimi API Platform
API references:
- Chat Completions API - Kimi API Platform
- Responses API - Kimi API Platform
- Messages API - Kimi API Platform
Pricing: Model Inference Pricing Explanation - Kimi API Platform
Choose between 5m and 1h
Cache Write supports two TTLs: 5m and 1h. When prompt_cache_options is omitted, the system uses the 5m TTL by default: prefixes that meet the hit conditions are automatically written to the cache and reused, and Cache Write charges apply.
- When follow-up requests usually arrive within 5 minutes, use
5m. - When requests may arrive more than 5 minutes apart but the prefix will be reused within 1 hour, use
1h. - When the same prefix is usually reused only after more than 1 hour, do not configure caching specifically for it.
How to improve cache hit rates
Caching matches request prefixes. When any part of a prefix changes, the content after that position cannot be reused.
We recommend:
- Put stable system prompts, tool definitions, reference material, and codebases at the front of the request.
- Put content that changes on every turn, such as user questions, tool results, and task state, at the end.
- Keep the order and exact text of fixed content unchanged within the same session. Do not put timestamps, random IDs, or other dynamic fields in the prefix.
- Keep the interval of scheduled tasks within the TTL, so that each hit renews the entry and keeps it active.
- Cache is isolated by organization (org): shared within an organization, not across organizations.
Caches cannot be cleared manually. A cached prefix expires automatically after it has been inactive for the selected TTL.
