Overview
Prompt caching lets you cache large, reusable prompt prefixes (system prompts, reference documents, tool definitions, and long conversation history) so follow-up requests can reuse the cached prefix instead of reprocessing it from scratch. Benefits:- Up to ~90% cost reduction on cached input tokens (cache hits)
- Lower latency on requests with large static prefixes
- Implicit caching (default): For providers/models that support provider-native prompt reuse (including OpenAI, Gemini, and many open-source provider routes), caching is applied automatically when eligible. No extra request fields are required.
- Explicit prompt caching (Claude): Claude caching uses request controls (
prompt_caching/promptCaching/body-levelcache_control, or inlinecache_control) to mark reusable prefixes and select TTLs. You can add these yourself, or your client can supply them automatically.
enabled: true and ttl: "1h" in the helper; you do not also need an anthropic-beta header.
NanoGPT also supports cache-capable provider routing with top-level caching: true. This is not a prompt annotation mode. It only requires routing to a provider that supports prompt/input caching and, by default, tries to keep later matching requests on the same provider.
Supported Models
Implicit caching (automatic)
NanoGPT automatically uses implicit caching on providers/models that support it, including OpenAI and Gemini model families plus many open-source provider/model routes. No cache-control flags are required for this mode. For normal automatic routing, NanoGPT also tries cache-affinity routing when it can help: it hashes the request shape with the user/API-key or session identity and tries to route later matching requests to the same provider. This applies only to eligible automatic routes. Explicit provider selection and routing preferences use their own routing rules. Claude routes also need caller cache intent, andprompt_caching.enabled: false disables that cache affinity.
If the provider that would be used does not support caching, NanoGPT does not apply cache-affinity routing for that request because there is no provider-side cache to reuse. Cache-affinity routing by itself does not add explicit cache-write annotations. For models/routes that require explicit writes, such as Claude, the request still needs a helper or cache markers, whether supplied by you or your client.
Cache-Capable Provider Routing
Set top-levelcaching: true when you want NanoGPT to route the request to any available provider that supports prompt/input caching. This is capability-based routing: you do not need to choose a provider. If no cache-capable provider is available for the model, the request fails rather than silently using a non-caching provider.
caching: true prefers the provider from your previous matching request (soft affinity). stickyProvider defaults to false; set it to true to disable cross-provider fallback. Neither setting guarantees a cache hit.
To disable cross-provider fallback for a cache-capable request:
stickyProvider is accepted as a camelCase alias for stickyprovider. true disables cross-provider fallback. false allows fallback, but explicit caching still keeps soft affinity to a previous eligible provider.
The affinity record expires after about one hour without a successful matching request. This is separate from the provider’s cache TTL and does not guarantee a warm cache for one hour.
Price routing is a different goal. provider.sort: "price", :price, or max_price without order can activate price ranking, which takes precedence over cache-affinity selection and can change providers between requests. See routing examples and precedence.
You can request the same behavior with a model suffix:
:cache and :cached are also accepted.
Use prompt_caching / promptCaching only when you need provider-specific cache-control annotations or TTL behavior. Top-level caching: true does not add Anthropic-style cache_control markers or configure cache TTLs.
Explicit prompt caching controls (Claude)
Explicit prompt-caching controls are available on Claude models, including these families (examples):
For legacy IDs,
anthropic/ aliases are also available (for example anthropic/claude-sonnet-4.5, anthropic/claude-opus-4.6:thinking). Use the exact ID returned by the model catalogue rather than constructing an alias.
Claude minimum cacheable tokens
Minimum cacheable sizes depend on the model version and upstream route. A minimum for an older Claude model does not establish the minimum for Claude 5.5. See the Claude prompt caching documentation for native model thresholds. If your cached prefix is smaller than the selected route’s minimum, the request can still succeed without creating a cache entry. Verify caching with write/read usage fields rather than a successful response alone.Subscription Billing for Chat Completions
OnPOST /api/v1/chat/completions, supported explicit cache controls select pay-as-you-go billing and bypass subscription coverage for that request. This includes the body helper and inline cache_control markers added by clients such as OpenCode.
This rule does not mean an implicit cache hit alone switches a request to pay-as-you-go, and it should not be used to infer billing on the separate Messages endpoint. See pay-as-you-go billing controls for other billing overrides.
How To Enable Explicit Prompt Caching (Claude)
Prompt caching works onPOST /api/v1/chat/completions.
You can enable it in 3 ways.
Option 1: body-level helper (promptCaching / prompt_caching / cache_control)
Add a top-level helper object:
explicitCacheControl (boolean, default false)
When true, no additional cache breakpoints are added automatically. Use it when you want to control every boundary yourself.
For a helper with ttl: "1h", existing message, system, and tool markers are updated to one hour whether this flag is true or false. You do not need explicitCacheControl: true to apply a one-hour TTL to OpenCode’s markers. Leaving it false also allows NanoGPT to add automatic breakpoints within the request’s block limit.
Also accepts the snake_case alias explicit_cache_control.
Aliases are accepted:
promptCachingprompt_cachingcache_control(body-level helper alias)
true instead of an object defaults to:
cutAfterMessageIndex is omitted, NanoGPT selects cache boundaries automatically.
The helper is sufficient for both TTLs. NanoGPT applies the cache settings upstream, so callers do not need to add beta headers alongside it.
Option 2: inline cache_control markers
Attach cache_control directly to content blocks you want cached:
Combining inline markers with body-level settings
cache_control marker is preserved and its TTL is set to 1h. The user message does not receive an auto-generated cache breakpoint.
Option 3: anthropic-beta header (Claude-compatible)
The Anthropic-compatible header is supported as an alternative on models with Claude cache controls. It is optional when using the body helper or inline markers:
ttl takes precedence over the header’s TTL. A body-level enabled: false disables explicit caching even if the header requests it.
Controlling What Gets Cached
Note:cutAfterMessageIndexlimits where new cache breakpoints can be placed automatically.explicitCacheControl: truedisables new automatic breakpoints, so the cut index has no effect on placement in that mode. Applying a one-hour helper TTL to existing markers does not depend on this flag.
cutAfterMessageIndex
Override automatic cache breakpoints by setting the last cached message index:
4. A breakpoint caches the prefix up to that boundary. It does not remove existing inline markers after that index.
You can also set this via request header:
Cache block limit
A maximum of 4cache_control breakpoints are allowed per request across system prompt, tools, and messages. Excess breakpoints are pruned automatically without removing prompt content.
Forcing a Cache Write
There is no separate “force write” flag. Enable prompt caching and send the request. The first eligible request writes cache automatically (if provider thresholds/availability allow it). Repeated requests with the same cached prefix read from cache.Usage Fields (How To Verify Cache Hits)
When caching is active (implicit or explicit), responses can include:cache_creation_input_tokens: tokens written to cache on this requestcache_read_input_tokens: tokens read from cache on this request (cache hit when> 0)prompt_tokens_details.cached_tokens: OpenAI-style cached token count
x_nanogpt_pricing includes cache pricing breakdown fields such as cacheCreationInputTokens, cacheReadInputTokens, cacheTTL, and cacheCost. cacheTTL describes the billing TTL; check write/read counts separately to establish cache use.
For streaming Chat Completions, request usage explicitly and inspect the final usage event:
Pricing
Cache writes and reads are billed differently by provider. Implicit-caching providers apply their cache pricing automatically when eligible. Explicit Claude caching uses the TTL settings below.Gemini Pro models (implicit caching, provider-native)
Example: writing 10,000 cached tokens costs
(10k × $2.00/M) + (10k × $0.375/M) = $0.02375. Reading 10,000 cached tokens costs 10k × $0.20/M = $0.002.
Gemini Flash models (implicit caching, provider-native)
For Gemini 2.0 models, cache reads are 25% of the base input rate (75% cheaper), not 10%.
Claude models (explicit caching)
The standard Claude cache multipliers are below. Model- or route-specific rates can differ; use the selected model’s pricing and the response’s billing breakdown when available.TTL Options (Explicit Claude Controls)
Structuring Prompts for Cache Hits
Cache hits require the cached prefix to be byte-identical across requests. Best practices:- Put static content first (system prompt, reference docs, tool definitions).
- Keep cached content identical across requests (no timestamps, request IDs, or dynamic inserts).
- Put dynamic content after the cache boundary (typically the latest user message).
Large tool catalogs and hosted search
An eagerly loaded function catalog is part of the initial prompt prefix. Adding, removing, or changing any eager definition changes that prefix and can prevent reuse of the earlier cached prompt. On compatible Responses models, hosted tool search keeps a small search-tool definition in the initial prefix and reveals relevant functions later. This can reduce initial input tokens and improve the opportunity for cache reuse across catalog changes. It does not guarantee a cache hit: exact-prefix rules, minimum token thresholds, TTLs, and the selected model route still apply.Cache Consistency with stickyProvider (Explicit Caching)
Each provider keeps its own cache. If a request fails over to another provider, the previous cache may be unavailable.
If cache consistency matters more than availability, set:
stickyProvider: false(default): the request may succeed even if routing changes, but you might rebuild caches.stickyProvider: true: if a fallback would be required, the request returns503instead.
promptCaching.stickyProvider unset unless you accept reduced availability to preserve the cache. It does not extend the TTL or guarantee a hit.
Top-level stickyprovider / stickyProvider and nested promptCaching.stickyProvider control the same fallback switch. If both are supplied, the top-level value takes precedence. With caching: true, soft affinity remains even when the switch is false. Privacy, price, and provider restrictions still constrain eligible routes.
NanoGPT Web UI
In the NanoGPT web UI, models with explicit prompt-caching controls show a prompt caching toggle where you can choose cache duration.Limitations and Caveats
- Provider-side minimum token thresholds still apply before a cache entry is created.
- A maximum of 4 cache breakpoints (
cache_control) are supported per request. - Some models report aggregate prompt usage differently; use
cache_creation_input_tokensandcache_read_input_tokensfor authoritative cached token counts. - On cache hits, a small non-zero
cache_creation_input_tokenscan appear due to per-request overhead and does not necessarily indicate a cache miss. - Implicit caching behavior (eligibility, TTL behavior, and exact discounts) is provider-dependent.