Anthropic prompt caching is easy to switch on and easy to misread. The first request can cost more, a later request can miss without an error, and a “cached” prefix may be too short to qualify.
Short answer: Put the cache breakpoint at the end of content that stays identical across requests, then confirm that a later response reports cache_read_input_tokens above zero using Anthropic's response usage fields. Price the first cache write and every read against your expected reuse count: on Claude Sonnet 5.5, one read inside the five-minute window repays the write premium; the one-hour cache needs two reads. These API figures do not tell you how Claude subscription limits are metered.
Prices and product details below were checked October 4, 2026.
How does Anthropic prompt caching work, and how do you enable it?
Claude prompt caching reuses a matching prompt prefix so the model can read cached input instead of processing it as fresh input again. For a first API test, use automatic caching: add the top-level cache_control field to a Messages request. In an existing Anthropic Messages API Python call, add this request argument (insert it into client.messages.create(...)):
cache_control={"type": "ephemeral"}
That top-level Anthropic cache control setting comes from the prompt caching documentation. Current examples use client.messages.create(...); the old client.beta.prompt_caching.messages.create(...) route is no longer required. For a static document or tool set, an explicit block-level breakpoint lets you control exactly where the reusable prefix ends. Anthropic allows up to four breakpoints.
To verify a hit, make two requests with the same model and unchanged prefix, with the second request starting before the cache expires. Read the response usage object:
| Field | What it tells you |
|---|---|
cache_creation_input_tokens |
Input tokens written to cache on this response |
cache_read_input_tokens |
Input tokens served from cache |
input_tokens |
Uncached input after the last breakpoint |
A first request should show cache creation. A later request should show cache reads; with automatic caching, a new tail may also be written. Anthropic defines total input as the sum of these three fields, so don't use input_tokens alone as total prompt size. If both cache counts are zero, the prompt wasn't cached. An API usage response proves the request-level behavior you observed; it does not establish the hit rate for your whole production workload.
Should you use the five-minute or one-hour cache?
Use the five-minute cache when calls reuse the same prefix within five minutes. Choose the one-hour option when real gaps regularly exceed that window and still fall within an hour. Anthropic starts the TTL clock when the request begins, not when its response ends; a four-minute stream leaves roughly one minute for the next request to arrive.
The break-even point depends on the number of successful reads, not on whether the prompt merely has a breakpoint. For a simple comparison, take one fixed 100,000-token prefix on Claude Sonnet 5.5. The current Anthropic pricing table lists $2 per million ordinary input tokens, $2.50 per million five-minute cache writes, $4 per million one-hour writes, and $0.20 per million reads. The read is 0.1× base input for this model.
For R cache reads, the cached prefix costs write rate + (R × read rate). Without caching, it costs base rate × (1 + R). That gives these totals for the prefix alone; output and changing request tails are excluded:
| Reads after the first write, within the TTL | No cache | 5-minute cache | 1-hour cache |
|---|---|---|---|
| 0 | $0.20 | $0.25 (125%) | $0.40 (200%) |
| 1 | $0.40 | $0.27 (67.5%) | $0.42 (105%) |
| 2 | $0.60 | $0.29 (48.3%) | $0.44 (73.3%) |
| 4 | $1.00 | $0.33 (33%) | $0.48 (48%) |
| 10 | $2.20 | $0.45 (20.5%) | $0.60 (27.3%) |
The percentages are cached cost divided by no-cache cost. On this model and under these assumptions, five-minute caching is cheaper after one read; one-hour caching becomes cheaper after two. The first one-hour write costs twice the ordinary input rate, so extending TTL is not automatically a saving. Anthropic's current cache-read multipliers differ for some models—5% for Opus 5.5 and 2.5% for Fable 5.1 and Mythos 5.1—so recalculate from the row for the model you actually call. Pricing checked October 2026.
A LinkedIn post by Roy Derks works through a hypothetical that explains the "caching raised my bill" complaints: an agent runs every 10 minutes with a 100,000-token reusable prompt, so each of its ten calls arrives after the five-minute entry has expired. Every call pays the 1.25× write and none earns the 0.1× read, so $10 of uncached input becomes $12.50, 25% more. That's a worked scenario, not a measured workload. Compare your own usage counters and call intervals before forecasting anything from it.
Why is my cache not getting hits?
A cache hit requires the same prefix through the marked breakpoint and an eligible prompt length.
| What you see | Likely reason | Fix |
|---|---|---|
| Both cache counters stay at zero | The marked prefix is shorter than the model minimum. For Claude Sonnet 5.5, the current minimum is 512 tokens; shorter requests run without a cache error. | Include enough stable content to clear the model's minimum, then inspect the next response's usage. |
| A full or large rewrite follows a small request change | The changed bytes are before the breakpoint. Anthropic's invalidation table covers edits to tool definitions, the web search, citations and speed toggles, tool_choice, images, and thinking settings; tool-definition changes invalidate the whole cache. |
Keep stable instructions and tool definitions first; move timestamps, per-request values, and the new user input after the static-prefix breakpoint. |
| A changing last block is written again on every call | Automatic caching moves the breakpoint to the last cacheable block; if that block changes, there is no reusable entry at the stable earlier boundary. | Put an explicit breakpoint on the last block that remains identical. |
| A gap unexpectedly causes a write | The cache lifetime expired, or a long response used up most of the TTL before the next request started. | Measure request-start intervals; use ttl: "1h" if your expected reuse window requires it and the extra write charge pencils out. |
| Tools or messages appear identical but reads fall | Ordering or serialization changed the prefix. Anthropic's troubleshooting guidance warns that unstable JSON key ordering inside tool_use blocks can break reuse. |
Keep tool and message ordering stable and serialize structured values deterministically. |
The request prefix is ordered as tools, then system, then messages. A change early in that sequence invalidates later cached content. A breakpoint is a boundary where Anthropic writes an entry; it doesn't search backward and create a missing cache entry for an earlier stable block. Its lookback can find only entries previous requests already wrote. See Anthropic's invalidation and breakpoint guidance when the usage fields show a miss you can't explain.
How does Claude Code prompt caching work?
Claude Code manages prompt caching automatically. You don't add cache_control to a normal Claude Code session. Its request layers put system prompt and project context before the growing conversation, so changing an earlier layer can invalidate everything after it.
The TTL depends on authentication: the current Claude Code prompt caching guide lists one hour for a subscription's main conversation, but five minutes for API keys and cloud providers by default. Claude Code v2.1.242 or later supports per-bucket TTL settings. To request one hour for both buckets with the documented environment variable, set ENABLE_PROMPT_CACHING_1H=1; individual settings or environment variables can take precedence, so check the guide's precedence list before assuming which TTL applied.
For a session summary, run /usage. Claude Code's Session block reports the main conversation's cache hit ratio and miss count after its first response. For request-level evidence, the guide documents cache_creation_input_tokens and cache_read_input_tokens; on v2.1.251 or later, its status-line data exposes the prompt_cache object. These counters are more useful than inferring a miss from a long pause alone.
Claude Code's Prompt cache (main) display can include a likely cause, but one session's transcript is not a universal diagnosis. For example, an Anthropic Claude Code issue reports two measured background-subagent resumes on v2.1.273, both diagnosed as messages_changed, with cache writes of 243,214 and 398,622 tokens. Treat that as a version- and scenario-specific report; use your own usage data and the current release notes before attributing a bill change to a general product defect.
What changes for prompt caching on Amazon Bedrock and OpenRouter?
Bedrock prompt caching and OpenRouter prompt caching follow the same principle—reuse an eligible stable prefix—but the API surface and routing behavior change by provider.
| Route | What differs | Practical check |
|---|---|---|
| Anthropic API | Add top-level automatic cache_control, or place explicit block-level breakpoints. Current docs list up to four breakpoints and model-specific minimum token lengths. |
Read Anthropic usage.cache_creation_input_tokens and usage.cache_read_input_tokens. |
| Amazon Bedrock | AWS documents both implicit and explicit prompt caching. Support, token minimum, TTL, and API fields vary by model. For Claude explicit caching, AWS offers a simplified single checkpoint at the end of static content and checks back through about 20 content blocks; exact checkpoint and request syntax depend on InvokeModel versus Converse. |
Check the Bedrock model card for the model and region, then inspect the provider response's usage fields. Don't copy an Anthropic API payload blindly. |
| OpenRouter | OpenRouter's guide documents top-level automatic caching and explicit per-block markers. It can use provider sticky routing after a cached request to send later calls to the same endpoint; manual provider ordering takes priority. Its Responses API exposes automatic caching, while Anthropic-style per-block controls are not exposed there. | Keep a stable session or opening messages and check which provider served each request; route changes can affect reuse. |
On Bedrock, Anthropic's current platform guide says legacy Bedrock integrations for Opus 4.6 and earlier reject the top-level automatic field; explicit breakpoints are the documented path there. AWS's current page identifies model-specific support, minimums, and TTLs. Follow the page for your exact model rather than carrying that older integration rule forward to every Bedrock Claude model.
If you're building a provider layer, keep fresh input, cache reads, and cache writes as separate usage values. We build Octolib at Muvon; its Anthropic adapter maps the API's cache_read_input_tokens and ephemeral write fields separately, so a gateway-level total doesn't hide the distinction. Our notes on LLM token usage and a unified provider layer explain the broader accounting problem. For proxy-level visibility, see the OctoHub LLM proxy; for model-routing cost tradeoffs, see our LLM routing guide.
You can now switch caching on, make two requests, verify the read field, and compare the measured reuse cadence with the dated pricing table. If those four checks don't agree, stop optimizing the estimated bill and inspect the prefix and provider route first.
— Don
We build Octolib at Muvon and keep cache reads and writes visible as separate usage fields. If you find a provider response that makes those numbers hard to reconcile, open an issue.



