Field Notes on Prompt Caching: Cache Breakpoints and the Timestamp That Invalidated Every Hit
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
Prompt caching is often sold as a effortless discount on LLM infrastructure bills. The promise is enticing: tag your prompt with a cache marker, sit back, and watch your input token costs plummet by up to 90%. However, during a production deployment of a customer support agent powered by long system instructions and API references, the actual bill did not move by a single cent.
Upon checking the usage payload returned by the Anthropic Messages API, the cause was starkly clear. The cache_creation_input_tokens counter was populated on every single call, while cache_read_input_tokens remained firmly at zero. The system was paying a 25% premium (1.25x base cost) to write cache entries that were never once read. A single dynamic timestamp string interpolated into the system prompt had invalidated the byte-level prefix match required for every subsequent hit.
In this technical guide, we will analyze how LLM prompt caching works under the hood across major providers like Anthropic Claude, OpenAI, and Google Gemini, break down the exact mathematical pricing models, detail the five silent cache invalidation antipatterns, and demonstrate how to architect robust, static-first prompts.
How Prompt Caching Works Under the Hood
Prompt caching is a provider-side mechanism that stores the processed key-value (KV) states of an input prefix in GPU memory or fast host storage. When a new inference request arrives, the API gateway checks whether the leading prefix of the request matches an existing, pre-computed token sequence.
If an identical prefix is found, the provider skips the expensive pre-fill phase (KV projection and attention matrix computation) for those tokens. Instead, it reads the pre-computed KV cache directly into inference memory.
Cold Request (Cache Write): [Tools + System Prompt + API Ref] ---> [KV Pre-fill] ---> [Model Generation]
Warm Request (Cache Read): [Cached KV State] + [User Query] -------------> [Model Generation]
Because pre-fill is computationally heavy, skipping it allows providers to offer massive discounts on cached input tokens. However, the cache mechanism works on a strict exact byte-identical match of the serialized prefix. If token #10 in a 4,000-token prompt changes, tokens #10 through #4000 are immediately invalidated.
When deploying high-throughput agentic workflows, accessing reliable and high-speed API endpoints is essential. Routing your requests through standard LLM aggregators like n1n.ai ensures seamless infrastructure scaling while managing costs across multi-provider deployments.
The Economics of Prompt Caching
To evaluate whether prompt caching makes financial sense, you must understand the three distinct line items in your input token bill:
- Fresh (Uncached) Input Tokens: Standard input token cost (1.0x base price).
- Cache Creation Tokens (Cache Write): Charged when a prompt prefix is compiled into a cache entry.
- 5-minute TTL: Billed at 1.25x base price.
- 1-hour TTL: Billed at 2.0x base price.
- Cache Read Tokens: Billed when a request hits a pre-existing cache entry (0.1x base price, representing a 90% discount).
The Break-Even Arithmetic
The financial break-even calculation is deterministic. Let be the base input price per token:
- 5-Minute Cache Window:
- Write cost:
- Read cost:
- Savings per hit:
- Extra cost paid on write:
- Hits required to break even: hits.
A single cache hit within the 5-minute window easily repays the 1.25x write surcharge.
- 1-Hour Cache Window:
- Write cost:
- Extra cost paid on write:
- Hits required to break even: hits.
Roughly two cache hits inside the 1-hour window make the longer TTL profitable.
| Caching Mode | Write Price Multiplier | Read Price Multiplier | Hits to Break Even | Primary Use Case |
|---|---|---|---|---|
| Uncached | 1.0x | N/A | N/A | One-off queries, small prompts (< 1024 tokens) |
| 5-Minute Cache | 1.25x | 0.1x | 1 hit | Interactive chat, agent reasoning loops, multi-turn support |
| 1-Hour Cache | 2.00x | 0.1x | 2 hits | Fixed documentation analysis, enterprise knowledge base search |
Every time a cache read occurs, the TTL window automatically resets. A steady stream of user traffic every 4 minutes will keep a 5-minute cache entry alive indefinitely without ever incurring another cache write charge.
The 5 Silent Cache Killers
Cache invalidation fails silently. The API does not throw an error when a cache miss occurs; it simply processes the request as a cache write or a fresh execution. Here are the five most common antipatterns that reduce your cache hit rate to zero.
1. The Interpolated System Timestamp
Developers frequently inject dynamic context directly into the system prompt:
// ❌ BAD: Invalidates the cache on every single request!
const systemPrompt = `You are a helpful customer support agent.
Current server time: ${new Date().toISOString()}`;
Because the ISO timestamp changes down to the millisecond, every request generates a brand-new binary prefix. The model writes the entry at 1.25x pricing and never reads it.
2. Non-Deterministic Serialization
If your code builds tool definitions by iterating over JavaScript Map objects, standard Python dict structures across processes, or un-sorted Set instances, the JSON string serialization order can vary across runs.
# ❌ BAD: Key order in raw dictionaries can be non-deterministic across processes
tools = list({"get_user": fn1, "search_db": fn2}.values())
Because Anthropic serializes tool definitions before the system prompt, an unstable tool array invalidates both the tools and the entire system prompt that follows it.
3. Model ID Swapping and Alias Shifts
Cache entries are strictly isolated per model identifier. Routing a request to claude-3-5-sonnet-20241022 will not hit a cache built under claude-3-5-sonnet-20240620. Furthermore, if an API gateway or provider alias silently points traffic to a different backend snapshot, all cached prefixes are dropped.
Using unified endpoints from platform aggregators like n1n.ai allows developers to maintain consistent model identifier routing and avoid subtle upstream alias shifts.
4. Front-Trimming Conversation History
When standard chat applications exceed token limits, developers often pop the oldest message turn from the top of the context window.
// ❌ BAD: Removing index 0 changes the root prefix of the messages array!
messages.shift();
Truncating from the front alters the leading tokens of the messages array, invalidating any multi-turn cache breakpoints attached to subsequent assistant turns.
5. Parallel Fan-Out Cold Starts (Throttling)
If an application fires 20 parallel requests with an identical cached system prompt during a cold start, all 20 requests will execute concurrently before the first request finishes writing the cache. You end up paying the 1.25x cache creation cost 20 times over.
Provider Rules & Minimum Token Thresholds
Not all prompts can be cached. Providers impose minimum token length requirements before the cache_control marker is respected.
- Anthropic Claude 3.5 Sonnet & Claude 3 Opus: Minimum prefix length of 1024 tokens.
- Anthropic Claude 3.5 Haiku: Minimum prefix length of 2048 tokens.
- OpenAI (GPT-4o, o1, o3-mini): Implicit caching triggered automatically on prompts exceeding 1024 tokens (in 128-token increments).
- Google Gemini 1.5 Pro/Flash: Implicit caching for long contexts, plus explicit
CachedContentmanagement for fixed targets (> 32,768 tokens).
If your total prompt prefix up to the breakpoint is 950 tokens, Anthropic silently ignores the cache_control block. The request executes at standard 1.0x input pricing, and cache_creation_input_tokens returns 0.
Practical Implementation: Anthropic Messages API
The Anthropic Messages API uses explicit cache breakpoints. The provider serializes requests in a mandatory sequence:
To maximize cache efficiency, place your static definitions at the top and mark the transition boundary with cache_control: { type: "ephemeral" }.
Single Breakpoint Architecture (Static Context)
Here is how to properly structure a single-breakpoint request in TypeScript:
import Anthropic from "@anthropic-ai/sdk";
const anthropic = new Anthropic(\{ apiKey: process.env.ANTHROPIC_API_KEY \});
async function runAgent(userQuestion: string, currentTimeString: string) \{
const response = await anthropic.messages.create(\{
model: "claude-3-5-sonnet-20241022