OpenRouter prompt caching: when it saves money and when it quietly costs more

Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
Our test harness said OpenRouter prompt caching worked on Anthropic models. Our write-up a few days later said it didn't. Both were wrong.
The test sent a long block of reference text to Claude Haiku 4.5 with cache_control on it, exactly where OpenRouter's docs say to put it. The harness saw cache fields in the response and marked the feature as working. When we looked at the numbers later, cache_write_tokens was zero, so we called it dropped. The answer had been in the docs the whole time: Haiku 4.5 doesn't cache prompts under 4,096 tokens, and ours was about 3,600.
Most prompt caching problems on OpenRouter show up the same way. The request succeeds, and the feature is on. The response just doesn't tell you clearly when caching didn't happen, or when it happened and cost more than it saved.
What OpenRouter means by prompt caching
OpenRouter uses the word for two features that have almost nothing in common.
Provider prompt caching belongs to the provider. It stores the processed prefix of your prompt so the next request with the same prefix skips re-reading it, and bills those tokens at a discount. The model still runs and writes a fresh answer. Response caching belongs to OpenRouter. It stores the whole response and returns it on an identical request, and the model never runs at all.
The prices, floors and TTLs below are about provider prompt caching. If you're deciding between exact-match and semantic response caching, our guide to when similar questions should share an answer covers that. OpenRouter's response cache has its own section near the end.
Nine providers, nine price tags
OpenRouter passes each provider's cache pricing through without a per-token markup. Its fees sit on credit purchases instead, which our OpenRouter pricing breakdown covers. Multipliers below are relative to the model's normal input price.

Reads run from a tenth of the input price to half of it, and writes from free to double. Five of the nine have no minimum prompt length listed in OpenRouter's docs. And with model fallbacks or the Auto Router, one request can land on a different row from the last, with a cold cache.
The break-even rule
Whether caching saves money comes down to how many times the same prefix gets sent while the cache is still warm. Using the multipliers above, and counting only the cached part of the prompt:
- Anthropic's 5-minute cache costs 1.25x on the first send and 0.1x on every reuse. Two sends cost 1.35x against 2x uncached, so it pays from the first reuse.
- Anthropic's 1-hour cache costs 2x on the first send. Two sends cost 2.1x, slightly worse than not caching. Three sends cost 2.2x against 3x. It pays from the second reuse.
- OpenAI on GPT-5.6, reading at 0.5x, costs 1.75x for two sends against 2x. It pays from the first reuse and loses 25% if there is no reuse at all.
- DeepSeek, and every provider with free writes, can't lose. The first send costs what it would have anyway.
The rest of each request, the user's question and anything after the last cached block, bills the same either way, so these ratios apply to the prefix only.

When caching costs more than it saves
For most of the table a cache write is free or close to it, so caching is a free option. GPT-5.6 changed that for OpenAI. From that family on, cache writes bill at 1.25x the input price under automatic caching too, and automatic caching is on whether you asked for it or not. Any prompt of 1,024 tokens or more that you send once pays for the write and never collects on the read.
Take a 2,000-token prompt sent once. Without caching it bills as 2,000 tokens of input. With caching on, you pay close to what a 2,500-token prompt would cost. On a nightly job summarising a million long, unique documents, where the document comes first and nothing at the start ever repeats, that's a 25% surcharge on most of your input.
The fix reads backwards. Set prompt_cache_options.mode to "explicit" and don't mark any content block with prompt_cache_breakpoint. Explicit mode switches off OpenAI's automatic breakpoints, and with no explicit ones there's no cacheable prefix, so nothing gets written. cache_write_tokens comes back as zero, which is the one place in this post where a zero means what you want it to.
When it pays: a worked example
Now the other side. Claude Sonnet 4.6 with a 5,000-token system prompt, 100 requests across a working day, arriving in 20 bursts of five with bursts about 20 minutes apart. Counting only the system prompt, in tokens billed at the normal input price:
The 5-minute cache saves two thirds. The 1-hour cache saves 88%, because a 20-minute gap kills the short cache every time and never reaches the long one. Anthropic refreshes the cache lifetime each time it's read, so the 1-hour entry stays warm all day.
Change the traffic and the answer changes. If those 100 requests arrived evenly, one every four minutes or so, the 5-minute cache would never expire either, and the 1-hour option would mean paying 10,000 for a write the short cache handles at 6,250. Pick the TTL from the gaps between your requests, not from how many you send.
Token floors fail silently
Every provider sets a minimum prompt length below which nothing caches. For Claude, per OpenRouter's docs:
- 4,096 tokens: Opus 4.5, 4.6, 4.7 and 4.8, and Haiku 4.5
- 2,048 tokens: Haiku 3.5
- 1,024 tokens: Sonnet 4, 4.5 and 4.6, and Opus 4 and 4.1
OpenAI's floor is 1,024. Gemini's is 1,024 on 2.5 Flash and 4,096 on 2.5 Pro. Under the floor the request succeeds, bills at full input price, and comes back with its cache fields set to zero. No error, no warning.
Two consequences of this are easy to miss.
Newer models have higher floors. A 3,000-token system prompt caches on Sonnet 4.6 and stops caching the day you move it to Opus 4.7, with no code change and nothing in the response to tell you. Moving from Haiku 3.5 to Haiku 4.5 does the same for anything between 2,048 and 4,096 tokens. A model upgrade is a caching change, so check your cache hit rate after one.
And you can't verify caching by looking for cache fields. That's what caught our own test. OpenRouter returns cached_tokens and cache_write_tokens on responses from every provider, including ones where caching isn't native, so the fields being there proves nothing. Zeros don't prove much either. They mean the marker was ignored or the prompt was under the floor, and on Haiku 4.5 at 3,600 tokens it was the floor.
A test that tells you something:
- Use a prompt comfortably above the floor for that exact model. Token counts vary between tokenizers, so leave headroom.
- Send it twice in quick succession.
- On providers that charge for writes, expect
cache_write_tokensabove zero on the first call. Expectcached_tokensabove zero on the second. - Read
cache_discountin the response body. It's negative on an Anthropic write and positive on a read.
TTLs don't survive routing
OpenRouter translates cache markers between providers, which is useful. A text block with Anthropic-style cache_control becomes a prompt_cache_breakpoint when routed to a GPT-5.6 model, and a prompt_cache_breakpoint becomes a default cache_control on Anthropic or Google.
The TTL doesn't come along. A ttl on cache_control is dropped on the way to OpenAI. prompt_cache_options, where OpenAI's TTL lives, only applies to OpenAI. prompt_cache_breakpoint has no TTL field, so whatever it becomes on the Anthropic side gets the 5-minute default.
Lifetimes also behave differently once they're set. Anthropic's refreshes on every read. Gemini's explicit cache lasts five minutes from the write and doesn't refresh. Gemini's implicit cache is documented as three to five minutes on average, varying. So one request body asking for an hour can get an hour, five fixed minutes, or whatever OpenAI applies, depending on where it lands. The Responses API adds one more gap: it doesn't expose per-block Anthropic cache_control, so use Chat Completions or the Messages endpoint if you need a TTL on a specific block.
Sticky routing, and when a pin overrides it
A cache only helps if your next request reaches the provider holding it. OpenRouter handles that with provider sticky routing: after a request that uses caching, follow-up requests for the same model go to the same provider endpoint.
Stickiness is tracked per account, per model and per conversation, where a conversation is identified by hashing your first system message and first non-system message. It expires after 10 minutes without a request.
Agents tend to break that default, because their opening messages change between turns. Pass a session_id in the body or an x-session-id header and OpenRouter uses it as the sticky key instead, starting from the first successful request rather than waiting for a cache hit. OpenAI's prompt_cache_key works as a fallback key if you already send it.
Setting provider.order turns sticky routing off and your list takes priority. While your first choice is healthy that costs little, since a pin keeps you on one provider anyway. The difference shows after a failover, when sticky routing would follow the provider that served you and a pin starts again from the top of your list. Pinning is still right when you need predictable capacity, as we argued in our post on OpenRouter rate limits. The Auto Router has a similar edge: if the model it picked drops out of its candidates on a later turn, the cache goes with it.
OpenRouter's own response cache
Separate from all of the above, OpenRouter will cache whole responses if you send X-OpenRouter-Cache: true. A hit returns the stored response without calling a provider, bills nothing, and doesn't count against provider rate limits.
We measured it across three runs of 100 identical re-requests each, on GPT-4o mini and Claude Haiku 4.5. Every re-request was a hit. Median latency was 2,867 ms on a miss and 59 ms on a hit, and Haiku hits were quicker at 39 ms against 64 to 69 ms for GPT-4o mini. On the run billed to credits, 100 misses cost $0.0174 and 100 hits cost nothing.
Entries are keyed on the API key plus model, endpoint type, streaming mode and full request body, so services with separate keys share nothing and any parameter change is a miss. Our TTL probe, one sample per interval, found entries alive at 120 seconds and gone by 300. There's no published TTL, no semantic matching and no documented purge. It absorbs retries and duplicate submissions. An hourly repeat misses.
How TrueFoundry handles caching
Our gateway caches at the response layer in two modes. Exact-match returns a stored response for an identical request. Semantic mode embeds the last message and returns a stored response when an earlier request was close enough, and serves exact-match hits too, so you run one mode, not both. It's switched on with a request header, backed by Redis, and composes with provider prompt caching: on a gateway miss, the provider's own prefix cache still applies.
On SaaS the embedding model is text-embedding-3-small at 1,536 dimensions and isn't configurable; on-prem you choose it gateway-wide. Semantic matching has a failure mode that exact-match doesn't. Our semantic caching guide opens with it: two customers ask about their delivery minutes apart, the questions embed to nearly the same vector, and the second customer gets an answer about the first one's order. That is a scope bug. Moving the similarity threshold changes how often it happens, and it does not stop a cache that is shared across customers.
The cache lookup adds roughly 3 to 4 ms, and the gateway handles 350+ RPS on a single vCPU.
A short checklist
- Check the minimum prompt length for the exact model you're calling, and check it again after every model upgrade.
- Put stable content first and anything per-request last, so the prefix actually repeats.
- On GPT-5.6 and later, switch caching off for prompts you'll only send once.
- Choose the TTL from the gaps between requests: 5 minutes for steady traffic, 1 hour for bursts more than five minutes apart.
- Send a
session_idfor agent sessions so routing stays on the provider with your cache. - Verify caching with non-zero
cached_tokens, never with the field being present.
Related reading
- Semantic caching: when text stops being the right cache key, on embedding choice, thresholds and tenant isolation
- Best AI gateways for LLM inference optimization in 2026, comparing how gateways handle caching and routing
- OpenRouter alternatives, for teams weighing a move to a governed gateway
- Caching (exact and semantic) in TrueFoundry, the configuration reference
Conclusion
OpenRouter prompt caching works, and the docs are clearer here than they are for a lot of the surrounding features. You still have nine pricing schemes, floors that rise on newer models, lifetimes that change in translation, and a response that looks the same whether caching happened or not. Check the floor, put the stable text first, and verify with numbers rather than fields.
To configure caching once for every provider, with semantic matching on top of the provider's own prefix cache, see how the TrueFoundry AI Gateway handles caching.
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
What is OpenRouter prompt caching?
It's OpenRouter passing through each provider's prompt caching, where a repeated prompt prefix is stored and billed at a discount on later requests. Some providers cache automatically, while Anthropic and Alibaba need cache_control markers on the content you want cached. Prices differ by provider, from free writes and 0.1x reads to 2x writes on Anthropic's 1-hour cache.
Does cache_control work through OpenRouter?
Yes, as long as the prompt clears the model's minimum length. Claude Haiku 4.5 and Opus 4.5 and later need 4,096 tokens, and Sonnet 4.x needs 1,024. Below that floor the request succeeds at full price with cache fields at zero, which looks exactly like the marker being ignored.
Why did my bill go up after turning on caching?
Most likely you're paying for cache writes on prompts that never get reused. GPT-5.6 and later bill writes at 1.25x even under automatic caching, and Anthropic's 1-hour cache bills writes at 2x. If the prefix doesn't repeat while the cache is warm, you pay for the write and never get the discount.
Posso implantar o TrueFoundry na minha própria VPC ou on-prem?
Sim. O TrueFoundry é executado na sua VPC, on-prem, em ambiente air-gapped ou híbrido, de modo que os prompts e as respostas nunca saem do seu domínio, mesmo quando você roteia entre muitos provedores.














.png)
.png)



.png)



.png)
.png)
.png)

.png)
.png)





