Blank white background with no objects or features visible.

Te ofrecemos acceso gratuito al informe completo Gartner Hype Cycle for AI Governance 2026. Consigue tu copia →

OpenRouter prompt caching: when it saves money and when it quietly costs more

Por Kshitij Gupta

Published: October 8, 2026

TL;DR:

OpenRouter passes through each provider's prompt caching, and each one prices it differently: reads from 0.1x to 0.5x of the input price, writes from free to 2x. Below a model's minimum prompt length nothing caches and nothing tells you. On GPT-5.6 and later, caching a prompt you only send once costs 25% more than not caching it.

Our test harness said OpenRouter prompt caching worked on Anthropic models. Our write-up a few days later said it didn't. Both were wrong.

The test sent a long block of reference text to Claude Haiku 4.5 with cache_control on it, exactly where OpenRouter's docs say to put it. The harness saw cache fields in the response and marked the feature as working. When we looked at the numbers later, cache_write_tokens was zero, so we called it dropped. The answer had been in the docs the whole time: Haiku 4.5 doesn't cache prompts under 4,096 tokens, and ours was about 3,600.

Most prompt caching problems on OpenRouter show up the same way. The request succeeds, and the feature is on. The response just doesn't tell you clearly when caching didn't happen, or when it happened and cost more than it saved.

What OpenRouter means by prompt caching

OpenRouter uses the word for two features that have almost nothing in common.

Provider prompt caching belongs to the provider. It stores the processed prefix of your prompt so the next request with the same prefix skips re-reading it, and bills those tokens at a discount. The model still runs and writes a fresh answer. Response caching belongs to OpenRouter. It stores the whole response and returns it on an identical request, and the model never runs at all.

The prices, floors and TTLs below are about provider prompt caching. If you're deciding between exact-match and semantic response caching, our guide to when similar questions should share an answer covers that. OpenRouter's response cache has its own section near the end.

Nine providers, nine price tags

OpenRouter passes each provider's cache pricing through without a per-token markup. Its fees sit on credit purchases instead, which our OpenRouter pricing breakdown covers. Multipliers below are relative to the model's normal input price.

ProviderHow it's triggeredCache readCache writeMinimum prompt
Anthropiccache_control, top level (automatic) or per block (up to 4)0.1x1.25x for 5 minutes, 2x for 1 hour1,024 to 4,096 tokens, by model
OpenAIautomatic; explicit breakpoints on GPT-5.6 and later0.25x or 0.5x, by modelfree before GPT-5.6, 1.25x from GPT-5.61,024 tokens
Google Geminiimplicit on 2.5 and later, or explicit cache_control0.25xfree implicit; input price plus 5 minutes of storage explicit1,024 on 2.5 Flash, 4,096 on 2.5 Pro
DeepSeekautomatic0.1x1xnot listed
Alibaba Qwenexplicit cache_control only0.1x1.25xnot listed
Grokautomatic0.25xfreenot listed
Moonshot AIautomatic0.25xfreenot listed
Groqautomatic, Kimi K2 models only0.5xfreenot listed
Z.AIautomaticabout 0.2xfree for nownot listed
Figure 1: cache read and write prices per provider, as multiples of normal input price.

Reads run from a tenth of the input price to half of it, and writes from free to double. Five of the nine have no minimum prompt length listed in OpenRouter's docs. And with model fallbacks or the Auto Router, one request can land on a different row from the last, with a cold cache.

Take control of your LLM traffic
Route models, enforce budgets, and monitor every request from your own infrastructure.

The break-even rule

Whether caching saves money comes down to how many times the same prefix gets sent while the cache is still warm. Using the multipliers above, and counting only the cached part of the prompt:

  • Anthropic's 5-minute cache costs 1.25x on the first send and 0.1x on every reuse. Two sends cost 1.35x against 2x uncached, so it pays from the first reuse.
  • Anthropic's 1-hour cache costs 2x on the first send. Two sends cost 2.1x, slightly worse than not caching. Three sends cost 2.2x against 3x. It pays from the second reuse.
  • OpenAI on GPT-5.6, reading at 0.5x, costs 1.75x for two sends against 2x. It pays from the first reuse and loses 25% if there is no reuse at all.
  • DeepSeek, and every provider with free writes, can't lose. The first send costs what it would have anyway.

The rest of each request, the user's question and anything after the last cached block, bills the same either way, so these ratios apply to the prefix only.

Figure 2: where each cache option crosses the no-caching line.

When caching costs more than it saves

For most of the table a cache write is free or close to it, so caching is a free option. GPT-5.6 changed that for OpenAI. From that family on, cache writes bill at 1.25x the input price under automatic caching too, and automatic caching is on whether you asked for it or not. Any prompt of 1,024 tokens or more that you send once pays for the write and never collects on the read.

Take a 2,000-token prompt sent once. Without caching it bills as 2,000 tokens of input. With caching on, you pay close to what a 2,500-token prompt would cost. On a nightly job summarising a million long, unique documents, where the document comes first and nothing at the start ever repeats, that's a 25% surcharge on most of your input.

The fix reads backwards. Set prompt_cache_options.mode to "explicit" and don't mark any content block with prompt_cache_breakpoint. Explicit mode switches off OpenAI's automatic breakpoints, and with no explicit ones there's no cacheable prefix, so nothing gets written. cache_write_tokens comes back as zero, which is the one place in this post where a zero means what you want it to.

When it pays: a worked example

Now the other side. Claude Sonnet 4.6 with a 5,000-token system prompt, 100 requests across a working day, arriving in 20 bursts of five with bursts about 20 minutes apart. Counting only the system prompt, in tokens billed at the normal input price:

SetupWhat gets billedTotal
No caching100 requests at 5,000500,000
5-minute cacheper burst, one write at 6,250 and four reads at 500, times 20165,000
1-hour cacheone write at 10,000, then 99 reads at 50059,500

The 5-minute cache saves two thirds. The 1-hour cache saves 88%, because a 20-minute gap kills the short cache every time and never reaches the long one. Anthropic refreshes the cache lifetime each time it's read, so the 1-hour entry stays warm all day.

Change the traffic and the answer changes. If those 100 requests arrived evenly, one every four minutes or so, the 5-minute cache would never expire either, and the 1-hour option would mean paying 10,000 for a write the short cache handles at 6,250. Pick the TTL from the gaps between your requests, not from how many you send.

Token floors fail silently

Every provider sets a minimum prompt length below which nothing caches. For Claude, per OpenRouter's docs:

  • 4,096 tokens: Opus 4.5, 4.6, 4.7 and 4.8, and Haiku 4.5
  • 2,048 tokens: Haiku 3.5
  • 1,024 tokens: Sonnet 4, 4.5 and 4.6, and Opus 4 and 4.1

OpenAI's floor is 1,024. Gemini's is 1,024 on 2.5 Flash and 4,096 on 2.5 Pro. Under the floor the request succeeds, bills at full input price, and comes back with its cache fields set to zero. No error, no warning.

Two consequences of this are easy to miss.

Newer models have higher floors. A 3,000-token system prompt caches on Sonnet 4.6 and stops caching the day you move it to Opus 4.7, with no code change and nothing in the response to tell you. Moving from Haiku 3.5 to Haiku 4.5 does the same for anything between 2,048 and 4,096 tokens. A model upgrade is a caching change, so check your cache hit rate after one.

And you can't verify caching by looking for cache fields. That's what caught our own test. OpenRouter returns cached_tokens and cache_write_tokens on responses from every provider, including ones where caching isn't native, so the fields being there proves nothing. Zeros don't prove much either. They mean the marker was ignored or the prompt was under the floor, and on Haiku 4.5 at 3,600 tokens it was the floor.

A test that tells you something:

  1. Use a prompt comfortably above the floor for that exact model. Token counts vary between tokenizers, so leave headroom.
  2. Send it twice in quick succession.
  3. On providers that charge for writes, expect cache_write_tokens above zero on the first call. Expect cached_tokens above zero on the second.
  4. Read cache_discount in the response body. It's negative on an Anthropic write and positive on a read.

TTLs don't survive routing

OpenRouter translates cache markers between providers, which is useful. A text block with Anthropic-style cache_control becomes a prompt_cache_breakpoint when routed to a GPT-5.6 model, and a prompt_cache_breakpoint becomes a default cache_control on Anthropic or Google.

The TTL doesn't come along. A ttl on cache_control is dropped on the way to OpenAI. prompt_cache_options, where OpenAI's TTL lives, only applies to OpenAI. prompt_cache_breakpoint has no TTL field, so whatever it becomes on the Anthropic side gets the 5-minute default.

Lifetimes also behave differently once they're set. Anthropic's refreshes on every read. Gemini's explicit cache lasts five minutes from the write and doesn't refresh. Gemini's implicit cache is documented as three to five minutes on average, varying. So one request body asking for an hour can get an hour, five fixed minutes, or whatever OpenAI applies, depending on where it lands. The Responses API adds one more gap: it doesn't expose per-block Anthropic cache_control, so use Chat Completions or the Messages endpoint if you need a TTL on a specific block.

Sticky routing, and when a pin overrides it

A cache only helps if your next request reaches the provider holding it. OpenRouter handles that with provider sticky routing: after a request that uses caching, follow-up requests for the same model go to the same provider endpoint.

Stickiness is tracked per account, per model and per conversation, where a conversation is identified by hashing your first system message and first non-system message. It expires after 10 minutes without a request.

Agents tend to break that default, because their opening messages change between turns. Pass a session_id in the body or an x-session-id header and OpenRouter uses it as the sticky key instead, starting from the first successful request rather than waiting for a cache hit. OpenAI's prompt_cache_key works as a fallback key if you already send it.

Setting provider.order turns sticky routing off and your list takes priority. While your first choice is healthy that costs little, since a pin keeps you on one provider anyway. The difference shows after a failover, when sticky routing would follow the provider that served you and a pin starts again from the top of your list. Pinning is still right when you need predictable capacity, as we argued in our post on OpenRouter rate limits. The Auto Router has a similar edge: if the model it picked drops out of its candidates on a later turn, the cache goes with it.

OpenRouter's own response cache

Separate from all of the above, OpenRouter will cache whole responses if you send X-OpenRouter-Cache: true. A hit returns the stored response without calling a provider, bills nothing, and doesn't count against provider rate limits.

We measured it across three runs of 100 identical re-requests each, on GPT-4o mini and Claude Haiku 4.5. Every re-request was a hit. Median latency was 2,867 ms on a miss and 59 ms on a hit, and Haiku hits were quicker at 39 ms against 64 to 69 ms for GPT-4o mini. On the run billed to credits, 100 misses cost $0.0174 and 100 hits cost nothing.

Entries are keyed on the API key plus model, endpoint type, streaming mode and full request body, so services with separate keys share nothing and any parameter change is a miss. Our TTL probe, one sample per interval, found entries alive at 120 seconds and gone by 300. There's no published TTL, no semantic matching and no documented purge. It absorbs retries and duplicate submissions. An hourly repeat misses.

How TrueFoundry handles caching

Our gateway caches at the response layer in two modes. Exact-match returns a stored response for an identical request. Semantic mode embeds the last message and returns a stored response when an earlier request was close enough, and serves exact-match hits too, so you run one mode, not both. It's switched on with a request header, backed by Redis, and composes with provider prompt caching: on a gateway miss, the provider's own prefix cache still applies.

On SaaS the embedding model is text-embedding-3-small at 1,536 dimensions and isn't configurable; on-prem you choose it gateway-wide. Semantic matching has a failure mode that exact-match doesn't. Our semantic caching guide opens with it: two customers ask about their delivery minutes apart, the questions embed to nearly the same vector, and the second customer gets an answer about the first one's order. That is a scope bug. Moving the similarity threshold changes how often it happens, and it does not stop a cache that is shared across customers.

The cache lookup adds roughly 3 to 4 ms, and the gateway handles 350+ RPS on a single vCPU.

A short checklist

  • Check the minimum prompt length for the exact model you're calling, and check it again after every model upgrade.
  • Put stable content first and anything per-request last, so the prefix actually repeats.
  • On GPT-5.6 and later, switch caching off for prompts you'll only send once.
  • Choose the TTL from the gaps between requests: 5 minutes for steady traffic, 1 hour for bursts more than five minutes apart.
  • Send a session_id for agent sessions so routing stays on the provider with your cache.
  • Verify caching with non-zero cached_tokens, never with the field being present.
Try TrueFoundry AI Gateway
Connect your models and start managing LLM traffic through one API.

Related reading

Conclusion

OpenRouter prompt caching works, and the docs are clearer here than they are for a lot of the surrounding features. You still have nine pricing schemes, floors that rise on newer models, lifetimes that change in translation, and a response that looks the same whether caching happened or not. Check the floor, put the stable text first, and verify with numbers rather than fields.

To configure caching once for every provider, with semantic matching on top of the provider's own prefix cache, see how the TrueFoundry AI Gateway handles caching.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Inscríbase
Tabla de contenido

Controle, implemente y rastree la IA en su propia infraestructura

Reserva 30 minutos con nuestro Experto en IA

Reserve una demostración

La forma más rápida de crear, gobernar y escalar su IA

Demo del libro
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Descubra más

No se ha encontrado ningún artículo.
October 8, 2026
|
5 minutos de lectura

La seguridad de los agentes es un problema de sistemas: desde la inyección de prompts hasta el control en tiempo de ejecución

No se ha encontrado ningún artículo.
October 8, 2026
|
5 minutos de lectura

Intervención humana en MCP: TrueFoundry frente a Kong

comparación
October 8, 2026
|
5 minutos de lectura

¿Qué es un marco de gobierno de la IA?

No se ha encontrado ningún artículo.
October 8, 2026
|
5 minutos de lectura

Loop Engineering at Enterprise Grade: From Laptop Loops to Governed Runtimes

Liderazgo intelectual
No se ha encontrado ningún artículo.

Blogs recientes

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Preguntas frecuentes

What is OpenRouter prompt caching?

It's OpenRouter passing through each provider's prompt caching, where a repeated prompt prefix is stored and billed at a discount on later requests. Some providers cache automatically, while Anthropic and Alibaba need cache_control markers on the content you want cached. Prices differ by provider, from free writes and 0.1x reads to 2x writes on Anthropic's 1-hour cache.

Does cache_control work through OpenRouter?

Yes, as long as the prompt clears the model's minimum length. Claude Haiku 4.5 and Opus 4.5 and later need 4,096 tokens, and Sonnet 4.x needs 1,024. Below that floor the request succeeds at full price with cache fields at zero, which looks exactly like the marker being ignored.

Why did my bill go up after turning on caching?

Most likely you're paying for cache writes on prompts that never get reused. GPT-5.6 and later bill writes at 1.25x even under automatic caching, and Anthropic's 1-hour cache bills writes at 2x. If the prefix doesn't repeat while the cache is warm, you pay for the write and never get the discount.

¿Puedo desplegar TrueFoundry en mi propia VPC o en on-prem?

Sí. TrueFoundry se ejecuta en su VPC, on-prem, en entornos air-gapped o híbridos, de modo que los prompts y las respuestas nunca salen de su dominio, incluso cuando enruta entre muchos proveedores.

Realice un recorrido rápido por el producto
Comience el recorrido por el producto
Visita guiada por el producto