Prompt Caching
Prompt caching reduces processing time and costs by reusing previously computed prefixes. When you send the same or similar prompts repeatedly (e.g. alarge system prompt, a shared context document, or a set of tool definitions), the provider can skip reprocessing the cached portion and only compute the new tokens. This leads to lower latency and reduced costs on subsequent requests.
You can send
cache_control for any model. The AI Gateway handles the per provider caching logic internally, forwarding it to providers that read cache_control and stripping it for providers that cache prefixes on their own. This means you can write one request and use it across providers without worrying about compatibility.How to Enable
Add"cache_control": {"type": "ephemeral"} to any system prompt, user message content block, or tool definition:
Enabling Caching with a Header
If you can’t edit the request body — for example when the payload is built by a framework or an off-the-shelf agent — send thex-tfy-cache-control header instead. The AI Gateway places a single cache breakpoint at the end of the prompt, so the entire prefix (tool definitions, system prompt, and earlier messages) is cached.
cache_control field you would have written in the body:
The header is a hint, and the AI Gateway only acts on it where it changes anything:
- On Claude models from Anthropic, Claude Platform on AWS, Azure AI Foundry, AWS Bedrock, Google Vertex, and Databricks, the breakpoint is injected in the format that provider reads.
- On OpenAI, Gemini, Groq, xAI and other providers that cache repeated prefixes on their own, nothing is injected and the request is forwarded unchanged — see Provider Support below.
- A
cache_controlin the body always wins. If the request already marks any block, the header is ignored, so hand-placed breakpoints are never overwritten.
The value must be valid JSON and, if you set
ttl, it must match the 5m / 1h / 2d format — otherwise the AI Gateway rejects the request with a 400.x-tfy-cache-control is not the same as x-tfy-cache-config, which turns on the AI Gateway’s own exact-match and semantic response cache. They work at different layers and can be used together.Provider Support
There are three caching styles. The first two are marker-based — you sendcache_control and the provider caches based on it (this uses the same explicit vs automatic distinction as the Messages API):
- Explicit — you place
cache_controlon individual blocks (asystemblock, a message content block, or atooldefinition). - Automatic — you place a single
cache_controlfield at the top level of the request, or send thex-tfy-cache-controlheader, and the AI Gateway puts the breakpoint at the end of the prompt for you. - Provider-managed — the provider caches repeated prefixes on its own with no markup, so the AI Gateway strips any
cache_controlyou send.
The examples above use explicit caching (block-level
cache_control), which works on every Claude provider. On chat completions, a top-level cache_control is never sent to the provider as-is — the AI Gateway relocates it onto the last cacheable block, which is the one form every marker-based provider honours. The x-tfy-cache-control header behaves the same way. See Messages API caching for how the same request looks on /messages.Anthropic (and Bedrock / Vertex / Azure / Databricks Claude)
Anthropic (and Bedrock / Vertex / Azure / Databricks Claude)
For Anthropic (direct), Google Vertex (Claude), Azure AI Foundry (Claude), and Databricks (Claude), Anthropic enforces a minimum content length for caching to take effect; shorter prompts accept the
cache_control is forwarded to the provider unchanged. You can also include an optional ttl field (e.g. "5m" or "1h") to control cache duration.For AWS Bedrock, the AI Gateway translates each cache_control block into Bedrock’s native Converse cachePoint format, so your request code stays the same. Converse cachePoint markers do not carry a duration, so any ttl you set is ignored on Bedrock chat completions.On the Messages API (
/messages), Bedrock Claude models are served through the InvokeModel API with the native Anthropic body instead of Converse. There, cache_control (and ttl) is forwarded unchanged rather than translated to cachePoint.cache_control hint but are not cached:For Amazon Titan and Nova models on Bedrock,
cache_control on tool definitions is automatically skipped since these models do not support cache points on tools.OpenAI / Azure OpenAI
OpenAI / Azure OpenAI
OpenAI and Azure OpenAI are provider-managed — they cache matching prefixes on their own. No
cache_control markup is needed, and the AI Gateway strips it before forwarding.You can optionally pass prompt_cache_key to group requests that share a common prefix, improving cache hit rates:prompt_cache_key is only supported for OpenAI and Azure OpenAI.Gemini / Groq / xAI
Gemini / Groq / xAI
These providers are provider-managed — they handle caching on their own end. No
cache_control markup is needed, and the AI Gateway strips it before forwarding.Cached token counts are still reported in the response usage when the provider returns them.Cache Usage in Responses
When caching is active, the responseusage object includes cached token counts:
prompt_tokens_details.cached_tokens: tokens served from cache. Available across all providers that report cache usage.cache_read_input_tokens/cache_creation_input_tokens: Anthropic style fields, present for Anthropic, Bedrock, and Groq when values are non zero.
Reasoning Models
TrueFoundry AI Gateway provides access to model reasoning processes through thinking/reasoning tokens, available for models from multiple providers includingAnthropic,OpenAI,Azure OpenAI,Groq, xAI and Vertex.
These models expose their internal reasoning process, allowing you to see how they arrive at conclusions. The thinking/reasoning tokens provide step-by-step insights into the model’s cognitive process.
Supported Reasoning Models
OpenAI
OpenAI
Supported models:
o4-mini, o4-preview, o3 model family, o1 model family, gpt-5-mini, gpt-5-nano, gpt-5To reach the model’s reasoning, the AI Gateway serves these requests through OpenAI’s Responses API upstream while keeping your request and response on the Chat Completions contract. See Reasoning on OpenAI and Azure OpenAI.
Azure OpenAI
Azure OpenAI
Supported models:
gpt-5, gpt-5-mini, gpt-5-nano, o3-pro, codex-mini, o4-mini, o3, o3-mini, o1, o1-miniAzure deployments need an
api-version of 2025-03-01-preview or newer to return reasoning. See Reasoning on OpenAI and Azure OpenAI.Anthropic
Anthropic
Supported models:
viaUsing Direct API Calls with Native
For more precise control with Anthropic models, you can use the native
Claude Opus 4.1 (claude-opus-4-1-20250805), Claude Opus 4 (claude-opus-4-20250514), Claude Sonnet 4 (claude-sonnet-4-20250514), Claude Sonnet 3.7 (claude-3-7-sonnet-20250219) via
Anthropic, AWS Bedrock, and Google Vertex AIUsing OpenAI SDK
For Anthropic models (from Anthropic, Google Vertex AI, AWS Bedrock), TrueFoundry automatically translates the
reasoning_effort parameter into Anthropic’s native thinking parameter format since Anthropic doesn’t support the reasoning_effort parameter directly.The translation uses the max_tokens parameter with the following ratios:none: 0% of max_tokenslow: 30% of max_tokensmedium: 60% of max_tokenshigh: 90% of max_tokens
Using Direct API Calls with Native thinking Parameter
For more precise control with Anthropic models, you can use the native thinking parameter directly:Groq
Groq
Supported models:
OpenAI GPT-OSS 20B (openai/gpt-oss-20b), OpenAI GPT-OSS 120B (openai/gpt-oss-120b), Qwen 3 32B (qwen/qwen3-32b), DeepSeek R1 Distil Llama 70B (deepseek-r1-distill-llama-70b)xAI
xAI
Supported models:
grok-3-mini (with reasoning_effort parameter), grok-4-0709, grok-4-1-fast-reasoning, grok-4-fast-reasoning (reasoning built-in)For grok-3-mini, you can use the reasoning_effort parameter to control reasoning depth. Other Grok models like grok-4-0709 have reasoning capabilities built-in but do not support the reasoning_effort parameter.The
reasoning_effort parameter is only supported for grok-3-mini. For other Grok models like grok-4-0709 and grok-4-1-fast-reasoning, reasoning is built-in and the reasoning_effort parameter should not be used. Reasoning tokens are included in the usage metrics for all reasoning-capable models.Parameter Restrictions: Reasoning models (like grok-4-0709 and grok-4-1-fast-reasoning) do not support presence_penalty, frequency_penalty, or stop parameters. Using these parameters with reasoning models will result in an error.Gemini
Gemini
Supported models: All Using Direct API Calls with Native
For more precise control with Gemini models, you can use the native
Gemini 2.5 Series Models.These models can be accessed from Google Vertex or Google Gemini ProvidersFor Gemini models (from Anthropic, Google Vertex AI, AWS Bedrock), TrueFoundry automatically translates the
reasoning_effort parameter into Gemini’s native thinking parameter format since Gemini doesn’t support the reasoning_effort parameter directly.The translation uses the max_tokens parameter with the following ratios:none: 0% of max_tokenslow: 30% of max_tokensmedium: 60% of max_tokenshigh: 90% of max_tokens
Using Direct API Calls with Native thinking Parameter
For more precise control with Gemini models, you can use the native thinking parameter directly:Reasoning on OpenAI and Azure OpenAI
OpenAI exposes a model’s reasoning only through its Responses API — a plain Chat Completions call togpt-5.x cannot return it. So when a request needs reasoning, the AI Gateway serves it through the provider’s Responses API upstream and converts the reply back.
Your code does not change. You keep sending Chat Completions requests and reading Chat Completions responses; the switch happens inside the AI Gateway.
When it applies
The model must support the Responses API. Beyond that, any one of these turns it on:
That last row is worth calling out: responses-only models used to be rejected on
/chat/completions. They now work, so you can reach them through the same OpenAI SDK client as every other model.
Use the header when you want reasoning behaviour without changing the body — for example when the payload is built by a framework or an off-the-shelf agent:
What comes back
thinking_blocks here carry id and encrypted_content instead of the thinking and signature fields you get from Claude and Gemini. Both shapes are replayed the same way — see Multi-Turn Conversations.
When streaming, reasoning_content arrives as deltas before the content, and the thinking_blocks entry is emitted as its own chunk once the reasoning item is complete.
Multi-turn conversations
Echo the assistant message back unchanged on the next turn. Without it the model starts each turn from scratch, which matters most in agent loops where the model builds on its earlier thinking.A thinking block is only replayable if it still has both its
id and encrypted_content. Blocks missing either are skipped, so serialize the assistant message as returned rather than rebuilding it by hand.Parameters that are dropped
The Responses API has no equivalent for some Chat Completions parameters. When a request is served this way, these are dropped and the request still goes through — reasoning is never silently disabled to preserve them:n (above 1), best_of, frequency_penalty, presence_penalty, logit_bias, logprobs, top_logprobs, stop, seed, audio, modalities (non-text), prediction
Inert values are ignored rather than dropped, so n: 1, logprobs: false, and zeroed penalties never change anything.
Azure OpenAI requirements
Two extra conditions apply on Azure:- The deployment’s
api-versionmust be2025-03-01-previewor newer. Older versions reject/responsesoutright, so the AI Gateway keeps those requests on Chat Completions. - The deployment’s foundation model must be listed as Responses-capable in the model catalogue. Azure deployment names are arbitrary, so the AI Gateway matches on the foundation model rather than the deployment name.