OpenRouter BYOK explained: cheaper, often faster, and changing

Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
If you've read that OpenRouter BYOK is free for your first million requests a month, you've read a figure OpenRouter retired on 14 July 2026. It's still on some of OpenRouter's own older documentation pages, which is why so many guides still repeat it.
The current terms are better for most teams, and BYOK is the cheapest way to use OpenRouter. It was also faster for us on Anthropic. A failing key still falls back to OpenRouter's own account unless you change that on the key and restrict providers on the request.
What is OpenRouter BYOK?
BYOK, bring your own key, means adding your own provider credentials to OpenRouter. Requests routed to that provider use your key, the provider bills you directly, and your own rate limits and contract terms apply. OpenRouter still handles routing, translation and analytics, and your prompts still pass through it.
Keys are added per workspace, in the workspace's BYOK settings or through the management API. As of 5 October, 72 of the 92 providers OpenRouter lists support BYOK, including Anthropic, OpenAI, Google Vertex and AI Studio, Amazon Bedrock, Azure, DeepSeek, Mistral and Groq. Bedrock takes either a Bedrock API key or AWS credentials, Vertex a service account JSON, and Azure a resource name and key.
What OpenRouter BYOK costs
On the Standard and Business plans, the first $25,000 of BYOK usage each month carries no OpenRouter fee. On Enterprise the allowance is custom; some guides still quote $200,000, but OpenRouter's FAQ and pricing page now just say custom. Above the allowance, OpenRouter charges 5% of what the same model and provider would cost on OpenRouter, deducted from your OpenRouter credits.
Compare that with buying credits, where you pay OpenRouter list price for inference plus a fee on every purchase: 5.5% on Standard with a $0.80 minimum, or 8% on Business.

BYOK is cheaper at every volume, and the gap is widest below the allowance and on the Business plan. Three details change the numbers.
First, the BYOK fee comes out of OpenRouter credits, and buying those credits carries the 5.5% purchase fee. So above the allowance, the effective fee is about 5.3% of the overage rather than 5%, and you need to keep a credit balance even if all your inference runs on your own keys. The BYOK figures in the table include this.
Second, the fee is calculated on OpenRouter's list price, not on what you pay your provider. If you've negotiated 20% off with Anthropic, $100,000 of list-price usage costs you $80,000, and the $3,956 fee is about 5% of your actual bill rather than 4%.
Third, BYOK spend is easy to lose track of. It doesn't count toward guardrail or workspace budgets unless you turn on include_byok_in_budgets, so a budget can look comfortably under its cap while provider spend climbs. And the Activity page estimates BYOK spend at market rates, not your discount, so OpenRouter's dashboard, its budgets and your provider invoice can show three different numbers.
Is OpenRouter BYOK faster?
On Anthropic, it was in our tests. We sent the same prompts to Claude Haiku 4.5 directly, through OpenRouter on credits, and through OpenRouter with our own Anthropic key. Credits added about 200 ms to the first token, while our own key added about 120 ms. After the first token, our key streamed at the same speed as going direct, while credits generated noticeably more slowly. On a typical 290-token answer, that made our own key about 0.6 seconds faster. The full measurement and its limits are in our post on OpenRouter latency on Anthropic.
On OpenAI it went the other way. Credits reached the first token 23 to 50 ms sooner than our own key, and both streamed at the same rate. Whether BYOK is faster depends on the provider. The account serving your request affects speed in ways you can't see from outside.
BYOK also changes whose rate limits apply. On credits, OpenRouter manages your provider limits; with your own key, you get your provider account's limits, which can be higher or lower. Our post on OpenRouter rate limits covers the rest.
Where requests go when your key fails
When the key fails, the request still has to go somewhere.
By default, OpenRouter tries your key first. If it fails or hits a rate limit, OpenRouter falls back to its own shared capacity for that provider, billed to your credits. Each key has a shared capacity fallback setting with three levels: use shared capacity (the default), never use shared capacity for the models the key applies to, and never use shared capacity for any model on that provider.
The strongest level stops OpenRouter from using its own Anthropic account. It doesn't stop OpenRouter from sending the request to a different provider. If your Anthropic key fails and you ask for a Claude model, OpenRouter can serve it from Bedrock or Vertex on its own account and your credits. OpenRouter's BYOK guide says so explicitly, and the fix it gives is to restrict providers on the request, for example with provider.only.

Two more routing rules change where that request lands. Keys sit in two sections: prioritized keys are tried in order before OpenRouter's endpoints, and fallback keys are tried only after them, so a "backup" key in the fallback section runs after OpenRouter's account, not before it. And BYOK keys override your provider order. If you send order: ["amazon-bedrock", "google-vertex"] but only hold a key for Vertex, Vertex is tried first. OpenRouter's guide says there is currently no way to change this.
So if the reason you adopted BYOK is to keep inference on a contracted provider account, set the strongest fallback level on the key and restrict providers on every request. Either alone isn't enough.
BYOK and data policies
Your own key doesn't change which endpoints your data policies allow. If you've set data_collection: "deny" or enforced zero data retention, those rules still apply to your key's endpoints.
What BYOK adds is a way to tell OpenRouter about your own agreements. Each key has a provider agreement section where you can declare that your account has zero data retention, and on OpenAI and Azure keys, which region your account processes data in. Both are useful if you've negotiated those terms directly. Both are your own attestation: OpenRouter's documentation says it doesn't verify them, guardrails that restrict data regions don't recognise region declarations yet, and neither setting is in the management API yet.
On key storage, OpenRouter says provider keys are "securely encrypted". It doesn't document the encryption scheme, whether customer-managed keys are supported, or how rotation is handled.
Running BYOK in practice
BYOK is configured per workspace, with no per-request switch. Any API key in a workspace that has provider keys attached will route through BYOK, whether or not that's what the caller intended. When we compared credits against BYOK, it took two separate workspaces.
When we tested in September, the dashboard didn't show which billing path a given API key used. The reliable check is a single request that asks for routing metadata:
import os, requests
r = requests.post(
"https://openrouter.ai/api/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}",
"X-OpenRouter-Experimental-Metadata": "enabled"},
json={"model": "openai/gpt-4o-mini",
"messages": [{"role": "user", "content": "hi"}],
"max_tokens": 5,
"provider": {"order": ["openai"], "allow_fallbacks": False}},
)
meta = r.json().get("openrouter_metadata") or {}
print("is_byok:", meta.get("is_byok"))If it prints False for a key you expected to be on BYOK, you're paying credits. The header is labelled experimental, so check it still works before you build on it. When BYOK requests fail, the Activity page's raw metadata includes provider_responses, which shows the status code from each provider attempted, and is the quickest way to tell a revoked key from a rate limit or a permissions problem.
What's changing
BYOK pricing changed this summer, and OpenRouter has said it will change again.
- 14 July 2026: the free allowance was re-based from request count (1 million a month on pay-as-you-go) to list-price dollars ($25,000 a month). The 5% fee itself didn't change.
- Pending since June 2025: OpenRouter's own announcement of its current fee structure, last updated in June 2026, says the 5% BYOK fee "will be removed in the future and replaced with a fixed monthly subscription", with pricing to be decided.
- September 2026: the self-serve Business plan raised the credit fee to 8% for teams that need in-region routing, while keeping the same BYOK allowance, which widens BYOK's advantage on that plan.
- August 2026: Stripe agreed to acquire OpenRouter. As of early October the deal hadn't been reported as closed, but future pricing will be set by the new owner.
Model your costs on the current terms, and check them again each quarter. If the subscription arrives, the break-even against credits will move, especially for teams below the $25,000 allowance who currently pay OpenRouter nothing.
When BYOK is worth it
BYOK is worth it when you already have accounts with the providers you use, or you've negotiated rates or committed spend you want to draw down. It is also the faster route when latency on Claude matters. If your compliance terms live in your provider contract, it works only if you also restrict providers on each request.
Credits fit better when you want one bill and no provider relationships, which is what the credit fee pays for, or when you're experimenting across providers you'll never contract with. They also fit if you don't want to manage keys, rotation and per-provider limits yourself.
How TrueFoundry approaches this
The TrueFoundry AI Gateway calls providers with your own credentials, because it runs in your VPC or data centre and there is no shared provider account under it. A failing key falls back only to whatever your own routing config allows. The gateway exposes 1,000+ LLMs through one OpenAI-compatible API, adds roughly 3 to 4 ms of overhead, and handles 350+ RPS on a single vCPU.
If you're staying on OpenRouter, TrueFoundry can sit in front of it as a provider and keep budgets, rate limits and access control in one place you run. Our overview of how OpenRouter works covers when that setup fits.
Related reading
- OpenRouter alternatives, comparing the options for production teams
- OpenRouter free models, on what's free and what it costs you instead
- OpenRouter prompt caching, on when caching saves money and when it costs more
- What is an LLM gateway?, the general architecture
Conclusion
OpenRouter BYOK is the cheapest way to use OpenRouter at any volume, and on Claude it was the faster route in our tests. By default a failing key falls back to OpenRouter's capacity, and the request can still land on another provider unless you restrict providers on that request. Set the strongest fallback level, restrict providers on the request, turn on BYOK in your budgets, and re-check the terms when the promised subscription arrives.
To call providers with your own keys and no shared account underneath, see how the TrueFoundry AI Gateway works with your own provider keys.
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
Is OpenRouter BYOK free?
OpenRouter charges nothing for BYOK usage up to $25,000 a month at list price on the Standard and Business plans, and a custom amount on Enterprise. Above that it charges 5% of what the same usage would cost on OpenRouter, deducted from your credits. Your provider bills you separately for the inference itself.
Does BYOK make OpenRouter faster?
On Anthropic it did in our tests: our own key added about 120 ms to the first token against about 200 ms on credits, and then streamed at direct speed. On OpenAI, credits was slightly faster. Which route is faster depends on the provider and the account behind it.
What happens when my BYOK key hits a rate limit?
By default OpenRouter falls back to its own shared capacity for that provider, billed to your credits. Setting the key to never use shared capacity blocks that, but OpenRouter can still serve the request from a different provider unless you restrict providers on the request, for example with provider.only.
Can I deploy TrueFoundry in my own VPC or on-prem?
Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.
How many LLMs does TrueFoundry support?
1,000+ LLMs through a single OpenAI-compatible API. Switching models means changing the model name in the request — same URL, same credentials — which is what makes swapping in a new provider a configuration change rather than an integration project.










.png)
.png)
.png)
.png)
.png)





.png)



.png)





