What BYOK means on an AI Gateway
.png)
Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
BYOK, or bring your own key, means you keep your own accounts with model providers like OpenAI, Anthropic, and AWS Bedrock, and you plug those provider keys into the gateway once. Your applications never touch the raw provider keys again. They talk to one TrueFoundry AI Gateway endpoint using a single gateway key, and the gateway forwards each call to the right provider with your credentials.
This is the opposite of a reseller model where you buy tokens from a middleman. With BYOK you keep your own provider contracts, pricing, and rate limits, and you add centralized governance on top.
Why route provider keys through the gateway
Handing raw provider keys to every application and notebook is how key sprawl starts. Keys leak into code, get copied into environments, and become impossible to rotate without breaking something.
- One key for callers. Applications authenticate with a TrueFoundry key, not the provider key. As the docs put it, to access models through the gateway you use TrueFoundry API keys, not the original provider keys.
- Central rotation. Provider keys live in one place. Rotate them in the gateway and every application keeps working.
- Access control. You decide who can use which model account, with manager and user roles per account.
- One interface. Instead of managing separate SDKs, endpoints, and keys for OpenAI, Anthropic, Bedrock, and self-hosted models, applications talk to one gateway endpoint and use one gateway key.
Adding your provider keys: model accounts
In TrueFoundry, a provider key lives inside a model account. The quick start docs describe a model account as one account of a model provider, for example OpenAI, Anthropic, or AWS Bedrock. You can add multiple accounts per provider, each with their own API keys, and each account can have multiple models.
To add one, select the provider you want, then add models after providing the API key. You can add multiple keys from the same provider by creating separate model accounts, which is useful when you want to separate spend or rate limits by environment or team.

Once submitted, your model accounts and models appear under the Models tab. To add more later, go to AI Gateway, then Models, select a model account on the left, and click Add Model in the top right.

How callers authenticate: virtual keys
Once your provider keys are stored, applications never see them. The gateway access control docs describe two token types that callers use instead.
- Personal Access Tokens (PATs) are tied to a user and are recommended for developers during development.
- Virtual Account Tokens (VATs) are tied to a virtual identity and are recommended for production applications. This is the virtual key that your services ship with.
Access is configured at the model account level with two roles. A Model Account Manager can modify settings, add or remove models, and manage access permissions. A Model Account User can use all models within the account but cannot change settings or permissions. Tenant admins automatically have access to all models across the platform.
To call the gateway, you need three things from the Code Snippet tab of the Playground: the gateway base URL, an API key, and a model id. SaaS users point to the gateway base URL, authenticate with a PAT or VAT, and use the standard OpenAI client.

from openai import OpenAI
client = OpenAI(
api_key="your_truefoundry_api_key",
base_url="https://gateway.truefoundry.ai",
)
What you get once keys live in the gateway
- A single OpenAI compatible endpoint for every provider and every self-hosted model.
- Central key rotation without touching application code.
- Per account access control with manager and user roles.
- Centralized observability and policy, since every call flows through one place.
As one customer put it, the gateway eliminated the overhead of managing keys, routing logic, and scattered observability, and introducing new models became just configuration.
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
What does BYOK mean for an AI gateway?
BYOK means you keep your own provider accounts and keys and plug them into the gateway once. Applications then call the gateway with a single gateway key, and the gateway uses your stored provider keys to reach OpenAI, Anthropic, Bedrock, and others.
Do my applications ever see the provider key?
No. Callers authenticate with TrueFoundry keys, not the original provider keys. You store the provider key in a model account, and applications use a Personal Access Token or a Virtual Account Token instead.
Can I add more than one key for the same provider?
Yes. You can add multiple accounts per provider, each with their own API keys. This lets you separate spend, rate limits, or environments while keeping a single gateway endpoint.
What is a virtual key?
A Virtual Account Token is tied to a virtual identity rather than a person, which makes it the right choice for production applications. It lets a service authenticate to the gateway without embedding a user credential or a raw provider key.














.png)




.png)



.png)
.png)
.png)

.png)
.png)





