What is OpenRouter? A technical guide to what happens to your request

Diseñado para la velocidad: ~ 10 ms de latencia, incluso bajo carga
¡Una forma increíblemente rápida de crear, rastrear e implementar sus modelos!
- Gestiona más de 350 RPS en solo 1 vCPU, sin necesidad de ajustes
- Listo para la producción con soporte empresarial completo
Five model providers means five SDKs, five API keys, five invoices and five sets of rate limits. OpenRouter collapses that into one endpoint. Point the OpenAI SDK at a different base URL, change the model string, and the same client can call Claude, Gemini, Llama or DeepSeek.
That much is in every quickstart. Underneath, OpenRouter is a marketplace. It picks which company serves each request, holds your money, and pays the providers. We traced what happens between your code and the model, with numbers from our own testing where we have them.
What is OpenRouter?
OpenRouter is a hosted HTTP API that routes inference requests to third-party model providers through a single endpoint and credential, and bills you for model usage by default. It accepts three request formats: OpenAI Chat Completions, the OpenAI Responses API, and Anthropic's Messages API at /api/v1/messages. There are client SDKs for TypeScript, Python and Go, but they're a typed layer over the same REST API, so the plain OpenAI SDK with a changed baseURL works just as well.
OpenRouter is two-sided. On one side it recruits providers, sets the contract they integrate against, and pays them. On the other it holds your prepaid credits and consolidates billing across all of them. It sits in your payment path as well as your request path, which is what makes it convenient, and also why leaving takes more than changing a URL.
You are renting access to other companies' models. There is no self-hosted or air-gapped build of OpenRouter, and the core product is closed source. The SDKs and docs are public. It started as Window, a Chrome extension for using several models at once, and since then has added server-side tools such as web search, web fetch and sandboxed shell execution that models can call mid-request.
How a request flows through OpenRouter
One request goes through these steps:
- Your app sends it to
https://openrouter.ai/api/v1with one API key, in any of the three formats. - OpenRouter's edge, running on Cloudflare Workers, authenticates the key and checks your credit balance.
- The model string, in
vendor/modelform, resolves to a list of provider endpoints that serve that model. - OpenRouter picks one. By default it skips providers with an outage in the last 30 seconds and weights the rest by the inverse square of their price. Your
sort,order,onlyand data-policy settings override that. - The request is translated for that provider and forwarded, and the response is normalised back to the format you sent.
- Usage is billed: deducted from your credits at the provider's list price, or charged to your own provider key if you've attached one.
- If the provider fails, OpenRouter tries the next one. The
attemptfield in the response is the only in-band sign that happened.

One model name, eleven different deployments
The model string does not tell you who runs the model. The vendor prefix says whose model it is, not whose servers it's on. When we tested anthropic/claude-3-haiku, it was served by Amazon Bedrock.
For open-weight models the gap gets wider. On 5 October, OpenRouter's API listed 11 endpoints for meta-llama/llama-3.3-70b-instruct. Across those 11:
- three declared quantisations (bf16, fp16 and fp8), plus five endpoints that don't declare one
- context windows from 12,288 to 131,072 tokens
- maximum output from 2,048 to 128,000 tokens
- input prices from $0.10 to $1.04 per million tokens
- tool calling declared on 7 of them

Same name, materially different products. If your prompts run long, the 12K endpoint and the 131K endpoint aren't interchangeable, and the default router chooses partly on price. OpenRouter does filter by capability when a request needs it: when we sent eight requests with tool_choice: required and no provider pin, every one landed on a host that declares tool support. But declared isn't the same as working. One host that declares tool support returned no tool call at all when we pinned to it and forced one.
Read /api/v1/models/{model}/endpoints before trusting a model string, and pin or filter providers wherever quantisation, context length or tool behaviour matters.
Key components
The catalog
OpenRouter's API listed 466 models across 65 vendor namespaces on 5 October, up from 428 when we first pulled it on 8 September. Embedding models sit on a separate endpoint, /api/v1/embeddings/models, which listed 33 more. Add those and you're close to the "500+ models" on OpenRouter's pricing page. Provider counts line up less neatly: the providers endpoint returned 112 entries, one of them called FakeProvider, against "80+" in the marketing. Query the API rather than trusting a count in a blog post, this one included.
Routing
The default route spreads traffic. It does not optimise for price or speed. In our September test on Llama 3.3 70B, with eight requests per strategy, the default was the slowest option at a median 935 ms and cost twice as much as sort: price. sort: throughput was fastest at 289 ms, ahead of sort: latency at 347 ms, and the cheapest and most expensive strategies were 10x apart for the same prompt. Eight requests is enough to show the directives work and differ. It is too small a sample to rank hosts. If cost or speed matters to you, set a sort.
Two details about fallbacks. A misspelt provider slug in order is dropped silently rather than rejected, so a typo just narrows your list. And as noted above, attempt is the only signal a fallback fired.
API compatibility
The OpenAI SDK works as a drop-in, and for most chat traffic that's all there is to it. The edge cases are parameters a particular route doesn't support, which tend to be accepted with a 200 and not applied. We sent OpenAI's prediction parameter to GPT-4o mini, which declares support for it, and got a normal response with no sign it took effect. If a parameter changes your output or your bill, check the response for evidence that it worked.
Billing
You pay the provider's list price per token. OpenRouter does not add a markup on top. Its revenue is a fee when you buy credits: 5.5% on the Standard plan with a $0.80 minimum, 8% on Business, and 5% for crypto. With your own provider key, the first $25,000 a month of list-price usage carries no fee, and 5% after that. Credits can expire a year after purchase. The pricing has changed twice in fourteen months, so check OpenRouter's pricing page before modelling costs. Our OpenRouter pricing breakdown goes further.
Data controls
Prompts and responses always pass through OpenRouter. You cannot run the service yourself or keep it air-gapped. The controls available on the hosted service are zero data retention routing, a data_collection: "deny" filter that only routes to providers that don't collect your data, and in-region routing. In-region routing used to need an Enterprise contract. Since September it's on the self-serve Business plan: send requests to eu.openrouter.ai or us.openrouter.ai and they're processed only in that region, failing with an error if no in-region endpoint exists. The documented trade-offs: OpenRouter's own prompt and response logging is skipped on regional endpoints, and Batch, web fetch, sandboxed shell, file and image tools, and the Auto Router aren't available there.
Rate limits and caching
Both are covered in their own posts. OpenRouter rate limits covers why your free-model ceiling depends on what you've paid, and OpenRouter prompt caching covers why the same caching feature has nine different prices.
How much latency does OpenRouter add?
OpenRouter's docs say it runs at the edge to keep overhead minimal. Its homepage used to put a number on that, around 25 ms, and plenty of guides still repeat it. As of 5 October the homepage just says "minimal latency". The docs add one detail: when your credit balance runs low, OpenRouter expires its edge caches more aggressively to keep billing accurate, which adds latency until you top up.
We measured it. The same 100 prompts went three ways: direct to the provider, through OpenRouter on credits, and through OpenRouter with our own provider key. Requests were interleaved and paired per prompt, so prompt content cancels out of the difference, and each one was pinned to the model's first-party provider and checked per request.

On GPT-4o mini we couldn't measure any overhead. The median added time to first token was 19.7 ms on credits, but the middle half of results ran from 85 ms faster to 157 ms slower, so the true figure is below what our setup can resolve. On Claude Haiku 4.5 it was unambiguous: 206 ms added on credits and 124 ms with our own key, with the interquartile range well clear of zero. Without streaming, across two runs, credits added 837 and 883 ms to the full response while our own key stayed within 42 ms of going direct.
Credits and your own key pass through the same gateway and the same translation, so what differs is the upstream account. Our best explanation, and it is an inference, is contention on OpenRouter's shared Anthropic capacity. On a busy provider, bringing your own key buys latency as well as billing control.
All of this ran from one laptop on consumer wifi through OpenRouter's Los Angeles edge, in September 2026. Read the differences as real and the absolute numbers as specific to our setup.
Who runs OpenRouter
OpenRouter was founded in 2023 by Alex Atallah, previously a co-founder of OpenSea. On 19 August 2026, Stripe announced an agreement to acquire it. As of early October the deal hadn't been reported as closed, and neither company has disclosed a price; press reports put it above $7 billion. OpenRouter has said its product and commitments won't change.
OpenRouter's revenue is a fee on payments, its pricing will be set by a payments company, and its neutrality between providers is a stated commitment rather than something built into its structure. A team about to build on it for two years should have that in the evaluation, next to the routing and latency numbers, and then decide.
OpenRouter vs an LLM gateway you run
OpenRouter gets compared with LLM gateways like LiteLLM or ours. Both put one API in front of many models. They are different products. Our explainer on what an LLM gateway is covers the general idea.
Two rows need a sentence each. OpenRouter's control plane and data plane are the same system, so when it's down, everything behind it is too. Its own engineering blog describes a roughly 50-minute database outage in August 2025 that took routing down, and argues that aggregate uptime still beats any single provider; in our own testing across roughly 4,300 requests we didn't see a single 4xx or 5xx. And on identity: keys are owned by the member who created them, and deactivating that person through SCIM deactivates their keys, which is the right default for people and an awkward one for production workloads.
When OpenRouter is a good fit
OpenRouter fits prototyping across many models, when you want one key and one bill and would rather skip a contract with every provider you try. It also fits open-weight models, where choosing the host by price or throughput saves real money, as long as your prompts can leave your network, with or without a regional constraint.
It is a poor fit when prompts have to stay on your network, when workloads need machine identities rather than keys owned by people, or when one vendor's outage taking every model offline is unacceptable. The same applies if latency matters on a contended provider and you can't bring your own key, or if procurement will want direct contracts with each provider anyway.
How TrueFoundry approaches this
The TrueFoundry AI Gateway runs in your VPC, on-prem or air-gapped, and calls providers directly with your own contracts and keys. It exposes 1,000+ LLMs through one OpenAI-compatible API, adds roughly 3 to 4 ms of overhead, and handles 350+ RPS on a single vCPU.
You can keep OpenRouter in that setup. It is supported as a provider in TrueFoundry, so the catalog stays available for experiments while your own rate limits, budgets, access control and logging sit in front of it. For teams already on OpenRouter, that is usually the lowest-risk first step.
Related reading
- OpenRouter alternatives, comparing the options for production teams
- LiteLLM vs OpenRouter, self-hosted proxy against hosted router
- Requesty vs OpenRouter, two hosted routers side by side
- What is an LLM router?, on routing between models rather than providers
- What is an LLM gateway?, the general architecture
Conclusion
OpenRouter is one API in front of a marketplace. It picks who serves each request, holds your credits, and pays the providers. Before you treat a model string as a stable product, read the endpoints list, set a routing sort, and measure the overhead on the providers you actually use.
If you want that single API inside your own infrastructure, see how the TrueFoundry AI Gateway works.
TrueFoundry AI Gateway ofrece una latencia de entre 3 y 4 ms, gestiona más de 350 RPS en una vCPU, se escala horizontalmente con facilidad y está listo para la producción, mientras que LitellM presenta una latencia alta, tiene dificultades para superar un RPS moderado, carece de escalado integrado y es ideal para cargas de trabajo ligeras o de prototipos.



Controle, implemente y rastree la IA en su propia infraestructura
Blogs recientes
Preguntas frecuentes
What is OpenRouter?
OpenRouter is a hosted, OpenAI-compatible API that gives you access to hundreds of AI models from many providers through one endpoint and one API key. It picks which provider serves each request, bills you at the provider's list price, and charges a fee when you buy credits. You cannot run it yourself, so prompts always pass through OpenRouter's infrastructure.
Is OpenRouter free?
Partly. OpenRouter offers free variants of some models, capped at 20 requests per minute and 50 requests per day, rising to 1,000 a day once you've bought at least 10 credits. Paid models are billed at the provider's list price, plus a fee on each credit purchase. Our post on OpenRouter rate limits has the detail.
Does OpenRouter add latency?
It depends on the provider. In our paired testing, the added time to first token on GPT-4o mini was too small to measure, while Claude Haiku 4.5 on OpenRouter credits added about 206 ms, or about 124 ms with our own Anthropic key. OpenRouter itself no longer publishes an overhead figure.
¿Puedo desplegar TrueFoundry en mi propia VPC o en on-prem?
Sí. TrueFoundry se ejecuta en su VPC, on-prem, en entornos air-gapped o híbridos, de modo que los prompts y las respuestas nunca salen de su dominio, incluso cuando enruta entre muchos proveedores.










.webp)
.webp)
.png)



.png)
.png)
.png)

.png)
.png)
.png)
.png)
.png)
.png)





