Blank white background with no objects or features visible.

Te ofrecemos acceso gratuito al informe completo Gartner Hype Cycle for AI Governance 2026. Consigue tu copia →

What is OpenRouter? A technical guide to what happens to your request

Por Kshitij Gupta

Published: October 7, 2026

TL;DR:

OpenRouter is a hosted, OpenAI-compatible API that sends your model requests to hundreds of third-party providers through one endpoint and one key, and bills you at provider list prices plus a fee on credit purchases. The model name you send doesn't decide who serves it; routing does. We followed a request through that path and measured what it adds.

Five model providers means five SDKs, five API keys, five invoices and five sets of rate limits. OpenRouter collapses that into one endpoint. Point the OpenAI SDK at a different base URL, change the model string, and the same client can call Claude, Gemini, Llama or DeepSeek.

That much is in every quickstart. Underneath, OpenRouter is a marketplace. It picks which company serves each request, holds your money, and pays the providers. We traced what happens between your code and the model, with numbers from our own testing where we have them.

What is OpenRouter?

OpenRouter is a hosted HTTP API that routes inference requests to third-party model providers through a single endpoint and credential, and bills you for model usage by default. It accepts three request formats: OpenAI Chat Completions, the OpenAI Responses API, and Anthropic's Messages API at /api/v1/messages. There are client SDKs for TypeScript, Python and Go, but they're a typed layer over the same REST API, so the plain OpenAI SDK with a changed baseURL works just as well.

OpenRouter is two-sided. On one side it recruits providers, sets the contract they integrate against, and pays them. On the other it holds your prepaid credits and consolidates billing across all of them. It sits in your payment path as well as your request path, which is what makes it convenient, and also why leaving takes more than changing a URL.

You are renting access to other companies' models. There is no self-hosted or air-gapped build of OpenRouter, and the core product is closed source. The SDKs and docs are public. It started as Window, a Chrome extension for using several models at once, and since then has added server-side tools such as web search, web fetch and sandboxed shell execution that models can call mid-request.

How a request flows through OpenRouter

One request goes through these steps:

  1. Your app sends it to https://openrouter.ai/api/v1 with one API key, in any of the three formats.
  2. OpenRouter's edge, running on Cloudflare Workers, authenticates the key and checks your credit balance.
  3. The model string, in vendor/model form, resolves to a list of provider endpoints that serve that model.
  4. OpenRouter picks one. By default it skips providers with an outage in the last 30 seconds and weights the rest by the inverse square of their price. Your sort, order, only and data-policy settings override that.
  5. The request is translated for that provider and forwarded, and the response is normalised back to the format you sent.
  6. Usage is billed: deducted from your credits at the provider's list price, or charged to your own provider key if you've attached one.
  7. If the provider fails, OpenRouter tries the next one. The attempt field in the response is the only in-band sign that happened.
Figure 1: the path of one request through OpenRouter.

One model name, eleven different deployments

The model string does not tell you who runs the model. The vendor prefix says whose model it is, not whose servers it's on. When we tested anthropic/claude-3-haiku, it was served by Amazon Bedrock.

For open-weight models the gap gets wider. On 5 October, OpenRouter's API listed 11 endpoints for meta-llama/llama-3.3-70b-instruct. Across those 11:

  • three declared quantisations (bf16, fp16 and fp8), plus five endpoints that don't declare one
  • context windows from 12,288 to 131,072 tokens
  • maximum output from 2,048 to 128,000 tokens
  • input prices from $0.10 to $1.04 per million tokens
  • tool calling declared on 7 of them
Figure 2: one model string on OpenRouter, eleven deployments.

Same name, materially different products. If your prompts run long, the 12K endpoint and the 131K endpoint aren't interchangeable, and the default router chooses partly on price. OpenRouter does filter by capability when a request needs it: when we sent eight requests with tool_choice: required and no provider pin, every one landed on a host that declares tool support. But declared isn't the same as working. One host that declares tool support returned no tool call at all when we pinned to it and forced one.

Read /api/v1/models/{model}/endpoints before trusting a model string, and pin or filter providers wherever quantisation, context length or tool behaviour matters.

Key components

The catalog

OpenRouter's API listed 466 models across 65 vendor namespaces on 5 October, up from 428 when we first pulled it on 8 September. Embedding models sit on a separate endpoint, /api/v1/embeddings/models, which listed 33 more. Add those and you're close to the "500+ models" on OpenRouter's pricing page. Provider counts line up less neatly: the providers endpoint returned 112 entries, one of them called FakeProvider, against "80+" in the marketing. Query the API rather than trusting a count in a blog post, this one included.

Routing

The default route spreads traffic. It does not optimise for price or speed. In our September test on Llama 3.3 70B, with eight requests per strategy, the default was the slowest option at a median 935 ms and cost twice as much as sort: price. sort: throughput was fastest at 289 ms, ahead of sort: latency at 347 ms, and the cheapest and most expensive strategies were 10x apart for the same prompt. Eight requests is enough to show the directives work and differ. It is too small a sample to rank hosts. If cost or speed matters to you, set a sort.

Two details about fallbacks. A misspelt provider slug in order is dropped silently rather than rejected, so a typo just narrows your list. And as noted above, attempt is the only signal a fallback fired.

API compatibility

The OpenAI SDK works as a drop-in, and for most chat traffic that's all there is to it. The edge cases are parameters a particular route doesn't support, which tend to be accepted with a 200 and not applied. We sent OpenAI's prediction parameter to GPT-4o mini, which declares support for it, and got a normal response with no sign it took effect. If a parameter changes your output or your bill, check the response for evidence that it worked.

Billing

You pay the provider's list price per token. OpenRouter does not add a markup on top. Its revenue is a fee when you buy credits: 5.5% on the Standard plan with a $0.80 minimum, 8% on Business, and 5% for crypto. With your own provider key, the first $25,000 a month of list-price usage carries no fee, and 5% after that. Credits can expire a year after purchase. The pricing has changed twice in fourteen months, so check OpenRouter's pricing page before modelling costs. Our OpenRouter pricing breakdown goes further.

Data controls

Prompts and responses always pass through OpenRouter. You cannot run the service yourself or keep it air-gapped. The controls available on the hosted service are zero data retention routing, a data_collection: "deny" filter that only routes to providers that don't collect your data, and in-region routing. In-region routing used to need an Enterprise contract. Since September it's on the self-serve Business plan: send requests to eu.openrouter.ai or us.openrouter.ai and they're processed only in that region, failing with an error if no in-region endpoint exists. The documented trade-offs: OpenRouter's own prompt and response logging is skipped on regional endpoints, and Batch, web fetch, sandboxed shell, file and image tools, and the Auto Router aren't available there.

Rate limits and caching

Both are covered in their own posts. OpenRouter rate limits covers why your free-model ceiling depends on what you've paid, and OpenRouter prompt caching covers why the same caching feature has nine different prices.

How much latency does OpenRouter add?

OpenRouter's docs say it runs at the edge to keep overhead minimal. Its homepage used to put a number on that, around 25 ms, and plenty of guides still repeat it. As of 5 October the homepage just says "minimal latency". The docs add one detail: when your credit balance runs low, OpenRouter expires its edge caches more aggressively to keep billing accurate, which adds latency until you top up.

We measured it. The same 100 prompts went three ways: direct to the provider, through OpenRouter on credits, and through OpenRouter with our own provider key. Requests were interleaved and paired per prompt, so prompt content cancels out of the difference, and each one was pinned to the model's first-party provider and checked per request.

Figure 3: time to first token added by OpenRouter, paired against calling the provider directly.

On GPT-4o mini we couldn't measure any overhead. The median added time to first token was 19.7 ms on credits, but the middle half of results ran from 85 ms faster to 157 ms slower, so the true figure is below what our setup can resolve. On Claude Haiku 4.5 it was unambiguous: 206 ms added on credits and 124 ms with our own key, with the interquartile range well clear of zero. Without streaming, across two runs, credits added 837 and 883 ms to the full response while our own key stayed within 42 ms of going direct.

Credits and your own key pass through the same gateway and the same translation, so what differs is the upstream account. Our best explanation, and it is an inference, is contention on OpenRouter's shared Anthropic capacity. On a busy provider, bringing your own key buys latency as well as billing control.

All of this ran from one laptop on consumer wifi through OpenRouter's Los Angeles edge, in September 2026. Read the differences as real and the absolute numbers as specific to our setup.

Who runs OpenRouter

OpenRouter was founded in 2023 by Alex Atallah, previously a co-founder of OpenSea. On 19 August 2026, Stripe announced an agreement to acquire it. As of early October the deal hadn't been reported as closed, and neither company has disclosed a price; press reports put it above $7 billion. OpenRouter has said its product and commitments won't change.

OpenRouter's revenue is a fee on payments, its pricing will be set by a payments company, and its neutrality between providers is a stated commitment rather than something built into its structure. A team about to build on it for two years should have that in the evaluation, next to the routing and latency numbers, and then decide.

OpenRouter vs an LLM gateway you run

OpenRouter gets compared with LLM gateways like LiteLLM or ours. Both put one API in front of many models. They are different products. Our explainer on what an LLM gateway is covers the general idea.

OpenRouterSelf-hosted LLM gateway
Where it runsOpenRouter's infrastructureYour VPC or data centre
Provider contractsOpenRouter's, or your own keysYours
BillingCredits at list price plus a purchase feeDirect from each provider
Where prompts goThrough OpenRouter, then to the providerFrom your network straight to the provider
Failure domainAn OpenRouter outage takes every provider offline at onceYour gateway; provider outages stay isolated
Workload identityKeys belong to the person who created themDepends on the gateway
Time to first requestMinutesHours to days
Cost of leavingHours to move traffic, weeks to replace the provider contractsContracts are already yours; config needs migrating

Two rows need a sentence each. OpenRouter's control plane and data plane are the same system, so when it's down, everything behind it is too. Its own engineering blog describes a roughly 50-minute database outage in August 2025 that took routing down, and argues that aggregate uptime still beats any single provider; in our own testing across roughly 4,300 requests we didn't see a single 4xx or 5xx. And on identity: keys are owned by the member who created them, and deactivating that person through SCIM deactivates their keys, which is the right default for people and an awkward one for production workloads.

When OpenRouter is a good fit

OpenRouter fits prototyping across many models, when you want one key and one bill and would rather skip a contract with every provider you try. It also fits open-weight models, where choosing the host by price or throughput saves real money, as long as your prompts can leave your network, with or without a regional constraint.

It is a poor fit when prompts have to stay on your network, when workloads need machine identities rather than keys owned by people, or when one vendor's outage taking every model offline is unacceptable. The same applies if latency matters on a contended provider and you can't bring your own key, or if procurement will want direct contracts with each provider anyway.

How TrueFoundry approaches this

The TrueFoundry AI Gateway runs in your VPC, on-prem or air-gapped, and calls providers directly with your own contracts and keys. It exposes 1,000+ LLMs through one OpenAI-compatible API, adds roughly 3 to 4 ms of overhead, and handles 350+ RPS on a single vCPU.

You can keep OpenRouter in that setup. It is supported as a provider in TrueFoundry, so the catalog stays available for experiments while your own rate limits, budgets, access control and logging sit in front of it. For teams already on OpenRouter, that is usually the lowest-risk first step.

Related reading

Conclusion

OpenRouter is one API in front of a marketplace. It picks who serves each request, holds your credits, and pays the providers. Before you treat a model string as a stable product, read the endpoints list, set a routing sort, and measure the overhead on the providers you actually use.

If you want that single API inside your own infrastructure, see how the TrueFoundry AI Gateway works.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Inscríbase
Tabla de contenido

Controle, implemente y rastree la IA en su propia infraestructura

Reserva 30 minutos con nuestro Experto en IA

Reserve una demostración

La forma más rápida de crear, gobernar y escalar su IA

Demo del libro
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Descubra más

No se ha encontrado ningún artículo.
October 7, 2026
|
5 minutos de lectura

What is OpenRouter? A technical guide to what happens to your request

No se ha encontrado ningún artículo.
October 7, 2026
|
5 minutos de lectura

OpenRouter rate limits: your free tier depends on what you've already paid

No se ha encontrado ningún artículo.
Analyzing the differences between Mint MCP and Vercel AI governance
October 7, 2026
|
5 minutos de lectura

Mint MCP frente a Vercel AI Gateway: ¿Qué plataforma se adapta mejor a los equipos de IA empresariales?

No se ha encontrado ningún artículo.
Comparing Solo.io and Mint MCP gateways
October 7, 2026
|
5 minutos de lectura

Mint MCP frente a Solo.io: ¿Qué pasarela MCP se adapta mejor a los equipos de IA empresarial?

No se ha encontrado ningún artículo.
No se ha encontrado ningún artículo.

Blogs recientes

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Preguntas frecuentes

What is OpenRouter?

OpenRouter is a hosted, OpenAI-compatible API that gives you access to hundreds of AI models from many providers through one endpoint and one API key. It picks which provider serves each request, bills you at the provider's list price, and charges a fee when you buy credits. You cannot run it yourself, so prompts always pass through OpenRouter's infrastructure.

Is OpenRouter free?

Partly. OpenRouter offers free variants of some models, capped at 20 requests per minute and 50 requests per day, rising to 1,000 a day once you've bought at least 10 credits. Paid models are billed at the provider's list price, plus a fee on each credit purchase. Our post on OpenRouter rate limits has the detail.

Does OpenRouter add latency?

It depends on the provider. In our paired testing, the added time to first token on GPT-4o mini was too small to measure, while Claude Haiku 4.5 on OpenRouter credits added about 206 ms, or about 124 ms with our own Anthropic key. OpenRouter itself no longer publishes an overhead figure.

¿Puedo desplegar TrueFoundry en mi propia VPC o en on-prem?

Sí. TrueFoundry se ejecuta en su VPC, on-prem, en entornos air-gapped o híbridos, de modo que los prompts y las respuestas nunca salen de su dominio, incluso cuando enruta entre muchos proveedores.

Realice un recorrido rápido por el producto
Comience el recorrido por el producto
Visita guiada por el producto