OpenRouter latency on Anthropic: why bringing your own key is faster

Conçu pour la vitesse : latence d'environ 10 ms, même en cas de charge
Une méthode incroyablement rapide pour créer, suivre et déployer vos modèles !
- Gère plus de 350 RPS sur un seul processeur virtuel, aucun réglage n'est nécessaire
- Prêt pour la production avec un support complet pour les entreprises
OpenRouter's homepage used to say it adds about 25 ms between your users and their inference, and plenty of guides still repeat the figure. The homepage itself now just says "minimal latency". On Claude we measured several times that before the first token arrived, and the larger cost came after it.
The measurements are for Anthropic: how much OpenRouter adds, where the time goes, and why our own Anthropic key removed most of it. One run used short prompts. The other used prompts up to 20,000 tokens. The limits of both runs are in the section on what we didn't test, and they bound the numbers as much as the medians do.
How we measured
Every request went three ways: directly to Anthropic's API, through OpenRouter billed to OpenRouter credits, and through OpenRouter using our own Anthropic key (BYOK). The same prompts ran on each, interleaved and with the order alternated per prompt, so network drift hit all three equally. Every comparison is paired: each prompt is compared against itself on another route, so prompt content cancels out of the difference.
OpenRouter requests were pinned to Anthropic's own endpoint rather than Bedrock or Vertex, with fallbacks off, and we checked on every response which endpoint had served it and whether it had used our key. Temperature was 0 and connections were reused.
There were two runs. On 7 September, two separate sessions of 100 short prompts each, 4 to 37 tokens long, drawn from public datasets. On 14 September, 99 prompts padded to three sizes: 200 input tokens with 100 output, 2,000 with 500, and 20,000 with 1,000. Everything ran from one laptop on consumer wifi through OpenRouter's Los Angeles edge. That inflates absolute latency, so we only report differences between routes, which the pairing makes valid.
How much latency OpenRouter adds to the first token

The first-token cost was steady. On credits it was 192 ms in the first session and 206 ms in the second, and 194 ms at 200 tokens a week later. With our own key it was 119, 124 and 102 ms. Two runs a week apart, on different prompt sets, landing within about 20 ms of each other is about as much replication as a client-side test can give you. Across four sessions where credits and our own key ran side by side, credits was slower to the first token every time, by a median of 50 to 115 ms.
It also grows with prompt size. From 200 to 20,000 input tokens, the added time to first token rose from 194 to 303 ms on credits and from 102 to 186 ms with our own key. The growth is sublinear, roughly 1.6 to 1.8 times for a hundred times the input, and it only became visible at 20,000 tokens; an earlier sweep that stopped at 5,000 found nothing. Within the short-prompt runs, where inputs only ranged from 4 to 37 tokens, there was no relationship at all.
On GPT-4o mini, the same measurement stayed inside the noise in every condition, including 20,000-token prompts. The median added time to first token was 19.7 ms on credits, but the middle half of results ran from 85 ms faster to 157 ms slower than going direct, so the real figure is smaller than our setup can resolve.
Two costs, not one
Time to first token is only part of a streamed response. Splitting each response into the wait for the first token and the time to stream the rest shows two separate costs.

With our own key, OpenRouter added about 120 ms before the first token and nothing after it. Our key streamed at 115 and 116 tokens per second in the two sessions, against 115 and 113 going direct. Whatever this cost is, it's paid once, up front.
On credits, the same answers streamed more slowly, at 94 and 95 tokens per second, about 17% slower. The responses weren't longer: median output was around 290 tokens on every route. On a 290-token answer that slower stream adds roughly 520 ms, on top of the first-token cost. The arithmetic closes: 290 tokens at 94 rather than 114 tokens per second is about 540 ms, plus about 200 ms to the first token, against the 711 and 730 ms we measured for the full response.
That generation penalty was real in both runs but not the same size. On 14 September the extra time after the first token was roughly 130, 215 and 280 ms for 100, 500 and 1,000-token answers, much smaller than on 7 September. Those figures come from medians rather than paired differences, so treat them as approximate, but the direction is clear: the first-token cost held steady across days and prompt sizes, and the generation penalty varied.
Why: the gateway and the account
The two costs point at two different causes.
The first-token cost appears with our own key as well as on credits, so it belongs to OpenRouter's path to Anthropic rather than to whose account pays. The most likely candidates are the work of translating an OpenAI-format request into Anthropic's Messages format and translating the stream back, and the network path from OpenRouter's edge to Anthropic. Growth with prompt size fits the first, since translation work scales with the request. Our data can't separate the two, and the absence of any measurable cost on GPT-4o mini, which needs no format translation, is consistent with either. Sending the same requests through OpenRouter's native Anthropic endpoint, /api/v1/messages, is the experiment that would settle it, and we haven't run it yet.
The generation penalty is different. Credits and our own key go through the same gateway, the same translation and the same Anthropic endpoint. What differs is the Anthropic account behind the request: OpenRouter's, shared across its customers, or ours. Our best explanation is that requests on OpenRouter's account were served with less capacity than requests on ours, and that the amount varies with demand, which would also explain why the penalty shrank between our two runs. That is an inference. Neither OpenRouter nor Anthropic publishes how accounts are tiered.
On OpenAI the pattern reversed. On GPT-4o mini, credits reached the first token 23 to 50 ms sooner than our own key in all three sessions where both ran, and in the session where we compared streaming speed, both generated at the same rate. The account serving your request does affect speed, and which account is faster depends on the provider. You can't see the tier you're on from the outside. You can only measure it.
Tails and full responses
The middle of the distribution is the typical request. Tails matter for user-facing work. In the two short-prompt sessions, the 95th-percentile time to first token was about 700 ms going direct, 950 to 985 ms on credits, and 885 to 905 ms with our own key. The 99th percentile widened much more on both OpenRouter routes, from about 750 to 765 ms direct to 1.2 to 2.6 seconds on credits, but at 100 requests per session the 99th percentile is effectively the second-slowest request, so read it as a sign of a heavier tail rather than a figure.
Without streaming, the picture is the same. Paired against direct, credits added 548 and 629 ms to the full response in two sessions, and our own key added 42 and 56 ms.
What we didn't test
These results cover one model, Claude Haiku 4.5, on Anthropic's own endpoint. We didn't test Sonnet or Opus, Claude on Bedrock or Vertex through OpenRouter, other OpenRouter edges, or different times of day. The short-prompt runs are reproducible from our published harness. The padded-prompt run was a separate session. We haven't run a concurrency sweep on Anthropic credits, which is the most direct test of the shared-capacity explanation, or the /api/v1/messages comparison described above. And everything ran from a single client, which is fine for differences between routes but says nothing about absolute latency from your infrastructure.
What to do about it
- If latency on Claude matters, bring your own Anthropic key. In our tests it removed all of the generation penalty and about 40% of the first-token cost. It's also usually cheaper.
- Check which endpoint actually served you. Send
X-OpenRouter-Experimental-Metadata: enabledand readopenrouter_metadatain the response, which tells you the serving endpoint and whether your own key was used. Without it you're trusting the pin. - Pin the provider if you need predictable latency. Claude served from Anthropic, Bedrock and Vertex are different paths.
- Measure on your own traffic, with pairing. Send each prompt down both routes, alternate the order, subtract per prompt, and report the median difference with its interquartile range. If the range crosses zero, you haven't measured an overhead, you've measured noise. Our overview of how OpenRouter routes requests covers the rest of the request path.
Rate limits are the other common reason Claude feels slow through OpenRouter, and our post on OpenRouter rate limits explains how to tell a slow response from a throttled one.
How TrueFoundry approaches this
The TrueFoundry AI Gateway runs in your own VPC or data centre and calls Anthropic directly with your own key and contract, so requests never share an upstream account with other customers and there's no third-party hop between your network and the provider. Our published figure is roughly 3 to 4 ms of gateway overhead, handling 350+ RPS on a single vCPU.
That 3 to 4 ms figure is TrueFoundry's published benchmark. We did not measure it with the harness behind this post. If you want to compare gateways, the paired method above is how we'd do it, and it works against any OpenAI-compatible endpoint, ours included.
Related reading
- OpenRouter prompt caching, including OpenRouter's response cache, which answered repeat requests in about 59 ms
- OpenRouter alternatives, comparing the options for production teams
- Best AI gateways for LLM inference optimization, on how gateways handle routing and latency
- What is an LLM gateway?, the general architecture
Conclusion
OpenRouter latency on Anthropic comes in two parts. A first-token cost of roughly 100 to 300 ms that grows with prompt size and appears whichever account you use, and, on OpenRouter credits, slower generation that added more than half a second to a typical answer in one run and less in another. Bringing your own key removed the slower generation in our tests, and left the first-token cost in place. The only way to know what that extra wait costs on your prompts is to measure it the way we did, paired.
To call Anthropic on your own key, from your own network, with no shared upstream account, see how the TrueFoundry AI Gateway connects to Anthropic directly.
TrueFoundry AI Gateway offre une latence d'environ 3 à 4 ms, gère plus de 350 RPS sur 1 processeur virtuel, évolue horizontalement facilement et est prête pour la production, tandis que LiteLM souffre d'une latence élevée, peine à dépasser un RPS modéré, ne dispose pas d'une mise à l'échelle intégrée et convient parfaitement aux charges de travail légères ou aux prototypes.



Gouvernez, déployez et suivez l'IA dans votre propre infrastructure
Blogs récents
Questions fréquemment posées
Does OpenRouter add latency?
It depends on the provider. In our paired testing, the added time to first token on GPT-4o mini was too small to measure, while Claude Haiku 4.5 on OpenRouter credits added about 206 ms, or about 124 ms with our own Anthropic key. OpenRouter itself no longer publishes an overhead figure.
S'intègre-t-elle à ma pile d'observabilité existante ?
Oui. La passerelle est compatible OpenTelemetry et s'intègre à Grafana, Datadog, Prometheus, ou à votre pile technologique préférée. Elle trace chaque requête, du prompt à l'exécution de l'outil et du modèle, vous offrant ainsi une journalisation unifiée sans avoir à remplacer ce que vous utilisez déjà.
Does BYOK make OpenRouter faster?
On Anthropic it did in our tests: our own key added about 120 ms to the first token against about 200 ms on credits, and then streamed at direct speed. On OpenAI, credits was slightly faster. Which route is faster depends on the provider and the account behind it.
S'intègre-t-il à ma stack d'observabilité et d'évaluation existante ?
Oui. La passerelle est conforme à OpenTelemetry et exporte les traces vers des backends externes ; c'est ainsi que les données de la passerelle parviennent à des plateformes telles que Braintrust, Langfuse ou Arize. TrueFoundry n'exécute pas lui-même de tâches de scoring hors ligne.









.png)

.png)
.png)
.png)
.png)




.png)



.png)





