Blank white background with no objects or features visible.

TrueFoundry Named Frost & Sullivan's 2026 Global Transformational Innovation Leader. Read report

Découvrez TrueForge : l'infrastructure d'agents open-source et indépendante des fournisseurs. Réduisez vos coûts de 50%. Explorer maintenant→

API Rate Limiting for LLMs: Count Tokens, Not Requests

Par Ashish Dubey

Published: September 28, 2026

Token-Based Rate Limiting and Budgeting — TrueFoundry

‍

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

INSCRIVEZ-VOUS
Table des matières

Gouvernez, déployez et suivez l'IA dans votre propre infrastructure

Réservez un séjour de 30 minutes avec notre Expert en IA

Réservez une démo

Le moyen le plus rapide de créer, de gérer et de faire évoluer votre IA

Démo du livre
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Découvrez-en plus

Aucun article n'a été trouvé.
September 28, 2026
|
5 min de lecture

Langfuse Alternatives: 7 Options Compared on Licence, Price and Limits

Aucun article n'a été trouvé.
September 28, 2026
|
5 min de lecture

Data Loss Prevention for LLM Traffic: Where It Has to Sit

Aucun article n'a été trouvé.
September 28, 2026
|
5 min de lecture

Data Masking in the AI Gateway: What Actually Works

Aucun article n'a été trouvé.
September 28, 2026
|
5 min de lecture

API Rate Limiting for LLMs: Count Tokens, Not Requests

Aucun article n'a été trouvé.
Aucun article n'a été trouvé.

Blogs récents

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Questions fréquemment posées

What is API rate limiting for LLM traffic, and how is it different?

Same idea — cap what a caller may consume in a window — but the unit changes. Classic API rate limiting counts requests, which works when requests cost roughly the same. LLM requests vary by orders of magnitude in tokens, so a token limit tracks real consumption far better. Run both: tokens govern spend rate, requests protect the connection from retry storms.

Should I use token-based rate limiting or a budget?

Both. A token limit is a throughput control on a rolling window: it stops one caller monopolising capacity right now. A budget is a financial control on a calendar period: it stops a team overspending this month. A token limit cannot express dollars, since the same token count costs different amounts on different models — and a budget says nothing about the current minute.

What happens to my quota when a request fails?

A failed request — rejected by the gateway, or failed by the provider with a 4XX or 5XX — counts as one request against every requests_per_* rule it matched. The one exception is a request already rejected with a 429 by rate limiting itself. Token limits are unaffected: a failed request reports no tokens.

Can I deploy TrueFoundry in my own VPC or on-prem?

Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.

What does the gateway add to request latency?

Roughly 3-4 ms of overhead, handling 350+ RPS on a single vCPU, across 1,000+ supported LLMs. The exception is the optional LLM classifier, which adds a real model call before the request is forwarded.

Does it integrate with my observability stack?

Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, or Prometheus. Each LLM classifier call produces its own span, so classifier latency is visible separately from the served model’s.

Faites un rapide tour d'horizon des produits
Commencer la visite guidée du produit
Visite guidée du produit