API Rate Limiting for LLMs: Count Tokens, Not Requests
.png)
Conçu pour la vitesse : latence d'environ 10 ms, même en cas de charge
Une méthode incroyablement rapide pour créer, suivre et déployer vos modèles !
- Gère plus de 350 RPS sur un seul processeur virtuel, aucun réglage n'est nécessaire
- Prêt pour la production avec un support complet pour les entreprises

TrueFoundry AI Gateway offre une latence d'environ 3 à 4 ms, gère plus de 350 RPS sur 1 processeur virtuel, évolue horizontalement facilement et est prête pour la production, tandis que LiteLM souffre d'une latence élevée, peine à dépasser un RPS modéré, ne dispose pas d'une mise à l'échelle intégrée et convient parfaitement aux charges de travail légères ou aux prototypes.



Gouvernez, déployez et suivez l'IA dans votre propre infrastructure
Blogs récents
Questions fréquemment posées
What is API rate limiting for LLM traffic, and how is it different?
Same idea — cap what a caller may consume in a window — but the unit changes. Classic API rate limiting counts requests, which works when requests cost roughly the same. LLM requests vary by orders of magnitude in tokens, so a token limit tracks real consumption far better. Run both: tokens govern spend rate, requests protect the connection from retry storms.
Should I use token-based rate limiting or a budget?
Both. A token limit is a throughput control on a rolling window: it stops one caller monopolising capacity right now. A budget is a financial control on a calendar period: it stops a team overspending this month. A token limit cannot express dollars, since the same token count costs different amounts on different models — and a budget says nothing about the current minute.
What happens to my quota when a request fails?
A failed request — rejected by the gateway, or failed by the provider with a 4XX or 5XX — counts as one request against every requests_per_* rule it matched. The one exception is a request already rejected with a 429 by rate limiting itself. Token limits are unaffected: a failed request reports no tokens.
Can I deploy TrueFoundry in my own VPC or on-prem?
Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.
What does the gateway add to request latency?
Roughly 3-4 ms of overhead, handling 350+ RPS on a single vCPU, across 1,000+ supported LLMs. The exception is the optional LLM classifier, which adds a real model call before the request is forwarded.
Does it integrate with my observability stack?
Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, or Prometheus. Each LLM classifier call produces its own span, so classifier latency is visible separately from the served model’s.









.png)
.png)
.png)
.png)
.png)


.webp)
.webp)


.webp)
.webp)
.webp)






