API Rate Limiting for LLMs: Count Tokens, Not Requests
.png)
Auf Geschwindigkeit ausgelegt: ~ 10 ms Latenz, auch unter Last
Unglaublich schnelle Methode zum Erstellen, Verfolgen und Bereitstellen Ihrer Modelle!
- Verarbeitet mehr als 350 RPS auf nur 1 vCPU — kein Tuning erforderlich
- Produktionsbereit mit vollem Unternehmenssupport

TrueFoundry AI Gateway bietet eine Latenz von ~3—4 ms, verarbeitet mehr als 350 RPS auf einer vCPU, skaliert problemlos horizontal und ist produktionsbereit, während LiteLM unter einer hohen Latenz leidet, mit moderaten RPS zu kämpfen hat, keine integrierte Skalierung hat und sich am besten für leichte Workloads oder Prototyp-Workloads eignet.



Steuern, implementieren und verfolgen Sie KI in Ihrer eigenen Infrastruktur
Aktuelle Blogs
Häufig gestellte Fragen
What is API rate limiting for LLM traffic, and how is it different?
Same idea — cap what a caller may consume in a window — but the unit changes. Classic API rate limiting counts requests, which works when requests cost roughly the same. LLM requests vary by orders of magnitude in tokens, so a token limit tracks real consumption far better. Run both: tokens govern spend rate, requests protect the connection from retry storms.
Should I use token-based rate limiting or a budget?
Both. A token limit is a throughput control on a rolling window: it stops one caller monopolising capacity right now. A budget is a financial control on a calendar period: it stops a team overspending this month. A token limit cannot express dollars, since the same token count costs different amounts on different models — and a budget says nothing about the current minute.
What happens to my quota when a request fails?
A failed request — rejected by the gateway, or failed by the provider with a 4XX or 5XX — counts as one request against every requests_per_* rule it matched. The one exception is a request already rejected with a 429 by rate limiting itself. Token limits are unaffected: a failed request reports no tokens.
Can I deploy TrueFoundry in my own VPC or on-prem?
Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.
What does the gateway add to request latency?
Roughly 3-4 ms of overhead, handling 350+ RPS on a single vCPU, across 1,000+ supported LLMs. The exception is the optional LLM classifier, which adds a real model call before the request is forwarded.
Does it integrate with my observability stack?
Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, or Prometheus. Each LLM classifier call produces its own span, so classifier latency is visible separately from the served model’s.









.png)
.png)
.png)
.png)
.png)


.webp)
.webp)


.webp)
.webp)
.webp)






