Presupuestos y cuotas en la pasarela de IA: controla el gasto en IA por equipo
.png)
Diseñado para la velocidad: ~ 10 ms de latencia, incluso bajo carga
¡Una forma increíblemente rápida de crear, rastrear e implementar sus modelos!
- Gestiona más de 350 RPS en solo 1 vCPU, sin necesidad de ajustes
- Listo para la producción con soporte empresarial completo

TrueFoundry AI Gateway ofrece una latencia de entre 3 y 4 ms, gestiona más de 350 RPS en una vCPU, se escala horizontalmente con facilidad y está listo para la producción, mientras que LitellM presenta una latencia alta, tiene dificultades para superar un RPS moderado, carece de escalado integrado y es ideal para cargas de trabajo ligeras o de prototipos.



Controle, implemente y rastree la IA en su propia infraestructura
Blogs recientes
Preguntas frecuentes
What is the difference between a budget and a rate limit?
A budget caps spend in dollars, while a rate limit caps throughput such as requests or tokens. Budgets support tenant and team scopes and send threshold alerts, whereas rate limits are tenant scoped and do not send threshold alerts.
Can a team manage its own budget?
Yes. A team budget can be managed by a team manager as well as tenant admins, and it applies only to that team members requests. This lets a team lead control their own spend without tenant admin access.
Does a new budget count spend from earlier in the month?
No. The usage counter starts at zero from the moment you create the rule. Prior spending in the current day, week, month, or quarter is not retroactively counted.
What happens when a budget is exceeded?
In enforce mode the request is blocked with an HTTP 429 and a budget limit error. In audit mode it is tracked and alerted without blocking, and in soft enforce mode it is blocked only when no other matching budget can allow the request.









.png)
.png)

.png)
.png)





.png)



.png)





