Budgets and Quotas on the AI Gateway: Control AI Spend by Team
.png)
Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
Why dollar budgets, not just rate limits
A GPU endpoint or a provider account bills whether or not anyone is watching. Rate limits cap throughput, but they do not stop a team from quietly spending ten thousand dollars on tokens in a month. Budgets cap spend in dollars, which is what finance actually cares about.
On the AI Gateway, budget limiting helps you control spending on LLM workloads by setting cost boundaries per tenant or per team, with alerts and enforcement built in.
Tenant budgets vs team budgets
The budget limiting docs describe two scopes, and the distinction decides who owns the cap.
- Tenant budget. Managed by tenant admins only, it applies to all requests across the tenant that match the rule filters. Use it for organization wide caps, such as a monthly spend limit on a specific model.
- Team budget. Managed by tenant admins and team managers, it applies only to requests from members of a single team. Use it when a team lead should manage their own team spend without tenant admin access.

Creating a budget rule, step by step
- Open the rule editor. Go to AI Gateway, then Policies, then Budget Limiting, and click Add Rule.
- Select the scope, either Tenant Budget or Team Budget.
- Add scope filters on subjects, models, provider accounts, or metadata. If you leave all filters empty, the rule matches every request within its scope.
- Set budget limits with one or more periods: daily, weekly, monthly, quarterly, or lifetime. A single rule can enforce multiple periods at once, and the request is blocked if any period is exceeded.
- Choose how the limit applies, whether aggregate, per user, per model, per virtual account, or per metadata. This cannot be changed after the rule is created.
- Pick an enforcement mode: enforce to block when exceeded, audit to warn only, or soft enforce to block only when no other matching budget can allow the request.

One important detail from the docs: when you create a budget rule, the usage counter starts at zero from that moment, regardless of how much was spent earlier in the period. Prior spending is not retroactively counted.
Alerts before you hit the cap
Budgets are most useful before they are breached. When creating a rule you can turn on budget milestone alerts and select thresholds at 75, 90, 95, or 100 percent, then choose a notification target such as email, Slack, PagerDuty, or Microsoft Teams. This gives owners time to react instead of discovering the cap only when requests start failing.
Seeing budgets in analytics
The analytics dashboard surfaces budget outcomes under routing metrics. Top level counters show model calls blocked by rate limit and model calls blocked by budget limit, and the budget charts show the checks rate, the exceeded rate, and an allowed versus blocked breakdown per budget rule.
You can view these by configs, users, virtual accounts, or teams. The teams view is what supports chargebacks and lets each team manage its own budget.

Budgets vs rate limiting
| | Budgets | Rate limiting |
| --- | --- | --- |
| Caps | Cost in dollars | Throughput (requests, tokens, tool calls) |
| Scopes | Tenant and team | Tenant only |
| Threshold alerts | Yes (75, 90, 95, 100%) | No |
| Periods | Daily to lifetime | Per minute, hour, day |
The two are complementary. Use rate limiting to protect capacity and budgets to protect the wallet. Both return an HTTP 429 when breached, and
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
What is the difference between a budget and a rate limit?
A budget caps spend in dollars, while a rate limit caps throughput such as requests or tokens. Budgets support tenant and team scopes and send threshold alerts, whereas rate limits are tenant scoped and do not send threshold alerts.
Can a team manage its own budget?
Yes. A team budget can be managed by a team manager as well as tenant admins, and it applies only to that team members requests. This lets a team lead control their own spend without tenant admin access.
Does a new budget count spend from earlier in the month?
No. The usage counter starts at zero from the moment you create the rule. Prior spending in the current day, week, month, or quarter is not retroactively counted.
What happens when a budget is exceeded?
In enforce mode the request is blocked with an HTTP 429 and a budget limit error. In audit mode it is tracked and alerted without blocking, and in soft enforce mode it is blocked only when no other matching budget can allow the request.










.png)
.png)

.png)
.png)





.png)



.png)





