Blank white background with no objects or features visible.

We’re sharing complimentary access to the full Gartner Hype Cycle for AI Governance 2026. Get your copy →

Budgets and Quotas on the AI Gateway: Control AI Spend by Team

By Ashish Dubey

Published: October 9, 2026

Why dollar budgets, not just rate limits

A GPU endpoint or a provider account bills whether or not anyone is watching. Rate limits cap throughput, but they do not stop a team from quietly spending ten thousand dollars on tokens in a month. Budgets cap spend in dollars, which is what finance actually cares about.

On the AI Gateway, budget limiting helps you control spending on LLM workloads by setting cost boundaries per tenant or per team, with alerts and enforcement built in.

Tenant budgets vs team budgets

The budget limiting docs describe two scopes, and the distinction decides who owns the cap.

  • Tenant budget. Managed by tenant admins only, it applies to all requests across the tenant that match the rule filters. Use it for organization wide caps, such as a monthly spend limit on a specific model.
  • Team budget. Managed by tenant admins and team managers, it applies only to requests from members of a single team. Use it when a team lead should manage their own team spend without tenant admin access.
Creating a budget rule under Policies, then Budget Limiting.

Creating a budget rule, step by step

  • Open the rule editor. Go to AI Gateway, then Policies, then Budget Limiting, and click Add Rule.
  • Select the scope, either Tenant Budget or Team Budget.
  • Add scope filters on subjects, models, provider accounts, or metadata. If you leave all filters empty, the rule matches every request within its scope.
  • Set budget limits with one or more periods: daily, weekly, monthly, quarterly, or lifetime. A single rule can enforce multiple periods at once, and the request is blocked if any period is exceeded.
  • Choose how the limit applies, whether aggregate, per user, per model, per virtual account, or per metadata. This cannot be changed after the rule is created.
  • Pick an enforcement mode: enforce to block when exceeded, audit to warn only, or soft enforce to block only when no other matching budget can allow the request.
Setting periods and the apply-as option on a budget rule.

One important detail from the docs: when you create a budget rule, the usage counter starts at zero from that moment, regardless of how much was spent earlier in the period. Prior spending is not retroactively counted.

Worried a runaway agent could blow the monthly budget?
We will set tenant and team caps with milestone alerts on your gateway, live.

Alerts before you hit the cap

Budgets are most useful before they are breached. When creating a rule you can turn on budget milestone alerts and select thresholds at 75, 90, 95, or 100 percent, then choose a notification target such as email, Slack, PagerDuty, or Microsoft Teams. This gives owners time to react instead of discovering the cap only when requests start failing.

Seeing budgets in analytics

The analytics dashboard surfaces budget outcomes under routing metrics. Top level counters show model calls blocked by rate limit and model calls blocked by budget limit, and the budget charts show the checks rate, the exceeded rate, and an allowed versus blocked breakdown per budget rule.

You can view these by configs, users, virtual accounts, or teams. The teams view is what supports chargebacks and lets each team manage its own budget.

Budget and rate-limit outcomes in the routing metrics view.

Budgets vs rate limiting

|  | Budgets | Rate limiting |

| --- | --- | --- |

| Caps | Cost in dollars | Throughput (requests, tokens, tool calls) |

| Scopes | Tenant and team | Tenant only |

| Threshold alerts | Yes (75, 90, 95, 100%) | No |

| Periods | Daily to lifetime | Per minute, hour, day |

The two are complementary. Use rate limiting to protect capacity and budgets to protect the wallet. Both return an HTTP 429 when breached, and

Attribute every dollar and cap spend by team
See tenant and team budgets, milestone alerts, and chargeback views on your own traffic.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Start free
Table of Contents

One Gateway for Every LLM, Agent and MCP Server

Book a 30-min with our AI expert

Book a Demo

The fastest way to build, govern and scale your AI

Book Demo
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Discover More

No items found.
October 9, 2026
|
5 min read

Budgets and Quotas on the AI Gateway: Control AI Spend by Team

No items found.
October 9, 2026
|
5 min read

Gateway Tracing and Request Logs: Debug Every LLM Call

No items found.
October 9, 2026
|
5 min read

How to Configure Guardrails on the AI Gateway

No items found.
October 9, 2026
|
5 min read

OpenRouter BYOK explained: cheaper, often faster, and changing

No items found.
No items found.

Recent Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Frequently asked questions

What is the difference between a budget and a rate limit?

A budget caps spend in dollars, while a rate limit caps throughput such as requests or tokens. Budgets support tenant and team scopes and send threshold alerts, whereas rate limits are tenant scoped and do not send threshold alerts.

Can a team manage its own budget?

Yes. A team budget can be managed by a team manager as well as tenant admins, and it applies only to that team members requests. This lets a team lead control their own spend without tenant admin access.

Does a new budget count spend from earlier in the month?

No. The usage counter starts at zero from the moment you create the rule. Prior spending in the current day, week, month, or quarter is not retroactively counted.

What happens when a budget is exceeded?

In enforce mode the request is blocked with an HTTP 429 and a budget limit error. In audit mode it is tracked and alerted without blocking, and in soft enforce mode it is blocked only when no other matching budget can allow the request.

Take a quick product tour
Start Product Tour
Product Tour