Skip to main content
Rate limiting controls how much traffic flows through the AI Gateway. You define named rate limits that match specific callers, models, provider accounts, MCP servers, or metadata, and the AI Gateway blocks requests once a limit is breached, or runs in a warn-only mode so you can size a limit before enforcing it.
Rate Limiting V2 runs alongside Rate Limiting V1. V1 is still supported and needs no migration. New rate limits should use V2, which adds multi-window limits, NOT IN exclusions, per-entity overrides, audit mode, and MCP tool-call limits. See Relationship to rate limiting V1 for how the two interact.

Rate limiting families

V2 splits rate limiting into two families. Each is a separate named configuration with its own limit units and its own filters.
Rate limits control throughput, not spend. To cap cost in dollars, use Budget Limiting instead. A model rate limit and an MCP rate limit are allowed to share the same name, because each family counts usage separately.

How rate limiting works?

Rate limiting consists of a set of independent named limits. Each one defines which traffic it applies to, and how much of it is allowed. During evaluation:
  1. Every matching limit is checked. The AI Gateway finds all rate limits whose filters match the incoming request.
  2. All matching limits must allow the request. If any matching limit is breached and is in enforcement mode, the request is blocked.
  3. Usage is counted against every matching limit. When a request matches several rate limits, usage increments on each of them.
Think of matching rate limits as an AND across limits: a request proceeds only when every limit it touches has room. This differs from V1, where only the first matching rule was enforced.
When several limits match, they are evaluated in alphabetical order by name. This affects only which limit is reported first on a blocked request. There is no priority field, and ordering never changes whether a request is blocked.
Enforcement is eventually consistent. The AI Gateway does not hold usage counters itself: it reports usage and reads back the blocked state. A burst of concurrent in-flight requests can therefore overshoot a limit slightly before blocking begins. Raising a limit clears an active block immediately, without waiting for the next reconciliation.
Rate Limiting V2 rules list showing model and MCP rate limits with usage and mode The rules list shows each rule’s partition under Applies To, its filters under Scope, current usage, the configured limits, and its enforcement mode. The family, Model or MCP, appears under the rule id.

Setting up a rate limit

To create a rate limit, go to AI Gateway → Policies → Rate Limiting v2, and click on + Add Rule. The form has two steps: Select type and Configure. In Select type, choose what you are rate limiting and give the rule a name: Rate limit type selection between Model and MCP with a rule name field Rate limits are tenant-scoped and are managed by tenant admins. Unlike Budget Limiting, there is no team-owned scope, and rate limits do not send threshold alerts.
Click Apply as YAML at any point in the Configure step to view or edit the rule as YAML instead of filling the form. This is also the only way to set the filters that the form does not expose, described below.

Scope filters

Scope defines which traffic the rate limit applies to. Use + Add Filters to add filters, each with an IN or NOT IN operator. Filters use AND logic, so a request must match every filter you add. The form offers three filters for model rate limits:
Two further filters exist in the configuration but are not in the form: provider_accounts matches a provider account name such as openai-main, and subjects.agents matches agent callers. Set either through Apply as YAML. See the YAML configuration reference.
Subject filters cover users, teams, virtual accounts, and agents. An IN match on any one of them satisfies the scope. A NOT IN match always wins: a caller excluded by one subject filter is not rescued by matching another. * works as a wildcard in both operators. in: ["*"] matches any caller of that kind, and not_in: ["*"] excludes all of them.
If you add no filters, the rate limit matches every request in the tenant. This is useful for a tenant-wide default limit.
Requests authenticated by a service account cannot be selected by a subject filter. A rate limit with no subject filters still applies to them.

Rate limits

Under Add Rate Limit, set a Limit amount and pick its period. Use + Add Period to add more. Every limit you set is enforced at the same time, so a request is blocked if any period on that rule is breached. Model rate limits support these periods: Rate limit period options for requests and tokens per minute, hour, and day All values must be positive integers. Token limits count input and output tokens together, so a single expensive request can consume a large share of the window.
Failed requests still consume request quota, because the request reached the provider. Requests the AI Gateway itself rejected with 429 are not counted, so a blocked caller does not dig their own hole deeper.

Apply limit as

Apply limit as controls how the limit is partitioned across matching traffic. For per metadata, you also pick the metadata key whose distinct values each get their own limit. A limit partitioned on project_id, for example, gives every project its own allowance. Configure step with per-user partitioning, two limit periods, and an overrides section
Apply limit as cannot be changed after the rate limit is created. The API rejects the change with a 403. To repartition a limit, create a new one and delete the old one.
Two combinations are rejected, because the partition would have no value to key on:
  • per user cannot be combined with the virtual_accounts or agents subject filters.
  • per virtual-account cannot be combined with the users, teams, or agents subject filters.
The form enforces this as you go. Selecting per user narrows the Subjects picker to users and teams only.
There is no per-agent partition. Agents can be filtered on through subjects.agents in YAML, but agent traffic caught by a per-user limit produces no counter. Use aggregate, per model, or per metadata to bucket it.

Overrides

On any per-entity partition, an Overrides (optional) section appears. Use + Add Override to add per-entity exceptions that replace specific limits for named entities. Every other matching entity keeps the base limits. Overrides merge with the base limits one period at a time. If a limit sets 100 requests per minute and 10,000 requests per day, an override of 500 requests per minute for one user leaves that user’s daily limit at 10,000. Two rules apply:
  • An override can only replace a period that the base limits already set. You cannot introduce a new period through an override.
  • Override target lists cannot overlap. No entity may appear in two overrides on the same rate limit.

Enforcing strategy

Enforcing Strategy controls what happens when a rate limit is breached. Enforcing Strategy options: Enforce, Soft Enforce, and Audit In Audit mode, usage is still tracked and the limit is reported on the response, so you can see what would have been blocked.
Soft Enforce is a condition across rate limits, not a gentler version of Enforce. Whether it blocks depends entirely on what else matched the request. A breached Enforce limit always blocks, no matter what other limits allow.
Use Audit to size a new limit against real traffic before enforcing it. Use Soft Enforce for an overflow cap that should bite only when nothing else is already holding the traffic back.

Log request body on block

Log Request Body on Block keeps the request body on the trace for requests this rate limit blocks. It is off by default, and blocked requests normally have their body stripped from the trace. Turn it on when you need to see what a blocked caller was actually sending. The body remains subject to your tenant logging and redaction settings, so a redacted field stays redacted.

Virtual models in when.models

You can put a virtual model id in the Models filter the same way as a concrete model. When a request uses that virtual model:
  • The AI Gateway matches the rate limit against the virtual model id and the concrete target it routes to. Either id can satisfy an IN list, and either id can trigger a NOT IN exclusion.
  • Per Model counters always key on the concrete model that served the request, never on the virtual model id.
Do not list a virtual model id in Models IN on a limit that uses Per Model. That configuration is rejected and the rate limit is ignored entirely. Use Aggregate, Per User, Per Virtual Account, or Per Metadata Value instead.

MCP rate limits

MCP rate limits cap tool calls through the MCP Gateway. They share the concepts above, including subject filters, metadata filters, overrides, and the three enforcement modes. This section covers only what differs.

MCP scope filters

MCP rate limits filter on Subjects, MCP Servers, and Metadata. They have no Models or provider account filter. MCP Servers accepts two forms of value: MCP limits support tool calls / minute, tool calls / hour, and tool calls / day, on the same rolling windows as the model periods. There is no token or request period, because a tool call is the countable event. Apply limit as offers per mcp in place of per model. It partitions the limit by server. Listing a server:tool value in an override gives that single tool its own counter, while every other tool on that server continues to share the server counter. This is how you carve out one expensive tool without splitting the whole server. MCP rate limit configure step with MCP Servers filter, tool call limits, and per-mcp partitioning Three behaviors are worth knowing:
  • Only tools/call is rate limited. tools/list, resources/*, and prompts/* are never counted.
  • A tool call that the upstream server refused with 403 or 429 is not counted, because the tool never ran. Any other failure is counted.
  • For a virtual MCP server, the call is counted once, against the source server it routes to. A rate limit may name either the virtual server or the source server, and both will match.

MCP block response

A blocked MCP tool call does not return HTTP 429. It returns HTTP 200 with a JSON-RPC result marked as a tool error, so the client hands the refusal to the model rather than failing the transport.
The text names the rate limit and the counter that was breached. No Retry-After header is returned.

Viewing rate limit usage

The rules list shows a Usage summary per rule: No usage when idle, a shared-counter bar for aggregate rules, or a count of entities in active use and how many are within their limit for per-entity rules. A rule that is currently over its limit is flagged there. Click View rate limit data on a rule for its Rate Limit Summary, which reports each configured period’s limit alongside the current cycle start and end, and a used-against-limit bar per period: Rate Limit Summary showing each period's limit, cycle start and end, and usage bars Entities are classified by their worst period: below 90% of the limit is within limit, 90% or more is at risk, and 100% or more is over.
Per-entity breakdowns track up to 500 entities per rate limit by default. Beyond that the breakdown is marked as incomplete, and it should be read as the busiest entities rather than an exact list of every one. Looking up a specific entity’s usage stays exact even when it is absent from the breakdown.

Rate limit exceeded response

When a model rate limit in enforcement mode is breached, the AI Gateway returns HTTP 429:
The response also includes an x-tfy-applied-rules header naming the rate limit that was breached:
A limit in audit mode reports itself on this header with "audit_mode":true and "violated":false, and the request still goes through. This is how you confirm an audit limit is matching the traffic you expect before switching it to enforce.
The AI Gateway does not return Retry-After, X-RateLimit-Limit, X-RateLimit-Remaining, or X-RateLimit-Reset headers. Clients should back off on 429 using their own retry policy. To read current usage programmatically, query the rate limit usage APIs rather than parsing response headers.
A budget breach also returns 429, distinguished by "type": "BudgetLimitError" and "error_origin_level": "budget_limit". See Budget Limiting.

Relationship to rate limiting V1

V1 and V2 are independent systems that both evaluate on every request. V1 is checked first, and if a V1 rule blocks the request, V2 is not consulted. If V1 allows the request, V2 evaluates its own matching limits. This has three practical consequences:
  • No migration is required. V1 configurations keep working unchanged, and V1 is not deprecated.
  • Counters are not shared. A V2 limit starts counting from zero when you create it, and it never inherits usage from a V1 rule, even one with the same intent.
  • Traffic can be covered twice. If a V1 rule and a V2 limit both match the same request, both count it, and either can block it. When you recreate a V1 rule in V2, delete or narrow the V1 rule to avoid enforcing the same cap twice.

Practical examples

Each example shows both the UI configuration and the equivalent YAML. Click Apply as YAML in the rule form to paste or edit YAML directly.
Give every user their own throughput allowance across both requests and tokens.
How it works: Each user gets an independent 100 requests per minute and 200,000 tokens per hour. A user who exhausts either window is blocked until it rolls forward, and other users are unaffected.
Cap how hard one team can hit an expensive model, as a single shared pool.
How it works: All requests from members of the research team to that model share one 600 requests per hour pool. Requests from the same team to other models do not count against it.
Apply a baseline limit to every user except a named service owner.
How it works: Every user except batch-owner@example.com gets 60 requests per minute. The excluded user does not match this limit at all, so they are governed only by whatever other rate limits match their traffic.
Protect a self-hosted or quota-limited provider account from being saturated.
How it works: Every request routed to any model on the selfhosted-vllm account counts against one shared pool, which keeps total load inside what the deployment can serve.
Give every project its own allowance, identified by a metadata header.
Requests must include the header:
How it works: Each distinct project_id gets its own 5,000 requests per day. Only production traffic matches, so non-production requests are not limited by this rule.
Set a default per-model limit, then raise it for a cheap, high-volume model.
Overrides: openai-main/gpt-4o-mini → 2,000 requests/minute. Its daily limit stays at the base 50,000, because an override replaces only the units it sets.
Watch what a limit would block before you enforce it.
How it works: No request is ever blocked. Usage is tracked per user, and requests that would have been blocked carry "audit_mode":true on the x-tfy-applied-rules header. Once the usage breakdown shows an acceptable number of users at risk, switch Mode to Enforce.
Add a tenant-wide backstop that bites only when no narrower limit is already holding the traffic.
How it works: A request is checked against both. The per-user limit blocks absolutely when breached. The backstop blocks only when it is breached and no other matching limit has room, so it acts as a ceiling on total tenant throughput without pre-empting the per-user limit’s own error.
Limit each MCP server’s tool calls, and give one expensive tool a tighter limit of its own.
Overrides: github:create_issue → 5 tool calls/minute.How it works: Every server gets 120 tool calls per minute. Because github:create_issue is named in an override, that one tool gets its own counter capped at 5 per minute, while all other github tools continue to share the server’s 120.
Keep a CI virtual account from exhausting a shared MCP server.
How it works: Only ci-bot calls to tools on the search server match. Both windows are enforced, so the account is capped on bursts and on sustained volume. Human users calling the same server are unaffected.

YAML configuration reference

This example uses a metadata partition so that it can show every subject filter at once. A per-user partition cannot be combined with the virtual_accounts or agents filters, and a per-virtual-account partition cannot be combined with users, teams, or agents. Those combinations are rejected with a 400.
Field reference:MCP rate limits use the same structure, with mcp_servers in place of models and provider_accounts, tool_calls_* units, and per-mcp in place of per-model: