Skip to main content
A streaming provider can answer 200 OK and only then report a failure inside the stream itself — an Anthropic overloaded_error in the first event, or an OpenAI response.failed event after response.created. Once the AI Gateway has written a 200 and the first bytes to your client, that request can no longer be retried or routed elsewhere: your application receives a successful-looking stream that carries no usable output. The first chunk gate is the mechanism that prevents this. For streaming requests, the AI Gateway reads the first chunk from the provider before it starts the response to your client. Only when that chunk looks like real generation does the gateway open the stream. If the chunk carries a failure, nothing has been sent yet, so the gateway can return a proper error status or fall back to the next target in your routing config.

When the gate is active

The gate is armed for streaming requests when any of these apply: The gate only applies to streaming requests. A non-streaming request already returns in one piece, so its status code is accurate and there is nothing to hold back.
When a request is both Anthropic and a native Responses call, the Responses validation takes precedence. A TTFT timeout applies on top of either.

Fail fast with a TTFT timeout

Send x-tfy-ttft-timeout-ms to cap how long the AI Gateway waits for the first chunk. The value is an integer in milliseconds.
If no chunk arrives within the window, the gateway cancels the upstream call and returns HTTP 408:
If the provider closes an empty stream before the window expires, the message instead reads TTFT timeout: stream ended before first chunk (timeout=30000ms). Behaviour worth knowing:
  • Fallback is unconditional. With a virtual model or routing config, a TTFT timeout falls back to the next eligible target even if 408 is not in your fallback status codes. The target is not put into cooldown, because a slow first token is not treated as target failure.
  • The upstream call is actually cancelled, not just abandoned, so you are not left paying for a request nobody is reading.
  • Only the first chunk is timed. Once streaming starts, this header no longer applies. Use x-tfy-request-timeout to bound the whole attempt.
  • The header is ignored on non-streaming requests, and a value that is not an integer is rejected with HTTP 400.
  • If every target times out, the last 408 is returned to you.
Set the timeout above your model’s normal first-token latency, not at it. Reasoning models and long prompts routinely take several seconds before the first token. A value that is too tight turns healthy requests into fallbacks.

Anthropic overload detection

Anthropic can answer 200 OK and place an overloaded_error in the first stream event. For every Anthropic streaming request, the gateway inspects the first non-empty chunk. If it carries that overload signal, the attempt fails with HTTP 503 instead of opening the stream. 503 is in the default fallback status codes, so with a routing config or virtual model the request moves to the next target automatically. Unlike the TTFT path, this does count toward the target’s cooldown and error metrics, because an overloaded provider is a genuinely unhealthy target. Detection is deliberately narrow: the gateway parses the stream event and matches the overload signal structurally, so a model that merely writes the words “overloaded error” in its output does not trip the check.
The gate itself runs on every Anthropic streaming request, including single-target ones. Falling back to a different model requires a routing config or virtual model with fallback targets configured.

In-band errors on the Responses API

Native /v1/responses streams begin with lifecycle events (response.created, response.in_progress, response.queued) before any output. A provider can send those and then report a failure, so the request looks successful until it isn’t. The gateway holds these lifecycle frames — up to four of them — and inspects what follows. When output generation starts, the stream is released to your client. When a response.failed or error event arrives instead, the request is failed with a status mapped from the provider’s own error type or code: When a provider sends both a type and a code, the code is used, since it is the more specific of the two. The error message you receive keeps the provider’s own wording and appends the taxonomy, for example Rate limit reached for gpt-4o (rate_limit_exceeded). The mapping is what makes fallback behave sensibly: 429, 500, 502 and 503 are fallback-eligible, so a capacity or server problem moves to the next target. The 400 cases are terminal on purpose — a prompt that is too long or malformed will fail the same way on every target, so retrying it wastes time and quota. Two provider-specific cases:
  • Microsoft Foundry splits the failure across two frames — a response.failed with no reason, then a separate error frame carrying it. The gateway waits one more frame for the detail. If it never arrives, the request fails with 502 and the message Provider reported a failure in the response stream without a reason.
  • An empty stream — the provider closes the connection having produced no output at all — fails with 502 and Provider closed the response stream before producing any output. This is classified as a provider failure rather than a timeout so that the load balancer treats it as fallback-eligible.
On a self-hosted AI Gateway, set the RESPONSES_IN_BAND_ERROR_FALLBACK_ENABLED environment variable to false to turn this check off. It defaults to true.With it disabled, a native Responses stream is released to your client as soon as it starts, and an in-band failure reaches your application as a 200 stream containing an error event. Only disable it if you have a client that depends on receiving those raw error events.

Observability

Failures caught by the gate are recorded as real failures, not as successful requests: Anthropic overload detections are logged at debug level rather than error, since they are an expected and handled condition.

Routing and fallbacks

Configure the fallback targets the gate routes to.

Request headers

Full reference for x-tfy-ttft-timeout-ms and x-tfy-request-timeout.