200 OK and only then report a failure inside the stream itself — an Anthropic overloaded_error in the first event, or an OpenAI response.failed event after response.created. Once the AI Gateway has written a 200 and the first bytes to your client, that request can no longer be retried or routed elsewhere: your application receives a successful-looking stream that carries no usable output.
The first chunk gate is the mechanism that prevents this. For streaming requests, the AI Gateway reads the first chunk from the provider before it starts the response to your client. Only when that chunk looks like real generation does the gateway open the stream. If the chunk carries a failure, nothing has been sent yet, so the gateway can return a proper error status or fall back to the next target in your routing config.
When the gate is active
The gate is armed for streaming requests when any of these apply:
The gate only applies to streaming requests. A non-streaming request already returns in one piece, so its status code is accurate and there is nothing to hold back.
When a request is both Anthropic and a native Responses call, the Responses validation takes precedence. A TTFT timeout applies on top of either.
Fail fast with a TTFT timeout
Sendx-tfy-ttft-timeout-ms to cap how long the AI Gateway waits for the first chunk. The value is an integer in milliseconds.
TTFT timeout: stream ended before first chunk (timeout=30000ms).
Behaviour worth knowing:
- Fallback is unconditional. With a virtual model or routing config, a TTFT timeout falls back to the next eligible target even if
408is not in your fallback status codes. The target is not put into cooldown, because a slow first token is not treated as target failure. - The upstream call is actually cancelled, not just abandoned, so you are not left paying for a request nobody is reading.
- Only the first chunk is timed. Once streaming starts, this header no longer applies. Use
x-tfy-request-timeoutto bound the whole attempt. - The header is ignored on non-streaming requests, and a value that is not an integer is rejected with HTTP 400.
- If every target times out, the last
408is returned to you.
Anthropic overload detection
Anthropic can answer200 OK and place an overloaded_error in the first stream event. For every Anthropic streaming request, the gateway inspects the first non-empty chunk. If it carries that overload signal, the attempt fails with HTTP 503 instead of opening the stream.
503 is in the default fallback status codes, so with a routing config or virtual model the request moves to the next target automatically. Unlike the TTFT path, this does count toward the target’s cooldown and error metrics, because an overloaded provider is a genuinely unhealthy target.
Detection is deliberately narrow: the gateway parses the stream event and matches the overload signal structurally, so a model that merely writes the words “overloaded error” in its output does not trip the check.
The gate itself runs on every Anthropic streaming request, including single-target ones. Falling back to a different model requires a routing config or virtual model with fallback targets configured.
In-band errors on the Responses API
Native/v1/responses streams begin with lifecycle events (response.created, response.in_progress, response.queued) before any output. A provider can send those and then report a failure, so the request looks successful until it isn’t.
The gateway holds these lifecycle frames — up to four of them — and inspects what follows. When output generation starts, the stream is released to your client. When a response.failed or error event arrives instead, the request is failed with a status mapped from the provider’s own error type or code:
When a provider sends both a type and a code, the code is used, since it is the more specific of the two. The error message you receive keeps the provider’s own wording and appends the taxonomy, for example
Rate limit reached for gpt-4o (rate_limit_exceeded).
The mapping is what makes fallback behave sensibly: 429, 500, 502 and 503 are fallback-eligible, so a capacity or server problem moves to the next target. The 400 cases are terminal on purpose — a prompt that is too long or malformed will fail the same way on every target, so retrying it wastes time and quota.
Two provider-specific cases:
- Microsoft Foundry splits the failure across two frames — a
response.failedwith no reason, then a separateerrorframe carrying it. The gateway waits one more frame for the detail. If it never arrives, the request fails with502and the messageProvider reported a failure in the response stream without a reason. - An empty stream — the provider closes the connection having produced no output at all — fails with
502andProvider closed the response stream before producing any output. This is classified as a provider failure rather than a timeout so that the load balancer treats it as fallback-eligible.
Disabling Responses in-band error detection (self-hosted)
Disabling Responses in-band error detection (self-hosted)
On a self-hosted AI Gateway, set the
RESPONSES_IN_BAND_ERROR_FALLBACK_ENABLED environment variable to false to turn this check off. It defaults to true.With it disabled, a native Responses stream is released to your client as soon as it starts, and an in-band failure reaches your application as a 200 stream containing an error event. Only disable it if you have a client that depends on receiving those raw error events.Observability
Failures caught by the gate are recorded as real failures, not as successful requests:
Anthropic overload detections are logged at debug level rather than error, since they are an expected and handled condition.
Related
Routing and fallbacks
Configure the fallback targets the gate routes to.
Request headers
Full reference for
x-tfy-ttft-timeout-ms and x-tfy-request-timeout.