> ## Documentation Index
> Fetch the complete documentation index at: https://www.truefoundry.com/llms.txt
> Use this file to discover all available pages before exploring further.

# First Chunk Gate for Streaming

> How the AI Gateway holds a streaming response until the first chunk is validated, so in-band provider failures can fall back instead of reaching your client.

A streaming provider can answer `200 OK` and only then report a failure inside the stream itself — an Anthropic `overloaded_error` in the first event, or an OpenAI `response.failed` event after `response.created`. Once the AI Gateway has written a `200` and the first bytes to your client, that request can no longer be retried or routed elsewhere: your application receives a successful-looking stream that carries no usable output.

The first chunk gate is the mechanism that prevents this. For streaming requests, the AI Gateway reads the first chunk from the provider **before** it starts the response to your client. Only when that chunk looks like real generation does the gateway open the stream. If the chunk carries a failure, nothing has been sent yet, so the gateway can return a proper error status or fall back to the next target in your routing config.

## When the gate is active

The gate is armed for streaming requests when any of these apply:

| Trigger | Applies to | Controlled by |
| - | - | - |
| Time to first token (TTFT) timeout | All providers | The `x-tfy-ttft-timeout-ms` request header. Off unless you send it. |
| Anthropic in-stream overload | Anthropic models | Always on. No configuration. |
| Native Responses API in-band errors | `/v1/responses` requests (OpenAI, Microsoft Foundry) | On by default. |

The gate only applies to streaming requests. A non-streaming request already returns in one piece, so its status code is accurate and there is nothing to hold back.

<Note>
  When a request is both Anthropic and a native Responses call, the Responses validation takes precedence. A TTFT timeout applies on top of either.
</Note>

## Fail fast with a TTFT timeout

Send `x-tfy-ttft-timeout-ms` to cap how long the AI Gateway waits for the first chunk. The value is an integer in milliseconds.

```bash theme={"dark"}
curl https://your-control-plane.truefoundry.cloud/api/llm/chat/completions \
  -H "Authorization: Bearer $TFY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "x-tfy-ttft-timeout-ms: 30000" \
  -d '{
    "model": "openai-main/gpt-4o",
    "stream": true,
    "messages": [{"role": "user", "content": "Summarise the Q3 incident report."}]
  }'
```

If no chunk arrives within the window, the gateway cancels the upstream call and returns **HTTP 408**:

```json theme={"dark"}
{
  "status": "failure",
  "message": "TTFT timeout: no first chunk within 30000ms",
  "error": {
    "message": "TTFT timeout: no first chunk within 30000ms",
    "type": "TTFTTimeoutError",
    "code": "408",
    "param": null
  },
  "error_origin_level": "ttft_timeout"
}
```

If the provider closes an empty stream before the window expires, the message instead reads `TTFT timeout: stream ended before first chunk (timeout=30000ms)`.

Behaviour worth knowing:

* **Fallback is unconditional.** With a [virtual model](/docs/ai-gateway/virtual-model) or [routing config](/docs/ai-gateway/load-balancing-overview), a TTFT timeout falls back to the next eligible target even if `408` is not in your fallback status codes. The target is not put into cooldown, because a slow first token is not treated as target failure.
* **The upstream call is actually cancelled**, not just abandoned, so you are not left paying for a request nobody is reading.
* **Only the first chunk is timed.** Once streaming starts, this header no longer applies. Use `x-tfy-request-timeout` to bound the whole attempt.
* **The header is ignored on non-streaming requests**, and a value that is not an integer is rejected with HTTP 400.
* If every target times out, the last `408` is returned to you.

<Tip>
  Set the timeout above your model's normal first-token latency, not at it. Reasoning models and long prompts routinely take several seconds before the first token. A value that is too tight turns healthy requests into fallbacks.
</Tip>

## Anthropic overload detection

Anthropic can answer `200 OK` and place an `overloaded_error` in the first stream event. For every Anthropic streaming request, the gateway inspects the first non-empty chunk. If it carries that overload signal, the attempt fails with **HTTP 503** instead of opening the stream.

`503` is in the default fallback status codes, so with a routing config or virtual model the request moves to the next target automatically. Unlike the TTFT path, this does count toward the target's cooldown and error metrics, because an overloaded provider is a genuinely unhealthy target.

Detection is deliberately narrow: the gateway parses the stream event and matches the overload signal structurally, so a model that merely writes the words "overloaded error" in its output does not trip the check.

<Note>
  The gate itself runs on every Anthropic streaming request, including single-target ones. Falling back to a different model requires a routing config or virtual model with fallback targets configured.
</Note>

## In-band errors on the Responses API

Native `/v1/responses` streams begin with lifecycle events (`response.created`, `response.in_progress`, `response.queued`) before any output. A provider can send those and then report a failure, so the request looks successful until it isn't.

The gateway holds these lifecycle frames — up to four of them — and inspects what follows. When output generation starts, the stream is released to your client. When a `response.failed` or `error` event arrives instead, the request is failed with a status mapped from the provider's own error type or code:

| Provider error type or code | HTTP status returned |
| - | - |
| `too_many_requests`, `rate_limit_exceeded`, `no_capacity`, `insufficient_quota`, `credit_balance_exhausted` | 429 |
| `server_error`, `api_error` | 500 |
| `overloaded_error` | 503 |
| `authentication_error` | 401 |
| `permission_error`, `permission_denied` | 403 |
| `not_found_error` | 404 |
| `invalid_request_error`, `invalid_prompt`, `context_length_exceeded`, `content_filter` | 400 |
| Anything unrecognised | 502 |

When a provider sends both a type and a code, the code is used, since it is the more specific of the two. The error message you receive keeps the provider's own wording and appends the taxonomy, for example `Rate limit reached for gpt-4o (rate_limit_exceeded)`.

The mapping is what makes fallback behave sensibly: `429`, `500`, `502` and `503` are fallback-eligible, so a capacity or server problem moves to the next target. The `400` cases are terminal on purpose — a prompt that is too long or malformed will fail the same way on every target, so retrying it wastes time and quota.

Two provider-specific cases:

* **Microsoft Foundry splits the failure across two frames** — a `response.failed` with no reason, then a separate `error` frame carrying it. The gateway waits one more frame for the detail. If it never arrives, the request fails with `502` and the message `Provider reported a failure in the response stream without a reason`.
* **An empty stream** — the provider closes the connection having produced no output at all — fails with `502` and `Provider closed the response stream before producing any output`. This is classified as a provider failure rather than a timeout so that the load balancer treats it as fallback-eligible.

<Accordion title="Disabling Responses in-band error detection (self-hosted)">
  On a self-hosted AI Gateway, set the `RESPONSES_IN_BAND_ERROR_FALLBACK_ENABLED` environment variable to `false` to turn this check off. It defaults to `true`.

  With it disabled, a native Responses stream is released to your client as soon as it starts, and an in-band failure reaches your application as a `200` stream containing an error event. Only disable it if you have a client that depends on receiving those raw error events.
</Accordion>

## Observability

Failures caught by the gate are recorded as real failures, not as successful requests:

| Where | What you see |
| - | - |
| Request traces | `error.type` of `ttft-timeout` or `provider-stream-error`, with the mapped status code rather than the provider's wire `200` |
| Span attributes | `gen_ai.response.time_to_first_chunk`, in seconds |
| Prometheus | `ai_gateway_gateway_http_status` labelled with `error_origin_level` (`ttft_timeout` for TTFT) |
| Gateway logs | TTFT activity is prefixed `[TTFT]`, including `Timeout fired for target <target> (no first chunk within <N>ms), falling back` |

Anthropic overload detections are logged at debug level rather than error, since they are an expected and handled condition.

## Related

<Columns cols={2}>
  <Card title="Routing and fallbacks" href="/docs/ai-gateway/load-balancing-overview">
    Configure the fallback targets the gate routes to.
  </Card>

  <Card title="Request headers" href="/docs/ai-gateway/request-headers">
    Full reference for `x-tfy-ttft-timeout-ms` and `x-tfy-request-timeout`.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.