> ## Documentation Index
> Fetch the complete documentation index at: https://www.truefoundry.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Cost Accounting for Disconnected Streams

> How the AI Gateway attributes tokens and cost when a client hangs up mid-stream, and how each provider's cancellation billing is handled.

When a client cancels a streaming request mid-flight — a user closes the browser tab, an agent stops generation, a pod is evicted — the provider has already generated tokens and, for most providers, already charged for them. If the AI Gateway dropped the request at that point, that spend would be invisible: absent from your cost dashboards and uncounted against budgets.

The AI Gateway instead attributes cost for interrupted streams, using a per-provider policy that matches how each provider actually bills a cancelled request.

## The three billing policies

| Policy | What the AI Gateway does | Cost attributed |
| - | - | - |
| **Estimate** (default) | Cancels the upstream call and bills the partial stream, estimating the output tokens the provider never got to report. | Partial, estimated |
| **Do not cancel** | Detaches your client but keeps reading the upstream stream to completion, so the provider's own usage block is used. | Exact, full completion |
| **Do not bill** | Cancels the upstream call and attributes nothing. | Zero |

Which policy applies is determined by the provider:

| Provider | Policy | Why |
| - | - | - |
| OpenAI | Do not cancel | OpenAI bills the full completion whether or not you keep reading, so draining the stream costs nothing extra and yields exact usage. |
| Google Vertex AI | Do not bill | Confirmed not to bill cancelled streams. |
| Azure OpenAI | Do not bill | Confirmed not to bill cancelled streams. |
| Azure AI Foundry | Do not bill | Confirmed not to bill cancelled streams. |
| Everything else — Anthropic, AWS Bedrock, Groq, Together, xAI, self-hosted models, and all other providers | Estimate | Providers generally charge for what they generated before the cancel, so billing the partial stream is the closer approximation. |

<Note>
  This mapping is fixed in the AI Gateway and is not configurable per provider or per request. It is based on measured provider behaviour; a provider moves off the default only once its cancellation billing has been confirmed.
</Note>

## How tokens are counted on the estimate path

Input and output tokens come from different places, with different accuracy.

**Input tokens are exact wherever the provider reports them early.** Anthropic sends usage twice — an input-only block in the opening event, and the final block at the end. A cancelled stream only ever receives the first, so the AI Gateway keeps that opening count and uses it instead of a guess. The same applies to Google Vertex AI and to Anthropic models served through Bedrock and Azure. This matters: on a real prompt, the provider's exact input count was 154 tokens where the character heuristic produced 111.

**Output tokens are always estimated.** The provider never reports them for a cancelled stream, so the AI Gateway approximates from the text it actually streamed, at roughly one token per 3.5 characters, plus a small per-message structural allowance. Tool-call arguments and reasoning content are included in the estimate. Image and file blocks are skipped, because their base64 payload would dwarf the real token count.

Cost is then calculated from those counts using the same [public or private pricing](/docs/ai-gateway/cost-tracking) as any other request.

<Warning>
  Estimated output tokens are an approximation, not a provider-reported figure. Expect the cost on a disconnected request to differ from what the provider invoices. Treat these rows as directional spend, and reconcile against provider billing when precision matters.
</Warning>

## Identifying disconnected requests

A disconnected request is recorded as a request that both carries cost and reports a client disconnect. There is no separate "partial" flag.

| Signal | Value |
| - | - |
| Request log `error_code` | `499` |
| Request log `error_detail` | `client-disconnected` |
| Request log `cost`, `input_tokens`, `output_tokens` | Non-zero on the estimate and do-not-cancel paths, zero on do-not-bill |
| Prometheus `ai_gateway_gateway_http_status` | `status_code="499"`, `error_origin_level="client_disconnected"` |
| Span attribute `tfy.model.tokens_estimated` | `true` when any part of the usage was estimated rather than provider-reported |
| Span `error.type` | `client-disconnected` |

`tfy.model.tokens_estimated` is the attribute to filter on when you want to separate measured spend from approximated spend in your own analysis. See [GenAI span attributes](/docs/ai-gateway/fetch-request-logs-span-attributes-genai) for how to read these from exported traces.

<Note>
  A client that disconnects immediately *after* a stream completes is not a disconnect. The AI Gateway marks the stream complete the moment the terminal event is written, so a late hang-up is recorded as an ordinary successful request.
</Note>

## Budgets and rate limits

**Budgets.** Attributed cost from a disconnected stream counts against [budget limits](/docs/ai-gateway/budget-limiting-v2) exactly like a completed request. Requests on the do-not-bill path carry zero cost and consume no budget.

**Rate limits.** A cancelled request consumes a request-quota slot under all three policies. [Rate limits](/docs/ai-gateway/ratelimiting-v2) are counted on errors as well as successes, so a client that repeatedly starts and aborts requests cannot bypass a requests-per-minute limit.

## Non-streaming requests

A non-streaming response arrives in one piece, so cancelling it leaves nothing partial to attribute. For that reason the AI Gateway finishes the upstream call when a client cancels a non-streaming request, and bills the provider's exact usage block. This applies to every provider, including those on the do-not-bill streaming policy.

<Accordion title="Changing non-streaming cancel behaviour (self-hosted)">
  On a self-hosted AI Gateway, set `COMPLETE_NON_STREAMING_ON_CLIENT_CANCEL` to `false` to abort the upstream call instead of completing it. It defaults to `true`.

  With it disabled, a cancelled non-streaming request is aborted and attributes no cost, at the risk of under-reporting spend that the provider still charges for.
</Accordion>

## Caching

A truncated stream is never written to the [response cache](/docs/ai-gateway/caching), even though it is billed. A cut-off completion stored as a cache entry would be replayed to later requests as if it were a whole answer.

On the do-not-cancel path the upstream stream is read to completion, so that response *is* cached normally.

## Current limitations

* **Native Responses API streams are not estimated.** A cancelled `/v1/responses` stream that never received its terminal usage event is recorded with zero tokens and zero cost. In practice the main Responses providers — OpenAI, Azure OpenAI and Azure AI Foundry — all sit on the do-not-cancel or do-not-bill policies, so the estimate path is rarely the one in play.
* **Gemini and Bedrock passthrough endpoints read provider usage only.** Requests served through those native passthrough routes attribute no cost when cancelled before the provider reports usage.
* **Proxy and text-to-speech passthrough requests are never billed on cancel.**
* **A network failure is not a client disconnect.** If the provider connection dies mid-stream, the request is recorded as an error, not as billable truncation. Likewise, when a provider reports an in-band failure and the client disconnects at the same time, the provider error takes precedence in the trace.

## Related

<Columns cols={2}>
  <Card title="Cost tracking" href="/docs/ai-gateway/cost-tracking">
    How pricing is resolved and how to break costs down.
  </Card>

  <Card title="First chunk gate" href="/docs/ai-gateway/first-chunk-gate">
    How streaming failures before any output are handled.
  </Card>
</Columns>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.