The three billing policies
Which policy applies is determined by the provider:
This mapping is fixed in the AI Gateway and is not configurable per provider or per request. It is based on measured provider behaviour; a provider moves off the default only once its cancellation billing has been confirmed.
How tokens are counted on the estimate path
Input and output tokens come from different places, with different accuracy. Input tokens are exact wherever the provider reports them early. Anthropic sends usage twice — an input-only block in the opening event, and the final block at the end. A cancelled stream only ever receives the first, so the AI Gateway keeps that opening count and uses it instead of a guess. The same applies to Google Vertex AI and to Anthropic models served through Bedrock and Azure. This matters: on a real prompt, the provider’s exact input count was 154 tokens where the character heuristic produced 111. Output tokens are always estimated. The provider never reports them for a cancelled stream, so the AI Gateway approximates from the text it actually streamed, at roughly one token per 3.5 characters, plus a small per-message structural allowance. Tool-call arguments and reasoning content are included in the estimate. Image and file blocks are skipped, because their base64 payload would dwarf the real token count. Cost is then calculated from those counts using the same public or private pricing as any other request.Identifying disconnected requests
A disconnected request is recorded as a request that both carries cost and reports a client disconnect. There is no separate “partial” flag.tfy.model.tokens_estimated is the attribute to filter on when you want to separate measured spend from approximated spend in your own analysis. See GenAI span attributes for how to read these from exported traces.
A client that disconnects immediately after a stream completes is not a disconnect. The AI Gateway marks the stream complete the moment the terminal event is written, so a late hang-up is recorded as an ordinary successful request.
Budgets and rate limits
Budgets. Attributed cost from a disconnected stream counts against budget limits exactly like a completed request. Requests on the do-not-bill path carry zero cost and consume no budget. Rate limits. A cancelled request consumes a request-quota slot under all three policies. Rate limits are counted on errors as well as successes, so a client that repeatedly starts and aborts requests cannot bypass a requests-per-minute limit.Non-streaming requests
A non-streaming response arrives in one piece, so cancelling it leaves nothing partial to attribute. For that reason the AI Gateway finishes the upstream call when a client cancels a non-streaming request, and bills the provider’s exact usage block. This applies to every provider, including those on the do-not-bill streaming policy.Changing non-streaming cancel behaviour (self-hosted)
Changing non-streaming cancel behaviour (self-hosted)
On a self-hosted AI Gateway, set
COMPLETE_NON_STREAMING_ON_CLIENT_CANCEL to false to abort the upstream call instead of completing it. It defaults to true.With it disabled, a cancelled non-streaming request is aborted and attributes no cost, at the risk of under-reporting spend that the provider still charges for.Caching
A truncated stream is never written to the response cache, even though it is billed. A cut-off completion stored as a cache entry would be replayed to later requests as if it were a whole answer. On the do-not-cancel path the upstream stream is read to completion, so that response is cached normally.Current limitations
- Native Responses API streams are not estimated. A cancelled
/v1/responsesstream that never received its terminal usage event is recorded with zero tokens and zero cost. In practice the main Responses providers — OpenAI, Azure OpenAI and Azure AI Foundry — all sit on the do-not-cancel or do-not-bill policies, so the estimate path is rarely the one in play. - Gemini and Bedrock passthrough endpoints read provider usage only. Requests served through those native passthrough routes attribute no cost when cancelled before the provider reports usage.
- Proxy and text-to-speech passthrough requests are never billed on cancel.
- A network failure is not a client disconnect. If the provider connection dies mid-stream, the request is recorded as an error, not as billable truncation. Likewise, when a provider reports an in-band failure and the client disconnects at the same time, the provider error takes precedence in the trace.
Related
Cost tracking
How pricing is resolved and how to break costs down.
First chunk gate
How streaming failures before any output are handled.