Skip to main content
The chat completions API supports all OpenAI compatible parameters to fine-tune model behavior and response characteristics.

Basic Parameters

Here’s how to use common parameters:
You can pass parameters supported by the selected model. Refer to the OpenAI API Reference for parameter definitions.
The AI Gateway automatically translates max_tokens and max_completion_tokens to the parameter accepted by the selected provider. OpenAI, Azure OpenAI, and non-Anthropic Bedrock Mantle models use max_completion_tokens. Anthropic models on Bedrock Mantle and other providers use max_tokens.
Parameter support can also vary between models from the same provider. The original GPT-5 reasoning models do not accept temperature, top_p, presence_penalty, or stop, so the Gateway omits these parameters. For GPT-5.1 and later models and gpt-5-chat variants, the Gateway forwards them for the provider to validate.

Reasoning parameters

Setting reasoning_effort to anything other than none (or sending a thinking field) asks the model to reason before answering. On OpenAI and Azure OpenAI this also changes how the AI Gateway reaches the provider: the request is served through the provider’s Responses API upstream, because OpenAI exposes reasoning only there. Your request and response stay on the Chat Completions contract.
The Responses API has no equivalent for some Chat Completions parameters, so when a request is served that way these are dropped:n (above 1), best_of, frequency_penalty, presence_penalty, logit_bias, logprobs, top_logprobs, stop, seed, audio, modalities (non-text), predictionSeveral of these appear in the example above. If your request depends on one — sampling with seed, cutting generation with stop, or scoring with logprobs — leave reasoning_effort unset and don’t send the x-tfy-openai-use-responses-api header. The request then stays on Chat Completions and the parameters are honoured.Inert values are ignored rather than dropped, so n: 1, logprobs: false, and zeroed penalties never change anything.
For the triggers, the response shape, and multi-turn replay, see Reasoning on OpenAI and Azure OpenAI. For Claude and Gemini reasoning, see Extended Thinking.