Basic Parameters
Here’s how to use common parameters:The AI Gateway automatically translates
max_tokens and max_completion_tokens
to the parameter accepted by the selected provider. OpenAI, Azure OpenAI, and
non-Anthropic Bedrock Mantle models use max_completion_tokens. Anthropic models
on Bedrock Mantle and other providers use max_tokens.Parameter support can also vary between models from the same provider. The original
GPT-5 reasoning models do not accept
temperature, top_p, presence_penalty, or
stop, so the Gateway omits these parameters. For GPT-5.1 and later models and
gpt-5-chat variants, the Gateway forwards them for the provider to validate.Reasoning parameters
Settingreasoning_effort to anything other than none (or sending a thinking field) asks the model to reason before answering. On OpenAI and Azure OpenAI this also changes how the AI Gateway reaches the provider: the request is served through the provider’s Responses API upstream, because OpenAI exposes reasoning only there. Your request and response stay on the Chat Completions contract.
For the triggers, the response shape, and multi-turn replay, see Reasoning on OpenAI and Azure OpenAI. For Claude and Gemini reasoning, see Extended Thinking.