Blank white background with no objects or features visible.

We’re sharing complimentary access to the full Gartner Hype Cycle for AI Governance 2026. Get your copy →

How to Configure Guardrails on the AI Gateway

By Ashish Dubey

Published: October 9, 2026

Why guardrails belong on the gateway

A raw model endpoint will happily return personal data, accept a prompt injection payload, or serve an unsafe response. Guardrails fix this by inspecting and, where needed, rewriting traffic before it reaches the model and before the response reaches the user.

Putting guardrails on the AI Gateway means every application gets the same protection automatically, without each team reimplementing safety. TrueFoundry ships built in guardrails for PII redaction and prompt injection detection, both fully managed with no third party keys to wire up.

The two-step model: register, then policy

As the guardrails docs put it, applying guardrails on the AI Gateway is a two step process.

  • Register guardrails. Go to AI Gateway, then Guardrails, create a guardrails group, and add the guardrail integrations you want, whether TrueFoundry built ins, external providers, or custom guardrails.
  • Configure policies. Go to AI Gateway, then Policies, then Guardrails, and create rules that decide when to apply which guardrails and on which hooks they run.
Selecting built-in guardrails such as PII and Prompt Injection from TrueFoundry Guardrails.

Building a guardrail rule

In Policies, then Guardrails, click Add Rule. Each rule has a unique Rule ID and a few sections that decide who it applies to and where it runs.

  • When request goes to (targets). Match on one or more models, or on MCP servers and even specific tools. Multiple target conditions combine with OR, so the rule matches if the request goes to any listed model or MCP server. No target means it matches any model.
  • From subjects. Filter by users, teams, or virtual accounts with IN or NOT IN conditions. No subject filter means the rule applies to all callers.
  • With metadata. Match on key value pairs sent in the X-TFY-METADATA header, for example environment production.
  • Apply on hooks. Pick the hook and select the guardrails to run on it. You can attach multiple guardrails to the same hook, and all of them run for matching requests.
The guardrail rule editor, with targets, subjects, metadata, and hooks.

The four hooks

  • LLM Input runs before the prompt is sent to the model.
  • LLM Output runs after the model responds, before the response is returned.
  • MCP Tool Pre-Invoke runs before an MCP tool is executed.
  • MCP Tool Post-Invoke runs after an MCP tool returns, before the result reaches the model.

All rules are evaluated for every request, and the guardrails from all matching rules are combined and applied together. If one rule applies PII detection on LLM Input and another applies prompt injection detection on LLM Input, both run.

Want the same safety on every LLM and tool call?
We will configure PII redaction and prompt-injection defense on your gateway, live.

Configuring PII and PHI redaction

The PII and PHI detection guardrail is a built in TrueFoundry guardrail that identifies and redacts personally identifiable information and protected health information. It is powered by Azure AI Language PII Detection under the hood and is fully managed, so there are no third party API keys.

It only supports mutate mode, which means it always redacts detected entities. In the config form you set a name, choose PII categories or keep the default of all categories, and pick an enforcing strategy. Detected values are replaced with asterisks, so a phone number and email in a message become masked before the model ever sees them.

Configuring the built-in PII and PHI detection guardrail with entity categories.

A common setup applies PII redaction on LLM Input to clean user messages, on LLM Output to clean responses, and on the MCP hooks to strip PII from tool parameters and tool results such as database rows.

Configuring prompt-injection defense

The prompt injection guardrail is a built in guardrail powered by Azure Prompt Shield. It detects direct prompt injection, jailbreak attacks such as the do anything now pattern, and indirect injection hidden inside document or context content. It analyzes the user prompt and the document content separately.

It only supports validate mode, which means it detects and blocks attacks but does not modify content. The config asks only for a name and an enforcing strategy. The docs recommend starting with the audit strategy to monitor detections in request traces, then switching to enforce once you trust it.

How the gateway runs guardrails on a request

On an LLM request, input mutation guardrails run first and block until they finish, for example redacting PII from the prompt. Input validation, such as prompt injection, then runs in the background while the model request is in flight. If input validation fails while the model is still running, the gateway cancels the model request right away so you do not pay for it.

After the model responds, output mutation can strip secrets, and output validation checks the final result before it reaches the client. For MCP tools, all pre tool guardrails run before the tool is called, and if any fail the tool never executes. Guardrails run on every tool call separately.

Guardrail latency and results surfaced in request traces.
Make guardrails standard across every team
See centralized PII and injection policies applied to all apps, with full tracing, on your infrastructure.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Start free
Table of Contents

One Gateway for Every LLM, Agent and MCP Server

Book a 30-min with our AI expert

Book a Demo

The fastest way to build, govern and scale your AI

Book Demo
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Discover More

No items found.
October 9, 2026
|
5 min read

Budgets and Quotas on the AI Gateway: Control AI Spend by Team

No items found.
October 9, 2026
|
5 min read

Gateway Tracing and Request Logs: Debug Every LLM Call

No items found.
October 9, 2026
|
5 min read

How to Configure Guardrails on the AI Gateway

No items found.
October 9, 2026
|
5 min read

OpenRouter BYOK explained: cheaper, often faster, and changing

No items found.
No items found.

Recent Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Frequently asked questions

What guardrails does TrueFoundry provide out of the box?

TrueFoundry ships built in PII and PHI detection, powered by Azure AI Language, and prompt injection detection, powered by Azure Prompt Shield. Both are fully managed, with no external credentials required. You can also add external providers or custom guardrails.

Where do guardrails run?

On four hooks: LLM Input, LLM Output, MCP Tool Pre-Invoke, and MCP Tool Post-Invoke. You attach guardrails to the hooks you want in a policy rule, and multiple guardrails can run on the same hook.

Does a blocked prompt still cost money?

If input validation fails while the model request is in flight, the gateway cancels that model request so you do not pay for it. Output validation runs after the model responds, so in that case the model cost is already incurred.

How should I roll out prompt-injection detection safely?

Start with the audit enforcing strategy so detections are recorded in request traces without blocking traffic. Once you are confident in the results, switch the strategy to enforce.

Take a quick product tour
Start Product Tour
Product Tour