How to Configure Guardrails on the AI Gateway
.png)
Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
Why guardrails belong on the gateway
A raw model endpoint will happily return personal data, accept a prompt injection payload, or serve an unsafe response. Guardrails fix this by inspecting and, where needed, rewriting traffic before it reaches the model and before the response reaches the user.
Putting guardrails on the AI Gateway means every application gets the same protection automatically, without each team reimplementing safety. TrueFoundry ships built in guardrails for PII redaction and prompt injection detection, both fully managed with no third party keys to wire up.
The two-step model: register, then policy
As the guardrails docs put it, applying guardrails on the AI Gateway is a two step process.
- Register guardrails. Go to AI Gateway, then Guardrails, create a guardrails group, and add the guardrail integrations you want, whether TrueFoundry built ins, external providers, or custom guardrails.
- Configure policies. Go to AI Gateway, then Policies, then Guardrails, and create rules that decide when to apply which guardrails and on which hooks they run.

Building a guardrail rule
In Policies, then Guardrails, click Add Rule. Each rule has a unique Rule ID and a few sections that decide who it applies to and where it runs.
- When request goes to (targets). Match on one or more models, or on MCP servers and even specific tools. Multiple target conditions combine with OR, so the rule matches if the request goes to any listed model or MCP server. No target means it matches any model.
- From subjects. Filter by users, teams, or virtual accounts with IN or NOT IN conditions. No subject filter means the rule applies to all callers.
- With metadata. Match on key value pairs sent in the X-TFY-METADATA header, for example environment production.
- Apply on hooks. Pick the hook and select the guardrails to run on it. You can attach multiple guardrails to the same hook, and all of them run for matching requests.

The four hooks
- LLM Input runs before the prompt is sent to the model.
- LLM Output runs after the model responds, before the response is returned.
- MCP Tool Pre-Invoke runs before an MCP tool is executed.
- MCP Tool Post-Invoke runs after an MCP tool returns, before the result reaches the model.
All rules are evaluated for every request, and the guardrails from all matching rules are combined and applied together. If one rule applies PII detection on LLM Input and another applies prompt injection detection on LLM Input, both run.
Configuring PII and PHI redaction
The PII and PHI detection guardrail is a built in TrueFoundry guardrail that identifies and redacts personally identifiable information and protected health information. It is powered by Azure AI Language PII Detection under the hood and is fully managed, so there are no third party API keys.
It only supports mutate mode, which means it always redacts detected entities. In the config form you set a name, choose PII categories or keep the default of all categories, and pick an enforcing strategy. Detected values are replaced with asterisks, so a phone number and email in a message become masked before the model ever sees them.

A common setup applies PII redaction on LLM Input to clean user messages, on LLM Output to clean responses, and on the MCP hooks to strip PII from tool parameters and tool results such as database rows.
Configuring prompt-injection defense
The prompt injection guardrail is a built in guardrail powered by Azure Prompt Shield. It detects direct prompt injection, jailbreak attacks such as the do anything now pattern, and indirect injection hidden inside document or context content. It analyzes the user prompt and the document content separately.
It only supports validate mode, which means it detects and blocks attacks but does not modify content. The config asks only for a name and an enforcing strategy. The docs recommend starting with the audit strategy to monitor detections in request traces, then switching to enforce once you trust it.
How the gateway runs guardrails on a request
On an LLM request, input mutation guardrails run first and block until they finish, for example redacting PII from the prompt. Input validation, such as prompt injection, then runs in the background while the model request is in flight. If input validation fails while the model is still running, the gateway cancels the model request right away so you do not pay for it.
After the model responds, output mutation can strip secrets, and output validation checks the final result before it reaches the client. For MCP tools, all pre tool guardrails run before the tool is called, and if any fail the tool never executes. Guardrails run on every tool call separately.

TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
What guardrails does TrueFoundry provide out of the box?
TrueFoundry ships built in PII and PHI detection, powered by Azure AI Language, and prompt injection detection, powered by Azure Prompt Shield. Both are fully managed, with no external credentials required. You can also add external providers or custom guardrails.
Where do guardrails run?
On four hooks: LLM Input, LLM Output, MCP Tool Pre-Invoke, and MCP Tool Post-Invoke. You attach guardrails to the hooks you want in a policy rule, and multiple guardrails can run on the same hook.
Does a blocked prompt still cost money?
If input validation fails while the model request is in flight, the gateway cancels that model request so you do not pay for it. Output validation runs after the model responds, so in that case the model cost is already incurred.
How should I roll out prompt-injection detection safely?
Start with the audit enforcing strategy so detections are recorded in request traces without blocking traffic. Once you are confident in the results, switch the strategy to enforce.










.png)
.png)

.png)
.png)





.png)



.png)





