Data Masking in the AI Gateway: What Actually Works
.png)
Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
What data masking actually is
Data masking replaces a sensitive value with a substitute that is safe to expose, while keeping the surrounding data usable. It comes from the database world: give QA a copy of production where every card number is 4111-****-****-1234 and they can test realistic shapes without holding real PII. Point that at an LLM gateway and it changes — you are masking free text, on its way to a third party, inside a conversation that has to keep making sense.
These get used as synonyms constantly. They are not synonyms.
Reversible vs irreversible is the distinction that matters. Tokenisation, FPE and encryption round-trip — someone holds a vault or a key, a new asset to protect. Masking, redaction and anonymisation are one-way: simpler, safer, and useless the moment something downstream needs the real value.
Pseudonymisation is not anonymisation. Under GDPR, pseudonymised data is still personal data and still in scope; only genuinely anonymised data falls outside the regulation (Recital 26). Replacing every name with a stable User-8812 is pseudonymisation, and calling it the latter will not survive review.
A third axis is easy to miss: consistency is separate from reversibility. A masker can be deterministic — the same input always yields the same placeholder — without being reversible, and determinism is what lets a model reason about “the same person” across a conversation.
What breaks when you mask LLM traffic
This is the part most write-ups skip, and it decides whether your design survives.
Model output quality degrades, unevenly. Turn Call our office at 312-555-1234 into Call our office at ************ and the model no longer knows that was a phone number — asterisks carry no type information. Labelled placeholders like [PHONE_NUMBER] keep the type, at the cost of telling anyone who sees the prompt what was there. Summarisation barely notices; extraction fails loudly; correlating two entities fails quietly, which is worse.
Tool calls need the real value. The sharpest failure. An agent that masks customer@acme.com on the way in cannot then call lookup_account(email=...). Masking the model path and leaving the tool path open creates the opposite hole: the model never saw the address, but the tool arguments carry it to a third-party MCP server. Decide, per tool, which is more trusted.
Structured output and JSON schemas break. If you are doing structured outputs with JSON Schema, masking can produce a value that no longer satisfies its own field: a format: email field cannot hold [REDACTED EMAIL], and an integer field cannot hold asterisks. Length-preserving masking is friendlier because it keeps the shape. Masking the output leg is worse — the model produced valid JSON, the masker rewrote a value inside it, and your parser throws on a response the model got right.
Multi-turn consistency. Turn one: “Priya Raman is the account owner.” Turn four: “does she still own it?” If turn one became asterisks, the model has no thread to follow. Deterministic pseudonymisation fixes this, but only if the masker keeps state across turns — which a stateless guardrail does not.
RAG retrieval against masked text. Mask documents before indexing and the embedding no longer contains the entity, so a query for it will not retrieve the chunk. Index unmasked and mask at retrieval time, and your vector store holds the unmasked corpus — the thing you were avoiding.
Applied narrowly, masking is worth it. Applied globally, it quietly makes the product worse and nobody connects the two.
Where teams get this wrong
Assuming masking round-trips. A team masks PII on the way to the model, assumes something unmasks it coming back, and builds on that. Unless you chose a reversible scheme and something holds keys, the caller sees the masked value too.
Masking the prompt and forgetting the tool path. Data moves in four places: into the model, out of the model, into a tool, out of a tool. Covering one is common. The leak then happens through a tool result nobody was inspecting.
Treating log redaction as a data-flow control. Redacting stored logs limits who inside your company reads prompts. The provider still received the original text.
Choosing irreversible masking for a reversible problem. If something downstream needs the value back, masking is the wrong technique. You need tokenisation or FPE, which means a vault or a key.
How masking works in TrueFoundry
Masking is not a separate product here. Every guardrail carries two settings — Operation Mode and Enforcement Strategy — masking is a value of the first.
Mutate vs Validate
Priority is the whole ordering model for masking. Three mutate guardrails on one hook run in order, each seeing the output of the one before: a secrets masker at 1 and a regex masker at 2 means the regex guardrail matches text that already reads ***REDACTED***.
Input mutation is blocking and runs first — the model request cannot start until the prompt is final. Input validation then runs in the background alongside it, and if validation fails mid-flight the gateway cancels the model request so you do not pay for it. Output mutation and validation run synchronously after the response arrives, by which point the model cost is incurred.

MCP tool calls add two synchronous hooks: pre-invoke, where a failure means the tool never runs, and post-invoke, where a failure withholds the result. Guardrails run on every tool call separately — five tools, five sets of checks.

Enforce vs Enforce But Ignore On Error
The second setting separates two failures: the guardrail found something, versus the guardrail broke.
Strategy
On violation
On guardrail error
Enforce
Block
Block (fail-closed)
Enforce But Ignore On Error
Block
Let through (graceful degradation)
Audit
Let through, log only
Let through
The docs call Enforce But Ignore On Error “the safest default” and recommend the ladder Audit → Enforce But Ignore On Error → Enforce.

Which built-ins can actually mask
“We have a PII guardrail” tells you nothing about whether it can mask.
PII / PHI Detection is the odd one: it cannot block, only mask. If the requirement is “reject any request containing an SSN”, this is not the guardrail.

Its documented enforcing-strategy options are only enforce and enforce_but_ignore_on_error — audit is not listed, unlike Secrets, Regex and SQL Sanitizer. Audit-first rollout is not uniform across the built-ins. [VERIFY] whether that is a real restriction or a docs gap.
Regex Pattern Matching gives labelled placeholders rather than asterisks. It ships 46 presets — SSN, credit cards, passports for ten countries, cloud keys, GitHub and Slack tokens, IPs — plus custom patterns, such as EMP-\d{6} → [REDACTED EMPLOYEE ID].

Among the third-party guardrail integrations, mutate-capable ones include Google Model Armor, CrowdStrike AIDR, F5 AI Security, TrojAI, Enkrypt AI, Pillar Security, Lasso (classifix), HiddenLayer (/redact-*), DeepKeep and Noma — the last two require Mutate. Validate-only: Cisco AI Defense, Gray Swan Cygnal, Patronus, NVIDIA NeMo, Guardrails AI, Arthur AI. Asking a validate-only vendor to mask is a dead end.
Google Model Armor carries a setup trap: Mutate needs both an SDP inspect template and a de-identify template in Google Cloud, linked to the Model Armor template, with the transformation set to “Replace with infoType name”. Without both, mutate fails.

The unmasking question, answered plainly
TrueFoundry’s own masking guardrails are one-way. The model sees the masked value and so does the caller. PII / PHI Detection, Secrets Detection and Regex Pattern Matching document no way to restore the original, and there is no gateway-level token vault or rehydration step in the docs as of September 2026.
There is exactly one documented round-trip, and it belongs to a third party. CrowdStrike AIDR supports format-preserving encryption: when Operation is Mutate and the input leg encrypted values with FPE, the gateway passes the resulting fpe_context back to AIDR on the output leg as input_fpe_context and calls POST {baseUrl}/v1/unredact, so the same keys are used end to end. A real reversible path — but it needs CrowdStrike, Mutate mode, and an fpe_context from the input leg.

So if your application needs the real value back, do not plan to get it from the gateway. Keep it in your application, send the masked version through, and re-join on an identifier you control. One related AIDR behaviour surprises people: once it raises a violation, Mutate does not soften it into a redaction — the request still 400s.
A worked example: an internal support copilot
A support copilot reads ticket text and looks up account records through an MCP server. Tickets routinely contain customer emails, phone numbers and internal employee IDs.
Decide what must never reach the provider. Emails and phone numbers: mask. Employee IDs: mask with a label, so the model can still tell two employees apart in one ticket. Credentials pasted in by mistake: mask on the way out too.
Pick modes and priorities. PII / PHI Detection on LLM Input, Mutate (its only mode), priority 1. Regex Pattern Matching on LLM Input, Mutate, priority 2, custom pattern EMP-\d{6} → [REDACTED EMPLOYEE ID]. Secrets Detection on LLM Output, Mutate. Start each in the loosest strategy it supports.
Decide the tool path deliberately. The account lookup needs the real email, so the masking rule targets the model, not the MCP server; Secrets Detection runs on MCP Post-Invoke instead, catching credentials that come back in a tool result. Write down why: the MCP server is inside your VPC and the model provider is not. That sentence is the actual control.

Attach it with a policy. Registering a guardrail does not apply it to traffic. Go to AI Gateway → Policies → Guardrails → Add Rule and scope it by targets (models, MCP servers, optionally per tool), subjects (users, teams, virtual accounts with IN / NOT IN), metadata and hooks. Matching rules are merged per hook.
Watch it before you trust it. Each guardrail is its own span in Monitor → Request Traces, with execution time, result, scope, input, output and mutations. Mutation rate near zero on a busy hook means your patterns are wrong; near 100% means you are about to learn what masking does to answer quality.

Gotchas worth knowing
Masking does not run on streamed output. Output guardrails are skipped entirely when "stream": true, because evaluating a response needs the complete text. Input guardrails are unaffected. Most chat interfaces stream by default, so output-side masking is inactive on the surface users actually see.
System prompts are never masked. The gateway strips them before sending content to any guardrail, so they are never inspected, blocked or redacted. The one exception is CrowdStrike AIDR, which sees the system prompt but never modifies it. Put a key in a system prompt and no masking guardrail will save you.
Log redaction is a different control from Mutate. TrueFoundry’s rule-based Logging Configuration redacts patterns from stored bodies with deny-wins semantics — one matching rule with Log off suppresses the body entirely. But the docs are explicit: “Redaction is applied only to the logged copy… The live request forwarded to the model provider is never modified.” It governs who inside your company reads prompts, not what leaves it. Our companion post on data loss prevention for LLM traffic covers egress.

Ordering can turn a block into a pass. DeepKeep documents this concretely: it applies first-listed precedence among the rails that fired, not the most severe action. “If PII is listed before Adversarial, a jailbreak that also trips PII is redacted and allowed instead of blocked.” That generalises to any mutate chain — put rails that must block ahead of rails that mask.
PII / PHI Detection is TrueFoundry-hosted only. It runs on Azure AI Language and is unavailable on self-hosted or hybrid (“Gateway Plane only”) deployments. Use Azure PII with your own key, or CrowdStrike.
Related reading
- Data Loss Prevention for LLM Traffic — the detection and egress side
- PII Redaction: Gateway vs Application — where the control should live
- TrueFoundry AI Gateway Guardrails Explained — the full guardrail model
- LLM Structured Outputs and JSON Schema — why masking and schemas collide
- Building an AI Governance Framework — the policy layer above
Conclusion
Data masking is a good control applied badly more often than it is applied well. The failure is rarely the masker — it is the assumptions around it: that masking round-trips, that the tool path is covered, that log redaction is the same thing, that the model reasons just as well over asterisks.
Four decisions are worth writing down first. Which entity types must genuinely not leave. Which of the four hooks each rule runs on. Whether something downstream needs the value back, because that answer alone decides masking versus tokenisation. And what you will accept in output quality, measured rather than assumed.
Then build it narrow: start in Audit where the guardrail allows it, read a week of traces, promote to Enforce But Ignore On Error, and keep full Enforce for routes where fail-closed is right. A gateway is the right place for this work — the one point every model, team and agent passes through. But a masking rule you have not watched is a rule you do not understand.
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
What is data masking in an AI gateway?
Replacing sensitive values in a prompt, response or tool payload with a substitute, inline in the request path, before the data reaches a model provider or external tool. In TrueFoundry this is the Mutate operation mode: the guardrail rewrites the matching spans and the rewritten version continues downstream. Mutate guardrails run by priority, lower first, and can still block.
What is the difference between data masking and redaction?
Masking overwrites characters while often preserving format — 4111-****-****-1234 keeps the shape of a card number — and redaction removes the value entirely, usually leaving a label like [REDACTED SSN]. Both are irreversible. Read the examples, not the noun: TrueFoundry’s PII guardrail produces length-matched asterisks, its Regex guardrail labelled tokens.
Can masked data be unmasked on the way back?
Not by default. TrueFoundry’s PII, Secrets and Regex guardrails are one-way — model and caller both see the masked value, and no restore step is documented. The single documented round-trip is CrowdStrike AIDR’s format-preserving encryption, where the gateway passes an fpe_context from the input leg to AIDR’s /v1/unredact endpoint on the output leg. If you need the real value downstream, hold it in your application.
Can I deploy TrueFoundry in my own VPC or on-prem?
Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.
Does TrueFoundry support MCP and AI agents generally?
Yes. It includes an MCP Gateway, an Agent Gateway, and an MCP & Agents Registry with tool-level access control. Agents on LangGraph, CrewAI, AutoGen, or a custom framework can all be governed centrally.
Does it integrate with my existing observability stack?
Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, Prometheus, or your preferred stack. It traces every request from prompt to tool and model execution, so you get unified logging without ripping out what you already run.










.png)
.png)
.png)
.png)
.png)


.webp)
.webp)


.webp)
.webp)
.webp)






