Blank white background with no objects or features visible.

TrueFoundry Named Frost & Sullivan's 2026 Global Transformational Innovation Leader. Read report

Data Masking in the AI Gateway: What Actually Works

By Ashish Dubey

Published: September 28, 2026

⚡ TL;DR
  • Data masking, redaction, tokenisation, encryption and anonymisation are different techniques with different reversibility. Most docs use the words interchangeably.
  • The question that decides your architecture is reversible or not. Tokenisation and format-preserving encryption round-trip; masking and redaction do not.
  • Masking LLM traffic breaks things database masking never had to: tool calls that need the real value, JSON schemas, multi-turn consistency, RAG retrieval.
  • In TrueFoundry, masking is the Mutate mode: it rewrites content, can still block, and runs sequentially by priority, lower first. Validate blocks without touching the data. Input validation runs in parallel with the model request; output and MCP pre/post-tool guardrails run synchronously.
  • There is no general unmasking on the response path. The only documented round-trip is CrowdStrike AIDR’s format-preserving encryption via /v1/unredact.

What data masking actually is

Data masking replaces a sensitive value with a substitute that is safe to expose, while keeping the surrounding data usable. It comes from the database world: give QA a copy of production where every card number is 4111-****-****-1234 and they can test realistic shapes without holding real PII. Point that at an LLM gateway and it changes — you are masking free text, on its way to a third party, inside a conversation that has to keep making sense.

These get used as synonyms constantly. They are not synonyms.

Technique What it does to the value Reversible Who can reverse it Format kept
Masking Overwrites characters with a fixed symbol, often leaving a partial view No Nobody Often
Redaction Removes the value, leaving a label like [REDACTED SSN] No Nobody No
Tokenisation Swaps in a surrogate; the pairing lives in a vault Yes Whoever holds the vault Usually
Format-preserving encryption Encrypts but keeps the shape — 16 digits stay 16 digits Yes Whoever holds the key Yes
Standard encryption Ciphertext of a different shape entirely Yes Whoever holds the key No
Pseudonymisation Consistent artificial identifiers (Patient-4471) With the mapping Whoever holds the mapping Partially
Anonymisation Irreversibly removes identifiability No Nobody, by definition No

Reversible vs irreversible is the distinction that matters. Tokenisation, FPE and encryption round-trip — someone holds a vault or a key, a new asset to protect. Masking, redaction and anonymisation are one-way: simpler, safer, and useless the moment something downstream needs the real value.

Pseudonymisation is not anonymisation. Under GDPR, pseudonymised data is still personal data and still in scope; only genuinely anonymised data falls outside the regulation (Recital 26). Replacing every name with a stable User-8812 is pseudonymisation, and calling it the latter will not survive review.

A third axis is easy to miss: consistency is separate from reversibility. A masker can be deterministic — the same input always yields the same placeholder — without being reversible, and determinism is what lets a model reason about “the same person” across a conversation.

What breaks when you mask LLM traffic

This is the part most write-ups skip, and it decides whether your design survives.

Model output quality degrades, unevenly. Turn Call our office at 312-555-1234 into Call our office at ************ and the model no longer knows that was a phone number — asterisks carry no type information. Labelled placeholders like [PHONE_NUMBER] keep the type, at the cost of telling anyone who sees the prompt what was there. Summarisation barely notices; extraction fails loudly; correlating two entities fails quietly, which is worse.

Tool calls need the real value. The sharpest failure. An agent that masks customer@acme.com on the way in cannot then call lookup_account(email=...). Masking the model path and leaving the tool path open creates the opposite hole: the model never saw the address, but the tool arguments carry it to a third-party MCP server. Decide, per tool, which is more trusted.

Structured output and JSON schemas break. If you are doing structured outputs with JSON Schema, masking can produce a value that no longer satisfies its own field: a format: email field cannot hold [REDACTED EMAIL], and an integer field cannot hold asterisks. Length-preserving masking is friendlier because it keeps the shape. Masking the output leg is worse — the model produced valid JSON, the masker rewrote a value inside it, and your parser throws on a response the model got right.

Multi-turn consistency. Turn one: “Priya Raman is the account owner.” Turn four: “does she still own it?” If turn one became asterisks, the model has no thread to follow. Deterministic pseudonymisation fixes this, but only if the masker keeps state across turns — which a stateless guardrail does not.

RAG retrieval against masked text. Mask documents before indexing and the embedding no longer contains the entity, so a query for it will not retrieve the chunk. Index unmasked and mask at retrieval time, and your vector store holds the unmasked corpus — the thing you were avoiding.

Applied narrowly, masking is worth it. Applied globally, it quietly makes the product worse and nobody connects the two.

Where teams get this wrong

Assuming masking round-trips. A team masks PII on the way to the model, assumes something unmasks it coming back, and builds on that. Unless you chose a reversible scheme and something holds keys, the caller sees the masked value too.

Masking the prompt and forgetting the tool path. Data moves in four places: into the model, out of the model, into a tool, out of a tool. Covering one is common. The leak then happens through a tool result nobody was inspecting.

Treating log redaction as a data-flow control. Redacting stored logs limits who inside your company reads prompts. The provider still received the original text.

Choosing irreversible masking for a reversible problem. If something downstream needs the value back, masking is the wrong technique. You need tokenisation or FPE, which means a vault or a key.

How masking works in TrueFoundry

Masking is not a separate product here. Every guardrail carries two settings — Operation Mode and Enforcement Strategy — masking is a value of the first.

Mutate vs Validate


Validate Mutate
What it does Blocks if something is wrong. Does not touch the data. Rewrites the data. Can also block.
Execution LLM Input validation can run in parallel with the model request. LLM Output and MCP pre/post-tool validation run synchronously. Sequentially by priority — lower first. Priority defaults to 1.
Use it for Hard stops: injection, policy violations, unsafe code. Masking, redaction, de-identification.

Priority is the whole ordering model for masking. Three mutate guardrails on one hook run in order, each seeing the output of the one before: a secrets masker at 1 and a regex masker at 2 means the regex guardrail matches text that already reads ***REDACTED***.

Input mutation is blocking and runs first — the model request cannot start until the prompt is final. Input validation then runs in the background alongside it, and if validation fails mid-flight the gateway cancels the model request so you do not pay for it. Output mutation and validation run synchronously after the response arrives, by which point the model cost is incurred.

TrueFoundry AI Gateway LLM request flow showing input guardrails before the model call and output guardrails after the response
TrueFoundry AI Gateway LLM request flow showing input guardrails before the model call and output guardrails after the response

MCP tool calls add two synchronous hooks: pre-invoke, where a failure means the tool never runs, and post-invoke, where a failure withholds the result. Guardrails run on every tool call separately — five tools, five sets of checks.

MCP tool invocation flow with pre-tool guardrails before execution and post-tool guardrails before the result reaches the model
MCP tool invocation flow with pre-tool guardrails before execution and post-tool guardrails before the result reaches the model

Enforce vs Enforce But Ignore On Error

The second setting separates two failures: the guardrail found something, versus the guardrail broke.

Strategy

On violation

On guardrail error

Enforce

Block

Block (fail-closed)

Enforce But Ignore On Error

Block

Let through (graceful degradation)

Audit

Let through, log only

Let through

The docs call Enforce But Ignore On Error “the safest default” and recommend the ladder Audit → Enforce But Ignore On Error → Enforce.

Guardrail enforcing strategy dropdown showing Enforce, Enforce But Ignore On Error and Audit
Guardrail enforcing strategy dropdown showing Enforce, Enforce But Ignore On Error and Audit

Which built-ins can actually mask

“We have a PII guardrail” tells you nothing about whether it can mask.

Guardrail Operation modes What the masked value becomes
PII / PHI Detection Mutate only — always redacts Asterisks matching the original length: 312-555-1234 → ************
Secrets Detection Validate or Mutate One configurable token, default ***REDACTED***
Regex Pattern Matching Validate or Mutate Per-pattern labels: [REDACTED SSN], [REDACTED CREDIT CARD]; custom default [REDACTED]
SQL Sanitizer Validate (default) or Mutate Comments stripped; in mutate mode DROP is logged, not blocked
Code Safety Linter Validate only n/a
Metadata Validation Validate only n/a

PII / PHI Detection is the odd one: it cannot block, only mask. If the requirement is “reject any request containing an SSN”, this is not the guardrail.

PII and PHI Detection guardrail configuration form with entity categories and enforcing strategy
PII and PHI Detection guardrail configuration form with entity categories and enforcing strategy

Its documented enforcing-strategy options are only enforce and enforce_but_ignore_on_error — audit is not listed, unlike Secrets, Regex and SQL Sanitizer. Audit-first rollout is not uniform across the built-ins. [VERIFY] whether that is a real restriction or a docs gap.

Regex Pattern Matching gives labelled placeholders rather than asterisks. It ships 46 presets — SSN, credit cards, passports for ten countries, cloud keys, GitHub and Slack tokens, IPs — plus custom patterns, such as EMP-\d{6} → [REDACTED EMPLOYEE ID].

Regex Pattern Matching guardrail form with preset patterns and custom pattern fields
Regex Pattern Matching guardrail form with preset patterns and custom pattern fields

Among the third-party guardrail integrations, mutate-capable ones include Google Model Armor, CrowdStrike AIDR, F5 AI Security, TrojAI, Enkrypt AI, Pillar Security, Lasso (classifix), HiddenLayer (/redact-*), DeepKeep and Noma — the last two require Mutate. Validate-only: Cisco AI Defense, Gray Swan Cygnal, Patronus, NVIDIA NeMo, Guardrails AI, Arthur AI. Asking a validate-only vendor to mask is a dead end.

Google Model Armor carries a setup trap: Mutate needs both an SDP inspect template and a de-identify template in Google Cloud, linked to the Model Armor template, with the transformation set to “Replace with infoType name”. Without both, mutate fails.

Google Cloud de-identify transformation rule set to replace detected values with the infoType name
Google Cloud de-identify transformation rule set to replace detected values with the infoType name

The unmasking question, answered plainly

TrueFoundry’s own masking guardrails are one-way. The model sees the masked value and so does the caller. PII / PHI Detection, Secrets Detection and Regex Pattern Matching document no way to restore the original, and there is no gateway-level token vault or rehydration step in the docs as of September 2026.

There is exactly one documented round-trip, and it belongs to a third party. CrowdStrike AIDR supports format-preserving encryption: when Operation is Mutate and the input leg encrypted values with FPE, the gateway passes the resulting fpe_context back to AIDR on the output leg as input_fpe_context and calls POST {baseUrl}/v1/unredact, so the same keys are used end to end. A real reversible path — but it needs CrowdStrike, Mutate mode, and an fpe_context from the input leg.

CrowdStrike AIDR guardrail configuration form in the TrueFoundry AI Gateway
CrowdStrike AIDR guardrail configuration form in the TrueFoundry AI Gateway

So if your application needs the real value back, do not plan to get it from the gateway. Keep it in your application, send the masked version through, and re-join on an identifier you control. One related AIDR behaviour surprises people: once it raises a violation, Mutate does not soften it into a redaction — the request still 400s.

A worked example: an internal support copilot

A support copilot reads ticket text and looks up account records through an MCP server. Tickets routinely contain customer emails, phone numbers and internal employee IDs.

Decide what must never reach the provider. Emails and phone numbers: mask. Employee IDs: mask with a label, so the model can still tell two employees apart in one ticket. Credentials pasted in by mistake: mask on the way out too.

Pick modes and priorities. PII / PHI Detection on LLM Input, Mutate (its only mode), priority 1. Regex Pattern Matching on LLM Input, Mutate, priority 2, custom pattern EMP-\d{6} → [REDACTED EMPLOYEE ID]. Secrets Detection on LLM Output, Mutate. Start each in the loosest strategy it supports.

Decide the tool path deliberately. The account lookup needs the real email, so the masking rule targets the model, not the MCP server; Secrets Detection runs on MCP Post-Invoke instead, catching credentials that come back in a tool result. Write down why: the MCP server is inside your VPC and the model provider is not. That sentence is the actual control.

Guardrail registry listing guardrails set to Mutate with the Enforce strategy
Guardrail registry listing guardrails set to Mutate with the Enforce strategy

Attach it with a policy. Registering a guardrail does not apply it to traffic. Go to AI Gateway → Policies → Guardrails → Add Rule and scope it by targets (models, MCP servers, optionally per tool), subjects (users, teams, virtual accounts with IN / NOT IN), metadata and hooks. Matching rules are merged per hook.

Watch it before you trust it. Each guardrail is its own span in Monitor → Request Traces, with execution time, result, scope, input, output and mutations. Mutation rate near zero on a busy hook means your patterns are wrong; near 100% means you are about to learn what masking does to answer quality.

Request Traces view with a guardrail span selected, showing latency, result and scope
Request Traces view with a guardrail span selected, showing latency, result and scope
Ready to try it on one route?
Register a PII guardrail, attach a policy to one model, and read the mutation spans.

Gotchas worth knowing

Masking does not run on streamed output. Output guardrails are skipped entirely when "stream": true, because evaluating a response needs the complete text. Input guardrails are unaffected. Most chat interfaces stream by default, so output-side masking is inactive on the surface users actually see.

System prompts are never masked. The gateway strips them before sending content to any guardrail, so they are never inspected, blocked or redacted. The one exception is CrowdStrike AIDR, which sees the system prompt but never modifies it. Put a key in a system prompt and no masking guardrail will save you.

Log redaction is a different control from Mutate. TrueFoundry’s rule-based Logging Configuration redacts patterns from stored bodies with deny-wins semantics — one matching rule with Log off suppresses the body entirely. But the docs are explicit: “Redaction is applied only to the logged copy… The live request forwarded to the model provider is never modified.” It governs who inside your company reads prompts, not what leaves it. Our companion post on data loss prevention for LLM traffic covers egress.

Logging configuration redaction settings applied to stored request and response bodies
Logging configuration redaction settings applied to stored request and response bodies

Ordering can turn a block into a pass. DeepKeep documents this concretely: it applies first-listed precedence among the rails that fired, not the most severe action. “If PII is listed before Adversarial, a jailbreak that also trips PII is redacted and allowed instead of blocked.” That generalises to any mutate chain — put rails that must block ahead of rails that mask.

PII / PHI Detection is TrueFoundry-hosted only. It runs on Azure AI Language and is unavailable on self-hosted or hybrid (“Gateway Plane only”) deployments. Use Azure PII with your own key, or CrowdStrike.

Related reading

Conclusion

Data masking is a good control applied badly more often than it is applied well. The failure is rarely the masker — it is the assumptions around it: that masking round-trips, that the tool path is covered, that log redaction is the same thing, that the model reasons just as well over asterisks.

Four decisions are worth writing down first. Which entity types must genuinely not leave. Which of the four hooks each rule runs on. Whether something downstream needs the value back, because that answer alone decides masking versus tokenisation. And what you will accept in output quality, measured rather than assumed.

Then build it narrow: start in Audit where the guardrail allows it, read a week of traces, promote to Enforce But Ignore On Error, and keep full Enforce for routes where fail-closed is right. A gateway is the right place for this work — the one point every model, team and agent passes through. But a masking rule you have not watched is a rule you do not understand.

Put masking on one route and read the traces

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Start free
Table of Contents

One Gateway for Every LLM, Agent and MCP Server

Book a 30-min with our AI expert

Book a Demo

The fastest way to build, govern and scale your AI

Book Demo
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Discover More

No items found.
September 28, 2026
|
5 min read

Langfuse Alternatives: 7 Options Compared on Licence, Price and Limits

No items found.
September 28, 2026
|
5 min read

Data Loss Prevention for LLM Traffic: Where It Has to Sit

No items found.
September 28, 2026
|
5 min read

Data Masking in the AI Gateway: What Actually Works

No items found.
September 28, 2026
|
5 min read

API Rate Limiting for LLMs: Count Tokens, Not Requests

No items found.
No items found.

Recent Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Frequently asked questions

What is data masking in an AI gateway?

Replacing sensitive values in a prompt, response or tool payload with a substitute, inline in the request path, before the data reaches a model provider or external tool. In TrueFoundry this is the Mutate operation mode: the guardrail rewrites the matching spans and the rewritten version continues downstream. Mutate guardrails run by priority, lower first, and can still block.

What is the difference between data masking and redaction?

Masking overwrites characters while often preserving format — 4111-****-****-1234 keeps the shape of a card number — and redaction removes the value entirely, usually leaving a label like [REDACTED SSN]. Both are irreversible. Read the examples, not the noun: TrueFoundry’s PII guardrail produces length-matched asterisks, its Regex guardrail labelled tokens.

Can masked data be unmasked on the way back?

Not by default. TrueFoundry’s PII, Secrets and Regex guardrails are one-way — model and caller both see the masked value, and no restore step is documented. The single documented round-trip is CrowdStrike AIDR’s format-preserving encryption, where the gateway passes an fpe_context from the input leg to AIDR’s /v1/unredact endpoint on the output leg. If you need the real value downstream, hold it in your application.

Can I deploy TrueFoundry in my own VPC or on-prem?

Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.

Does TrueFoundry support MCP and AI agents generally?

Yes. It includes an MCP Gateway, an Agent Gateway, and an MCP & Agents Registry with tool-level access control. Agents on LangGraph, CrewAI, AutoGen, or a custom framework can all be governed centrally.

Does it integrate with my existing observability stack?

Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, Prometheus, or your preferred stack. It traces every request from prompt to tool and model execution, so you get unified logging without ripping out what you already run.

Take a quick product tour
Start Product Tour
Product Tour