Blank white background with no objects or features visible.

TrueFoundry Named Frost & Sullivan's 2026 Global Transformational Innovation Leader. Read report

AI Red Teaming for Agents: Attacks, Campaigns, and Runtime Defense

By Ashish Dubey

Published: September 29, 2026

⚡ TL;DR
  • AI red teaming is adversarial testing of an agent as a system: you attack the prompt, the retrieved context, the tool surface and the budget — not the perimeter.
  • Two jobs get conflated constantly.Offline campaigns generate attacks and produce a report. >Runtime enforcement
  • Unlike a pen test, the attack surface is natural language, failure is probabilistic rather than binary, and a successful injection becomes a real action the moment the agent calls a tool.
  • Be clear what TrueFoundry is here. Its guardrails are runtime, inline enforcement across four hooks. TrueFoundry does not document offline campaign tooling, attack generation, an adversarial corpus, or scheduled scans.
  • Ordering beats coverage. Put PII before Adversarial Prompt Defense in DeepKeep and a jailbreak that also trips PII is redacted and allowed instead of blocked.

What AI red teaming actually means

Traditional red teaming borrows from military exercise: a team is chartered to attack your systems as an adversary would, then report what worked. AI red teaming keeps the charter and changes the target — you probe the model’s instruction-following, the trust boundary around retrieved content, and the actions the surrounding agent may take.

You are testing a system, not a model. A model that refuses a harmful request in isolation is a weak result if the agent around it will fetch a webpage, read hidden instructions and call a tool. The failures live in the seams.

The attack classes

Five families cover most of what a serious campaign produces. They do not share a control.

Attack class What the attacker does Control belongs
Jailbreak Role play, encoding, many-shot priming, obfuscation to defeat safety training Input classifier, model choice
Prompt injection (indirect) Instructions hidden in content the agent reads: a webpage, PDF, ticket, tool result Input + post-tool hooks, narrow tool scope
Data exfiltration Coaxing out secrets, PII or system prompts via the response, or a tool call to an attacker endpoint Output and post-tool hooks, egress control
Tool abuse / excessive agency Turning a legitimate tool into a weapon — DROP TABLE, a shell command, a payment call Pre-tool argument validation, allowlisting
Sponge attacks Inputs engineered to maximise compute — reasoning traps, huge contexts, non-converging loops Token rate limits, budgets

Indirect prompt injection changes the threat model. The canonical example is the 2025 Comet browser incident, cited in TrueFoundry’s agent-guardrails docs: a webpage carried hidden instructions written for the agent summarising it, and the agent followed them. No user tricked, no credential stolen — the content was the attack. Once an agent reads untrusted text, all of it is an untrusted instruction stream.

The OWASP Top 10 for LLM Applications mapped to gateway controls in the TrueFoundry documentation
The OWASP Top 10 for LLM Applications mapped to gateway controls in the TrueFoundry documentation

Sponge attacks get filed as billing and they are security — OWASP renamed “Model Denial of Service” to LLM10 Unbounded Consumption for exactly that reason. We mapped the full list in OWASP LLM Top 10: which risks a gateway actually fixes; four of the ten are request-path problems, and those four are what red teaming a live endpoint surfaces.

How agent red teaming differs from pen testing

Same discipline, different physics.


Traditional pen test AI red teaming
Target Network, host, app code, auth Prompt, context, tool surface, model behaviour
Outcome Binary — it works or it does not Probabilistic — works 3 times in 100, and the rate moves
Fix Patch, then closed Mitigate, then decay. New jailbreaks ship weekly
Scope Network and asset inventory Whatever the agent can reach — changes when someone adds an MCP server
Regression risk Low after the patch High. A model upgrade can reopen a whole class

Offline campaigns vs runtime enforcement

Two products, two problems.


Offline campaign Runtime enforcement (“LLM firewall”)
When Pre-launch, on a schedule, after a model change Every request, forever
Output A report: classes tried, success rates, examples A decision: allow, mutate, block — in milliseconds
Good at Finding failures you did not imagine Stopping known attacks at scale
Cannot Protect one production request Discover a novel attack class

Campaigns tell you what to enforce; enforcement makes the finding stop mattering. A campaign producing 40 pages and no change to the request path has not reduced risk. Enforcement without campaigns defends only against attacks someone already wrote a detector for. Run one pre-launch against the assembled system, on every model change including provider checkpoint updates, and on every tool-surface change — one new MCP server can expand blast radius more than a model swap.

Where teams get this wrong

Red teaming the model instead of the system. A team tests the base model, gets a clean result, then ships an agent with a database tool and a web fetcher. The model never failed; the system fails on the first page carrying hidden instructions.

Treating detection as prevention. Every injection detector is a classifier; classifiers have false negatives and attackers iterate cheaply. Assume injection occasionally succeeds and narrow what a compromised agent can do. Against a read-only agent that is an incident report; against an agent with a shell it is a breach.

Running the campaign once. A dated report describes a system that no longer exists: different checkpoint, new MCP servers, a changed system prompt. And guardrails rarely see the whole request — exclusions are platform-specific, so find yours before drawing the control matrix.

Want to see what your traffic is actually carrying?
Attach a guardrail group in Audit mode and read a week of traces before enforcing anything.

How this works in TrueFoundry

First, the boundary, plainly. TrueFoundry’s guardrails are runtime, inline enforcement: they validate or mutate content in the request path and return a decision. TrueFoundry does not ship or document offline red-team campaign tooling — no attack generation, no adversarial corpus, no scheduled scans, no red-team report. Across the AI Gateway guardrail docs “red teaming” appears essentially once, in OWASP LLM04 guidance recommending you “test your models against certain evaluation frameworks and red teaming” — an external recommendation, not a feature. DeepKeep is the partner page where red teaming surfaces as a capability, and it is DeepKeep’s, delivered vendor-side [VERIFY]. If you need campaigns, buy campaigns. What follows is the enforcement half.

Four hooks, and an execution order that saves you money

Guardrails attach to four hooks: LLM Input, LLM Output, MCP Tool Pre-Invoke and MCP Tool Post-Invoke.

TrueFoundry AI Gateway LLM request flow showing input guardrails before the model call and output guardrails after the response
TrueFoundry AI Gateway LLM request flow showing input guardrails before the model call and output guardrails after the response

The LLM path runs in seven documented steps, and step four is the interesting one. Input mutation runs first and blocks until done. Input validation then kicks off in the background, in parallel with the model request, so it does not sit on your time to first token — and if it fails mid-flight the gateway cancels the in-flight model request so you are not billed for it. Output mutation and validation run synchronously after the response arrives, so an output block still costs you the model call.

Validate vs Mutate, Enforce vs Enforce But Ignore On Error

Two independent settings per guardrail. Operation Mode is Validate or Mutate: Validate blocks without touching content, Mutate rewrites content and can still block, running by priority with the lower number first. Enforcement Strategy decides what happens on a violation, and separately on a guardrail error:

Strategy On violation On guardrail error
Enforce Block Block (fail closed)
Enforce But Ignore On Error Block Let through (graceful degradation)
Audit Let through, log Let through, log
Guardrail enforcing strategy dropdown in TrueFoundry showing Enforce, Enforce But Ignore On Error, and Audit
Guardrail enforcing strategy dropdown in TrueFoundry showing Enforce, Enforce But Ignore On Error, and Audit

Enforce But Ignore On Error is the fail-open/fail-closed switch, and the most consequential dropdown on the page: full Enforce means a vendor outage takes your application down with it. The documented ladder is Audit → Enforce But Ignore On Error → Enforce.

Built-ins and partner integrations

Ten guardrails are native, per the canonical truefoundry-guardrails page: Secrets Detection, Code Safety Linter, SQL Sanitizer, Regex Pattern Matching, Prompt Injection, PII Detection, Content Moderation, Metadata Validation, Cedar and OPA Guardrails. One doc inconsistency: the overview page lists nine, omitting Metadata Validation. Use the ten-item list.

Three run on managed Azure services — Content Moderation, PII/PHI Detection and Prompt Injection — and work only when TrueFoundry hosts the gateway, not on self-hosted or hybrid “Gateway Plane only” deployments. Bring-your-own-key alternatives are documented for each: Prisma AIRS for injection, CrowdStrike for PII, Bedrock Guardrails for moderation.

The adversarial-relevant partners, named from their own doc pages:

Partner Mode Adversarial-relevant detections
Gray Swan Cygnal Validate only violation score, violated_rules, ipi (indirect injection)
DeepKeep AI Firewall Mutate required Adversarial Prompt Defense, Credentials, PII, Toxic Language
HiddenLayer Validate or Mutate NONE / DETECT / REDACT / BLOCK
CrowdStrike AIDR Validate or Mutate Formerly Pangea. Only guardrail that sees the system prompt
Cisco AI Defense Validate only Prompt injection, code → SECURITY_VIOLATION
F5 AI Security (CalypsoAI) Validate or Mutate Prompt injection, jailbreak, PII, policy
TrojAI DEFEND Validate or Mutate PASS / BLOCK / REDACT / FLAG
Noma Security Mutate required Allow / Alert / Mask / Block
Lasso Security Validate or Mutate /classify hard stops, /classifix masking
Prisma AIRS, Enkrypt AI, Pillar Security, Google Model Armor Validate or Mutate Profile scanning, moderation, masking

That is thirteen security-focused partners. Adding the judge and moderation integrations — Patronus, NVIDIA NeMo, Guardrails AI, Arthur AI, Verra, Azure PII, Azure Prompt Shield, OpenAI Moderations, AWS Bedrock Guardrails — brings the roster to roughly twenty as of September 2026, plus a custom-guardrail contract for anything else.

Blocked request log in TrueFoundry showing a Gray Swan Cygnal guardrail violation
Blocked request log in TrueFoundry showing a Gray Swan Cygnal guardrail violation

Guardrails on MCP tool calls

TrueFoundry MCP tool call flow with pre-invoke guardrails on arguments and post-invoke guardrails on results
TrueFoundry MCP tool call flow with pre-invoke guardrails on arguments and post-invoke guardrails on results

Pre-invoke runs synchronously on the tool arguments; on failure the tool does not run. Post-invoke runs on the result; on failure the result is withheld from the model. Pre-invoke is where tool abuse dies — a SQL sanitizer reading the argument string, a code-safety linter on an execution tool. Post-invoke is where indirect injection dies, because a tool result is untrusted content.

Guardrails run on every tool call separately — five tools in a row means five full sets of checks. And X-TFY-GUARDRAILS-SCOPE, which narrows LLM guardrails to the last N messages, does not apply to MCP hooks. MCP tool poisoning and gateway defense covers the tool-description side.

Policies decide where a guardrail applies

Registering a guardrail does nothing to live traffic; a policy binds it. Path: AI Gateway → Policies → Guardrails → Add Rule, with four sections — When Request Goes To (models, MCP servers, optionally specific tools), From Subjects (users, teams, virtual accounts), With Metadata and Apply on Hooks.

TrueFoundry guardrail policies screen showing rules that bind guardrails to targets, subjects, and hooks
TrueFoundry guardrail policies screen showing rules that bind guardrails to targets, subjects, and hooks

All rules are evaluated on every request and matching ones are unioned per hook, so a permissive baseline and a strict production rule both apply. The X-TFY-GUARDRAILS header is the one bypass — it skips policies entirely, so treat it as privileged.

One scope limit as of September 2026: policy subjects are users, teams and virtual accounts — not agents. Policies keyed on agent identity, chain-aware conditions using delegation context, and hooks on the Agent Gateway itself are all documented as coming soon.

Metrics turn enforcement into evidence

The Metrics Dashboard’s guardrails view counts Requests evaluated, Mutated and Flagged (blocked), and charts requests per second by result — allowed, blocked, mutated, audit_mode_blocked — block and mutate rate per guardrail, and latency at P50 / P75 / P90 / P99, pivotable by Guardrails, Users, Virtual Accounts or Teams.

TrueFoundry guardrail metrics dashboard showing evaluated requests, block and mutate rates per guardrail, and latency percentiles
TrueFoundry guardrail metrics dashboard showing evaluated requests, block and mutate rates per guardrail, and latency percentiles

audit_mode_blocked is what makes an Audit-first rollout work: it counts what would have been blocked, so you can size the blast radius before enforcing.

A worked example: hardening an agent after a campaign

A campaign against a customer-support agent returns three findings: an indirect injection in a knowledge-base article makes the agent call an internal tool it should not; the agent echoes a customer phone number into a summary; one crafted question drives a twelve-tool loop.

Map findings to hooks. The documented baseline for agent traffic is the starting shape: PII redaction (Mutate) and prompt-injection detection (Validate) on input; secrets detection on output; SQL sanitizer and code-safety linter pre-tool; PII and secrets redaction post-tool. For the injection finding the load-bearing hook is post-tool, where the poisoned article enters the context.

Test in the Playground first. All four hooks are testable at AI Gateway → Playground; the Code button generates a snippet with X-TFY-GUARDRAILS already populated.

TrueFoundry Playground with LLM input guardrails selected for testing before deployment
TrueFoundry Playground with LLM input guardrails selected for testing before deployment

Bind a policy in Audit, scoped to this agent’s models and MCP servers, and leave it a week — a detector firing on 8% of legitimate tickets is one you cannot enforce yet. Then promote to Enforce But Ignore On Error and set a Custom Error Message, behind the Hide advanced fields toggle, supporting {{guardrail_message}} and {{failed_guardrails}}.

Verify in traces. Every guardrail is its own span in Monitor → Request Traces, with execution time, result, scope and mutations. Blocks return HTTP 400 with error.type = "guardrail_checks_failed".

TrueFoundry Request Traces view with a guardrail span showing execution time, result, and scope
TrueFoundry Request Traces view with a guardrail span showing execution time, result, and scope

Close the sponge finding somewhere else. The twelve-tool loop is a rate-limit and budget problem. No guardrail saves you from a loop that is well-formed on every iteration.

Ready to put enforcement behind your red-team findings?
Register a guardrail group, bind one policy, and watch it in traces.

Gotchas worth knowing

Rail ordering can convert a block into an allow. DeepKeep applies first-listed precedence among rails that fired — not the most severe action. Order PII Detector before Adversarial Prompt Defense and a jailbreak that also contains PII gets redacted and allowed rather than blocked: the redaction wins because it was listed first, and the jailbreak sails through with a phone number starred out. The documented fix is the recommended order — Credentials Leakage → PII → Adversarial → Toxic — and better still the two-firewall split: Adversarial Prompt Defense on the pre firewall only, credentials and PII on the post.

DeepKeep custom guardrail input configuration in TrueFoundry set to Mutate on the request path
DeepKeep custom guardrail input configuration in TrueFoundry set to Mutate on the request path

Output guardrails are skipped on streaming responses. With stream: true they do not run at all, because evaluating a response needs the complete text. Input guardrails always run. Since nearly every chat surface streams by default, this disables output-side defences where users see them.

System prompts are never inspected. The gateway strips them before sending content to any guardrail; CrowdStrike AIDR is the one documented exception, and it never modifies them. “We scan the whole prompt” is not a claim you can make.

If you write your own detector, return 2xx. The custom-guardrail contract is explicit: HTTP 2xx means the guardrail ran to completion, with the outcome in the body as verdict: false. Signalling a block with a 400 is the documented failure mode — the gateway may read it as an error and pass the request.

Related reading

Conclusion

Two halves that need each other. Campaigns find: episodic, expensive, the only way to discover a failure class nobody wrote a detector for. Enforcement holds: continuous, cheap per request, worthless against an attack nobody described to it. A programme with only the first produces reports; one with only the second is confident about the wrong things.

TrueFoundry sits squarely in the second half, and we would rather say so than imply otherwise. The gateway gives you one place to enforce the findings across every model, team and MCP tool call. Where it stops is at generating the attacks. If a vendor says their firewall makes red teaming unnecessary, ask which attack classes their detectors were trained on, and when — then run the campaign anyway, and enforce what it finds.

Put runtime enforcement behind your red-team findings

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Start free
Table of Contents

One Gateway for Every LLM, Agent and MCP Server

Book a 30-min with our AI expert

Book a Demo

The fastest way to build, govern and scale your AI

Book Demo
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Discover More

No items found.
September 29, 2026
|
5 min read

LLM as a Judge, Running Inline as a Gateway Guardrail

No items found.
September 29, 2026
|
5 min read

Langfuse vs LangSmith: Which LLM Observability Platform Fits

No items found.
September 29, 2026
|
5 min read

Datadog LLM Observability Pricing in 2026: What It Actually Costs

No items found.
September 29, 2026
|
5 min read

AI Red Teaming for Agents: Attacks, Campaigns, and Runtime Defense

No items found.
No items found.

Recent Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Frequently asked questions

What is AI red teaming?

Adversarial testing of an AI system — model, prompts, retrieved context and tools together — to find inputs that make it act against your intent. The classes are jailbreaks, direct and indirect prompt injection, data exfiltration, tool abuse, and resource-exhaustion “sponge” attacks. It differs from penetration testing because the attack surface is natural language, results are probabilistic, and a model upgrade can reopen a class you closed.

Is an LLM firewall the same as AI red teaming?

No. An LLM firewall is runtime enforcement: inline guardrails inspecting, rewriting or blocking on every request. Red teaming is offline discovery: campaigns that generate attacks and report what worked. Campaigns tell you what to enforce; enforcement makes the finding stop costing you.

Does TrueFoundry do AI red teaming?

Not in the campaign sense. It provides the runtime enforcement half: guardrails on four hooks, ten native built-ins, roughly twenty partner integrations, policy binding by model, MCP server, tool, user, team or virtual account, and per-guardrail metrics. It does not document attack generation, an adversarial corpus, scheduled scans or red-team reporting. Several partners are security vendors with campaign-side capability of their own.

Can I deploy TrueFoundry in my own VPC or on-prem?

Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.

Does TrueFoundry support MCP and AI agents generally?

Yes. It includes an MCP Gateway, an Agent Gateway, and an MCP & Agents Registry with tool-level access control. Agents on LangGraph, CrewAI, AutoGen, or a custom framework can all be governed centrally.

Does it integrate with my existing observability stack?

Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, Prometheus, or your preferred stack. It traces every request from prompt to tool and model execution, so you get unified logging without ripping out what you already run.

Take a quick product tour
Start Product Tour
Product Tour