AI Red Teaming for Agents: Attacks, Campaigns, and Runtime Defense

Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU â no tuning needed
- Production-ready with full enterprise support
What AI red teaming actually means
Traditional red teaming borrows from military exercise: a team is chartered to attack your systems as an adversary would, then report what worked. AI red teaming keeps the charter and changes the target â you probe the modelâs instruction-following, the trust boundary around retrieved content, and the actions the surrounding agent may take.
You are testing a system, not a model. A model that refuses a harmful request in isolation is a weak result if the agent around it will fetch a webpage, read hidden instructions and call a tool. The failures live in the seams.
The attack classes
Five families cover most of what a serious campaign produces. They do not share a control.
Indirect prompt injection changes the threat model. The canonical example is the 2025 Comet browser incident, cited in TrueFoundryâs agent-guardrails docs: a webpage carried hidden instructions written for the agent summarising it, and the agent followed them. No user tricked, no credential stolen â the content was the attack. Once an agent reads untrusted text, all of it is an untrusted instruction stream.

Sponge attacks get filed as billing and they are security â OWASP renamed âModel Denial of Serviceâ to LLM10 Unbounded Consumption for exactly that reason. We mapped the full list in OWASP LLM Top 10: which risks a gateway actually fixes; four of the ten are request-path problems, and those four are what red teaming a live endpoint surfaces.
How agent red teaming differs from pen testing
Same discipline, different physics.
Offline campaigns vs runtime enforcement
Two products, two problems.
Campaigns tell you what to enforce; enforcement makes the finding stop mattering. A campaign producing 40 pages and no change to the request path has not reduced risk. Enforcement without campaigns defends only against attacks someone already wrote a detector for. Run one pre-launch against the assembled system, on every model change including provider checkpoint updates, and on every tool-surface change â one new MCP server can expand blast radius more than a model swap.
Where teams get this wrong
Red teaming the model instead of the system. A team tests the base model, gets a clean result, then ships an agent with a database tool and a web fetcher. The model never failed; the system fails on the first page carrying hidden instructions.
Treating detection as prevention. Every injection detector is a classifier; classifiers have false negatives and attackers iterate cheaply. Assume injection occasionally succeeds and narrow what a compromised agent can do. Against a read-only agent that is an incident report; against an agent with a shell it is a breach.
Running the campaign once. A dated report describes a system that no longer exists: different checkpoint, new MCP servers, a changed system prompt. And guardrails rarely see the whole request â exclusions are platform-specific, so find yours before drawing the control matrix.
How this works in TrueFoundry
First, the boundary, plainly. TrueFoundryâs guardrails are runtime, inline enforcement: they validate or mutate content in the request path and return a decision. TrueFoundry does not ship or document offline red-team campaign tooling â no attack generation, no adversarial corpus, no scheduled scans, no red-team report. Across the AI Gateway guardrail docs âred teamingâ appears essentially once, in OWASP LLM04 guidance recommending you âtest your models against certain evaluation frameworks and red teamingâ â an external recommendation, not a feature. DeepKeep is the partner page where red teaming surfaces as a capability, and it is DeepKeepâs, delivered vendor-side [VERIFY]. If you need campaigns, buy campaigns. What follows is the enforcement half.
Four hooks, and an execution order that saves you money
Guardrails attach to four hooks: LLM Input, LLM Output, MCP Tool Pre-Invoke and MCP Tool Post-Invoke.

The LLM path runs in seven documented steps, and step four is the interesting one. Input mutation runs first and blocks until done. Input validation then kicks off in the background, in parallel with the model request, so it does not sit on your time to first token â and if it fails mid-flight the gateway cancels the in-flight model request so you are not billed for it. Output mutation and validation run synchronously after the response arrives, so an output block still costs you the model call.
Validate vs Mutate, Enforce vs Enforce But Ignore On Error
Two independent settings per guardrail. Operation Mode is Validate or Mutate: Validate blocks without touching content, Mutate rewrites content and can still block, running by priority with the lower number first. Enforcement Strategy decides what happens on a violation, and separately on a guardrail error:

Enforce But Ignore On Error is the fail-open/fail-closed switch, and the most consequential dropdown on the page: full Enforce means a vendor outage takes your application down with it. The documented ladder is Audit â Enforce But Ignore On Error â Enforce.
Built-ins and partner integrations
Ten guardrails are native, per the canonical truefoundry-guardrails page: Secrets Detection, Code Safety Linter, SQL Sanitizer, Regex Pattern Matching, Prompt Injection, PII Detection, Content Moderation, Metadata Validation, Cedar and OPA Guardrails. One doc inconsistency: the overview page lists nine, omitting Metadata Validation. Use the ten-item list.
Three run on managed Azure services â Content Moderation, PII/PHI Detection and Prompt Injection â and work only when TrueFoundry hosts the gateway, not on self-hosted or hybrid âGateway Plane onlyâ deployments. Bring-your-own-key alternatives are documented for each: Prisma AIRS for injection, CrowdStrike for PII, Bedrock Guardrails for moderation.
The adversarial-relevant partners, named from their own doc pages:
That is thirteen security-focused partners. Adding the judge and moderation integrations â Patronus, NVIDIA NeMo, Guardrails AI, Arthur AI, Verra, Azure PII, Azure Prompt Shield, OpenAI Moderations, AWS Bedrock Guardrails â brings the roster to roughly twenty as of September 2026, plus a custom-guardrail contract for anything else.

Guardrails on MCP tool calls

Pre-invoke runs synchronously on the tool arguments; on failure the tool does not run. Post-invoke runs on the result; on failure the result is withheld from the model. Pre-invoke is where tool abuse dies â a SQL sanitizer reading the argument string, a code-safety linter on an execution tool. Post-invoke is where indirect injection dies, because a tool result is untrusted content.
Guardrails run on every tool call separately â five tools in a row means five full sets of checks. And X-TFY-GUARDRAILS-SCOPE, which narrows LLM guardrails to the last N messages, does not apply to MCP hooks. MCP tool poisoning and gateway defense covers the tool-description side.
Policies decide where a guardrail applies
Registering a guardrail does nothing to live traffic; a policy binds it. Path: AI Gateway â Policies â Guardrails â Add Rule, with four sections â When Request Goes To (models, MCP servers, optionally specific tools), From Subjects (users, teams, virtual accounts), With Metadata and Apply on Hooks.

All rules are evaluated on every request and matching ones are unioned per hook, so a permissive baseline and a strict production rule both apply. The X-TFY-GUARDRAILS header is the one bypass â it skips policies entirely, so treat it as privileged.
One scope limit as of September 2026: policy subjects are users, teams and virtual accounts â not agents. Policies keyed on agent identity, chain-aware conditions using delegation context, and hooks on the Agent Gateway itself are all documented as coming soon.
Metrics turn enforcement into evidence
The Metrics Dashboardâs guardrails view counts Requests evaluated, Mutated and Flagged (blocked), and charts requests per second by result â allowed, blocked, mutated, audit_mode_blocked â block and mutate rate per guardrail, and latency at P50 / P75 / P90 / P99, pivotable by Guardrails, Users, Virtual Accounts or Teams.

audit_mode_blocked is what makes an Audit-first rollout work: it counts what would have been blocked, so you can size the blast radius before enforcing.
A worked example: hardening an agent after a campaign
A campaign against a customer-support agent returns three findings: an indirect injection in a knowledge-base article makes the agent call an internal tool it should not; the agent echoes a customer phone number into a summary; one crafted question drives a twelve-tool loop.
Map findings to hooks. The documented baseline for agent traffic is the starting shape: PII redaction (Mutate) and prompt-injection detection (Validate) on input; secrets detection on output; SQL sanitizer and code-safety linter pre-tool; PII and secrets redaction post-tool. For the injection finding the load-bearing hook is post-tool, where the poisoned article enters the context.
Test in the Playground first. All four hooks are testable at AI Gateway â Playground; the Code button generates a snippet with X-TFY-GUARDRAILS already populated.

Bind a policy in Audit, scoped to this agentâs models and MCP servers, and leave it a week â a detector firing on 8% of legitimate tickets is one you cannot enforce yet. Then promote to Enforce But Ignore On Error and set a Custom Error Message, behind the Hide advanced fields toggle, supporting {{guardrail_message}} and {{failed_guardrails}}.
Verify in traces. Every guardrail is its own span in Monitor â Request Traces, with execution time, result, scope and mutations. Blocks return HTTP 400 with error.type = "guardrail_checks_failed".

Close the sponge finding somewhere else. The twelve-tool loop is a rate-limit and budget problem. No guardrail saves you from a loop that is well-formed on every iteration.
Gotchas worth knowing
Rail ordering can convert a block into an allow. DeepKeep applies first-listed precedence among rails that fired â not the most severe action. Order PII Detector before Adversarial Prompt Defense and a jailbreak that also contains PII gets redacted and allowed rather than blocked: the redaction wins because it was listed first, and the jailbreak sails through with a phone number starred out. The documented fix is the recommended order â Credentials Leakage â PII â Adversarial â Toxic â and better still the two-firewall split: Adversarial Prompt Defense on the pre firewall only, credentials and PII on the post.

Output guardrails are skipped on streaming responses. With stream: true they do not run at all, because evaluating a response needs the complete text. Input guardrails always run. Since nearly every chat surface streams by default, this disables output-side defences where users see them.
System prompts are never inspected. The gateway strips them before sending content to any guardrail; CrowdStrike AIDR is the one documented exception, and it never modifies them. âWe scan the whole promptâ is not a claim you can make.
If you write your own detector, return 2xx. The custom-guardrail contract is explicit: HTTP 2xx means the guardrail ran to completion, with the outcome in the body as verdict: false. Signalling a block with a 400 is the documented failure mode â the gateway may read it as an error and pass the request.
Related reading
- OWASP LLM Top 10 â the taxonomy your findings map to
- Prompt injection defense in the LLM gateway â the highest-yield attack class
- TrueFoundry AI Gateway guardrails explained â the configuration walkthrough
- Benchmarking LLM guardrail providers â how the detectors compare
- Agentic AI security â the wider control model
Conclusion
Two halves that need each other. Campaigns find: episodic, expensive, the only way to discover a failure class nobody wrote a detector for. Enforcement holds: continuous, cheap per request, worthless against an attack nobody described to it. A programme with only the first produces reports; one with only the second is confident about the wrong things.
TrueFoundry sits squarely in the second half, and we would rather say so than imply otherwise. The gateway gives you one place to enforce the findings across every model, team and MCP tool call. Where it stops is at generating the attacks. If a vendor says their firewall makes red teaming unnecessary, ask which attack classes their detectors were trained on, and when â then run the campaign anyway, and enforce what it finds.
TrueFoundry AI Gateway delivers ~3â4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
What is AI red teaming?
Adversarial testing of an AI system â model, prompts, retrieved context and tools together â to find inputs that make it act against your intent. The classes are jailbreaks, direct and indirect prompt injection, data exfiltration, tool abuse, and resource-exhaustion âspongeâ attacks. It differs from penetration testing because the attack surface is natural language, results are probabilistic, and a model upgrade can reopen a class you closed.
Is an LLM firewall the same as AI red teaming?
No. An LLM firewall is runtime enforcement: inline guardrails inspecting, rewriting or blocking on every request. Red teaming is offline discovery: campaigns that generate attacks and report what worked. Campaigns tell you what to enforce; enforcement makes the finding stop costing you.
Does TrueFoundry do AI red teaming?
Not in the campaign sense. It provides the runtime enforcement half: guardrails on four hooks, ten native built-ins, roughly twenty partner integrations, policy binding by model, MCP server, tool, user, team or virtual account, and per-guardrail metrics. It does not document attack generation, an adversarial corpus, scheduled scans or red-team reporting. Several partners are security vendors with campaign-side capability of their own.
Can I deploy TrueFoundry in my own VPC or on-prem?
Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.
Does TrueFoundry support MCP and AI agents generally?
Yes. It includes an MCP Gateway, an Agent Gateway, and an MCP & Agents Registry with tool-level access control. Agents on LangGraph, CrewAI, AutoGen, or a custom framework can all be governed centrally.
Does it integrate with my existing observability stack?
Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, Prometheus, or your preferred stack. It traces every request from prompt to tool and model execution, so you get unified logging without ripping out what you already run.










.png)
.png)
.png)
.png)
.png)
.png)
.png)
.png)
.png)
.png)


.webp)
.webp)






