Blank white background with no objects or features visible.

Découvrez TrueForge : l'infrastructure d'agents open-source et indépendante des fournisseurs. Réduisez vos coûts de 50%. Explorer maintenant→

Agent Security Is a Systems Problem: From Prompt Injection to Runtime Control

Par Boyu Wang

Published: August 29, 2026

A major 2026 survey of LLM-agent security reaches a conclusion enterprise teams should take seriously: once a model can use tools, retain state, and act on behalf of someone else, security is no longer mainly a prompt-filtering problem. It becomes a systems problem spanning information flow, delegated authority, and persistent state.

Security Framework Notes and Key Takeaways
Source and independence note. This article is grounded in Yuchen Ling, Shengcheng Yu, Zhenyu Chen, and Chunrong Fang, Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation (arXiv:2606.10749, revised August 23, 2026), which synthesizes 247 papers. The mapping from the survey's security framework to TrueForge and TrueFoundry products is TrueFoundry editorial analysis; the paper does not evaluate or endorse TrueFoundry.
The infrastructure thesis: secure agents need more than model-level safety. The runtime should make trust boundaries and state transitions explicit; shared gateways should enforce identity, access, policy, credentials, budgets, guardrails, and evidence at model, tool, and agent boundaries; and systems of record should remain authoritative for business state and side effects.

Key Takeaways

  • Prompt injection is only one part of the problem. The survey finds tool-mediated control-flow hijacking remains prominent, while persistent-state corruption and multi-agent propagation are increasingly important.
  • Agent security has three coupled dimensions. Information enters the agent, authority lets it act, and persistent state lets an attack outlive the original interaction.
  • Controls need to compose across layers. A sandbox, guardrail, approval workflow, or gateway is useful, but none establishes security for the entire agent by itself.
  • TrueForge maps to the runtime boundary. It provides the execution loop, context management, tools, sandboxing, approvals, sessions, and events for agents built on it.
  • TrueFoundry Gateways map to shared control planes. AI Gateway governs routed model calls; MCP Gateway governs routed tool/data access; Agent Registry establishes identity, ownership, and access for registered agents, while broader Agent Gateway controls can add quotas, budgets, and traces for routed agent traffic.
  • Provenance-aware state remains an application responsibility. Persisting events is not the same as proving the integrity or trustworthiness of every piece of memory or business state.

1. Why Agent Security Is Different From LLM Safety

The survey starts from a simple architectural fact: an LLM agent does more than generate text. It may plan, invoke tools, browse, execute code, update state, and coordinate with other agents. Those capabilities turn model outputs into inputs to software control flow.

That changes the failure mode. A malicious instruction hidden in a web page can influence a later tool call. A poisoned tool result can become context for another decision. A compromised memory entry can survive the original session. A delegated credential can turn a bad plan into a real side effect. In a multi-agent system, the same contaminated information may propagate beyond the agent that first encountered it.

The survey therefore frames agent security around three interacting properties:

Three interacting security properties for AI agents: information flow, delegated authority, and persistent state.
Figure 1. TrueFoundry editorial synthesis of the survey's systems model. Risk compounds when untrusted information intersects with delegated authority and persistent state.

The useful implication is that secure-agent architecture should not ask only, “Did we block the malicious prompt?” It should also ask: what could that content influence, what authority could the resulting trajectory exercise, and what state could the trajectory leave behind?

2. The Attack Surface Runs Through the Whole Agent Lifecycle

Ling and coauthors use a lifecycle-based, systems-oriented framework rather than a flat list of attack names. Operationally, it is useful to separate the core action path—input, planning, decision, tool execution, and output—from cross-cutting surfaces such as memory, monitoring, and multi-agent coordination, because those surfaces can influence or observe multiple stages of the run.

Agent lifecycle from input through planning, decision, tool execution, and output, with memory, monitoring, and coordination shown as cross-cutting security surfaces.
Figure 2. The core lifecycle is Input → Planning → Decision → Tool Execution → Output. Memory, monitoring, and coordination are cross-cutting surfaces that can influence or observe multiple stages rather than simply occurring afterward.

This lifecycle perspective also explains why defenses can be “weakly compositional,” as the survey puts it. A content filter may block one class of malicious text but do nothing about over-privileged credentials. A sandbox may contain code execution but not prevent an authorized API call. An approval gate may stop one side effect but not a poisoned memory write. A trace may make an incident observable without preventing it.

Security therefore emerges from the composition of boundaries, privileges, state controls, and evidence.

3. Mapping the Survey to TrueForge and TrueFoundry

The cleanest product mapping is not “TrueFoundry solves agent security.” It is to ask which part of the survey's systems model each layer can realistically govern.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

INSCRIVEZ-VOUS
Table des matières

Gouvernez, déployez et suivez l'IA dans votre propre infrastructure

Réservez un séjour de 30 minutes avec notre Expert en IA

Réservez une démo

Le moyen le plus rapide de créer, de gérer et de faire évoluer votre IA

Démo du livre
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Découvrez-en plus

Aucun article n'a été trouvé.
August 29, 2026
|
5 min de lecture

Agent Security Is a Systems Problem: From Prompt Injection to Runtime Control

Aucun article n'a été trouvé.
August 28, 2026
|
5 min de lecture

From Agent Harness to System Intelligence: What Graph Engineering Changes in Production AI

Aucun article n'a été trouvé.
August 27, 2026
|
5 min de lecture

Wiring DeepKeep’s AI Firewall Into TrueFoundry AI Gateway as a Custom Guardrail

Aucun article n'a été trouvé.
August 27, 2026
|
5 min de lecture

What Is Vibe Coding? A Guide for Teams Shipping AI-Written Code

Aucun article n'a été trouvé.
Aucun article n'a été trouvé.

Blogs récents

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Faites un rapide tour d'horizon des produits
Commencer la visite guidée du produit
Visite guidée du produit