Context Engineering: Designing What Your AI Agent Sees

Conçu pour la vitesse : latence d'environ 10 ms, même en cas de charge
Une méthode incroyablement rapide pour créer, suivre et déployer vos modèles !
- Gère plus de 350 RPS sur un seul processeur virtuel, aucun réglage n'est nécessaire
- Prêt pour la production avec un support complet pour les entreprises
What Is Context Engineering?
Context engineering is the discipline of designing the full set of information a model receives at each step of an agent's execution. In an agent runtime, context is everything the model sees on a given step, which includes the system instructions, the skills that are in scope, the definitions of available tools, the conversation history, and the results returned by tools.
That is a much larger surface than a single prompt. And it is dynamic: every tool call adds a result back into the context on the next step, so the information environment keeps changing as the agent works. Context engineering is how you keep that environment focused, relevant, and within budget instead of letting it fill with noise.
What lives in an agent's context
- System instructions. The standing guidance for how the agent should behave.
- Skills. Reusable procedures the agent can pull in for specific tasks, rather than stuffing every playbook into the system prompt.
- Tool definitions. The tools, often exposed over MCP, that the agent is allowed to call, and their schemas.
- Conversation history. The running dialogue and the agent's own intermediate steps.
- Tool results. The output of each call, which re-enters the context and shapes the next step.
Context Engineering vs Prompt Engineering
The two are related but operate at different scopes, and conflating them is why some teams plateau. Prompt engineering asks "how do I word this one request." Context engineering asks "what should the model have in front of it, on every step, to do this job well."
Neither replaces the other. You still word instructions carefully, but for an agent that is a small slice of a much bigger design problem. As context windows grow, the temptation is to pour everything in, and that is exactly the trap context engineering avoids: more context is not better context, and an overloaded window degrades recall and raises cost.
Techniques That Actually Move the Needle
A few practices carry most of the value when engineering context for agents.
- Move procedures into skills, not the system prompt. Anything that reads like a workflow or playbook belongs in a skill the agent loads when relevant, which keeps the base context lean. Load short, always-relevant guidance (style guides, safety policies) eagerly, and leave long or occasional procedures to load on demand.
- Scope tools tightly. The more tools you expose, the more tokens their definitions consume and the more ways the agent can go wrong. Give an agent the tools its function needs and no more.
- Manage history deliberately. Summarize or truncate old turns so the window stays focused on what matters now, rather than carrying every step forward verbatim.
- Treat tool results as untrusted context. Results re-enter the model's context on the next step, so they need the same inspection an input does, both for relevance and for safety.
How TrueFoundry Helps You Engineer Context
Context engineering is only sustainable if the pieces are managed, not hand-assembled per agent. TrueFoundry gives you control points for each part of the context.

In the Agent Harness, you configure the context when you build the agent: the system prompt, the skills in scope, and which MCP servers are available. That makes the context an explicit, reviewable configuration rather than something buried in code.
- Skills. Skills come from the central Skills Registry, so the procedures in an agent's context are versioned, access-controlled, and reused rather than copy-pasted. You can preload short, always-relevant skills and leave long ones to load on demand, which is the core lean-context technique made into a setting. See Claude Skills for how the skill format works.
- Tools. The MCP Gateway decides which tools an agent can reach and can scope access down to individual tools, so tool definitions in the context stay limited to what the agent actually needs.
- Tool-result safety. Because tool results re-enter context, AI agent guardrails inspect them on the post-tool hook for injections and sensitive data before they influence the next step.
- Observability. Every step is traced, so you can see exactly what was in the context when the agent made a decision, which is what makes context engineering iterative instead of guesswork.
Underneath all of it, the AI Gateway routes each call across 1,000+ models with roughly 3 to 4 ms of overhead, so the model best suited to a given step is a routing choice, not a rewrite. The result is that context becomes something you design and govern in one place, rather than a side effect of how each agent happened to be coded.
Conclusion
Context engineering is the shift from wording one prompt to designing everything a model sees across an agent's entire run. Get the instructions, skills, tools, history, and tool results right, and the agent stays focused and affordable; get them wrong, and no amount of prompt tuning saves it. TrueFoundry turns those pieces into managed configuration through the Agent Harness, Skills Registry, MCP Gateway, and full tracing, so context is something you design and govern rather than something that happens to you.
See how TrueFoundry gives you control over your agents' context from one control plane. Book a demo or start free.
TrueFoundry AI Gateway offre une latence d'environ 3 à 4 ms, gère plus de 350 RPS sur 1 processeur virtuel, évolue horizontalement facilement et est prête pour la production, tandis que LiteLM souffre d'une latence élevée, peine à dépasser un RPS modéré, ne dispose pas d'une mise à l'échelle intégrée et convient parfaitement aux charges de travail légères ou aux prototypes.



Gouvernez, déployez et suivez l'IA dans votre propre infrastructure
Blogs récents
Questions fréquemment posées
Qu'est-ce que l'ingénierie de contexte ?
L'ingénierie de contexte est la pratique qui consiste à concevoir tout ce qu'un modèle voit à chaque étape de l'exécution d'un agent : instructions système, skills, définitions d'outils, historique de la conversation et résultats des outils. Elle va au-delà de la formulation d'un seul prompt et façonne tout l'environnement d'information dans lequel l'agent opère, afin qu'il reste concentré, pertinent et dans les limites de son budget de contexte.
Quelle est la différence entre l'ingénierie de contexte et l'ingénierie de prompts ?
L'ingénierie de prompts ajuste un seul message : sa formulation, ses exemples et son format. L'ingénierie de contexte conçoit l'ensemble du contexte qu'un agent voit à chaque étape, y compris les skills et les outils qui entrent dans son périmètre et la manière dont l'historique et les résultats des outils sont gérés. L'ingénierie de prompts corrige une réponse faible ; l'ingénierie de contexte corrige la dérive, la prolifération des outils et l'inflation du contexte dans les agents à plusieurs étapes.
Comment pratiquer l'ingénierie de contexte pour les agents d'IA ?
Gardez le contexte de base léger en déplaçant les procédures dans des skills que l'agent charge à la demande, délimitez strictement les outils pour que seules les définitions nécessaires soient présentes, gérez l'historique de la conversation afin que les anciens tours n'encombrent pas la fenêtre, et traitez les résultats des outils comme un contexte non fiable qui doit être inspecté. C'est le fait de gérer tout cela comme de la configuration plutôt que comme du code qui rend la démarche reproductible.
Une fenêtre de contexte plus grande dispense-t-elle de l'ingénierie de contexte ?
Non. Une fenêtre plus grande incite à tout inclure, mais un contexte surchargé dégrade le rappel et augmente les coûts. L'ingénierie de contexte consiste à placer la bonne information devant le modèle, ce qui compte davantage, et non moins, à mesure que les fenêtres s'agrandissent.
TrueFoundry prend-il en charge MCP et les skills pour gérer le contexte ?
Oui. L'Agent Harness vous permet de configurer les instructions système, les skills issues du Skills Registry central et les serveurs MCP qu'un agent peut atteindre, et la MCP Gateway délimite l'accès aux outils : les éléments du contexte d'un agent sont ainsi gouvernés de manière centralisée.
Puis-je l'exécuter dans mon propre VPC ?
Oui. TrueFoundry s'exécute dans votre VPC, on-premise, en environnement air-gapped ou en hybride, de sorte que le trafic vers Gemini 3 Pro et tous les autres modèles reste gouverné dans votre propre domaine.













.webp)


.webp)
.webp)
.webp)


.png)
.png)
.png)
.png)
.png)






.png)







