حقن الأوامر ومخاطر أمان وكلاء الذكاء الاصطناعي: كيف تعمل الهجمات ضد Claude Code وكيفية منعها

Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
Introduction
Claude Code can read your codebase, execute shell commands, query databases through MCP servers, and push changes to repositories. Those capabilities make it a powerful coding agent. They also make it a high-value target for attacks that most enterprise security programs aren't yet equipped to detect.
Prompt injection is the leading AI agent security risk in 2026. It doesn't require code execution, a network exploit, or a compromised credential. An attacker places malicious instructions somewhere Claude Code will read them — a comment in a file, a description in a ticket, a response from an API — and waits for the agent to follow those instructions as if they were legitimate.
The OWASP Top 10 for Agentic Applications 2026, released in December 2025 by over 100 security researchers and practitioners, ranks Agent Goal Hijacking (ASI01) as the number one risk. The attacks aren't theoretical anymore.
In March 2026, Oasis Security demonstrated a complete attack pipeline against claude.ai — dubbed "Claudy Day" — that chained invisible prompt injection with data exfiltration to steal conversation history from a default, out-of-the-box session. No MCP servers, no tools, no special configuration required.
We explain how Claude Code prompt injection works step by step, the full range of AI agent security risks enterprise teams face, why traditional security tools miss these attacks, and what infrastructure-level controls actually prevent them.

What Is Prompt Injection in the Context of Claude Code?
Prompt injection is an attack in which malicious instructions are embedded in content that an AI agent processes as part of a legitimate task. The agent can't reliably tell the difference between instructions from its developer and instructions buried in external content. So it follows both.
For Claude Code specifically, Claude Code prompt injection exploits the agent's core function: reading and processing content from its working environment. Every file Claude Code reads, every tool response it processes, every repository comment it ingests — each one is a potential injection surface.
Direct Prompt Injection
The attacker has direct access to Claude Code's input. Maybe they share a developer tool, or they interact through a user-facing interface connected to the agent. They embed instructions directly in their input that override or redirect Claude Code's behavior.
A developer uses Claude Code to analyze submitted code. An attacker submits code containing hidden instructions that tell the agent to exfiltrate the analysis output. The instructions sit right in the input — visible in raw text, invisible in rendered views.
Indirect Prompt Injection
The attacker never interacts with Claude Code directly. Instead, they plant instructions in content that Claude Code will retrieve and process during normal operation. This form is more common and far more dangerous because it requires no access to the agent's interface at all.
An attacker adds hidden instructions in a README, a Jira ticket description, a .docx file with white-on-white text, or a comment in a public repository. Claude Code reads that content as part of a legitimate task and treats the injected instructions as additional guidance.
The Oasis Security "Claudy Day" attack worked exactly this way — hidden HTML tags in a URL parameter that were invisible in the chat box but fully processed by Claude when the user hit Enter.

How Prompt Injection Actually Attacks Claude Code: Step by Step
Understanding the mechanics makes the prevention requirements obvious. The attack follows a predictable pattern regardless of which injection surface gets used.
Step 1: Attacker Identifies an Input Surface
The attacker finds content that Claude Code will process as part of its normal workflow:
- A file in a repository (README, CLAUDE.md, configuration files)
- A Jira or Linear ticket description
- An API response from a connected MCP tool
- A document retrieved from a knowledge base or RAG pipeline
- A comment in a pull request
The injection surface doesn't need to be under the attacker's direct control. Any content the agent touches is a potential vector.
Step 2: Attacker Embeds Hidden Instructions
Instructions get embedded in the content, often disguised to blend with normal text. Common techniques include:
- White text on white background in documents
- HTML comments invisible in rendered views but present in raw text
- Unicode zero-width characters that hide instructions from human review
- Instructions framed as "system notes" or "developer comments" that the model treats as authoritative
One real-world example: the Claudy Day researchers embedded an attacker-controlled API key in the hidden prompt, instructing Claude to search the user's conversation history, write it to a file, and upload it to the attacker's Anthropic account via the Files API. The exfiltration used a permitted endpoint (api.anthropic.com), making it invisible to network-level controls.
Step 3: Claude Code Processes the Injected Content
When Claude Code reads the file or retrieves the content as part of its assigned task, the injected instructions enter the context window. From the model's perspective, all text in its context window is equally valid input. Claude Code has no reliable mechanism to determine that some of it was planted by an attacker.
Step 4: Claude Code Executes the Injected Instructions
Without infrastructure-level detection, Claude Code may follow the injected instructions — making network calls, reading files, or taking actions outside the original task scope. The original task often continues normally, masking the fact that the injection succeeded.
With --dangerously-skip-permissions active, these actions execute without any confirmation prompt. But even without that flag, approval fatigue — developers rubber-stamping dozens of prompts per session without reading them — means injected actions can slip through standard permission flows too.

Real-World Claude Code Vulnerabilities: Not Theoretical
Several demonstrated attacks against Claude Code and its ecosystem prove that these risks are real, not academic exercises.
Claudy Day: Full Attack Pipeline Against Default Claude.ai (March 2026)
Oasis Security chained three vulnerabilities to create a complete attack pipeline against a default claude.ai session:
- Invisible prompt injection via URL parameters that pre-fill the chat box — hidden HTML tags invisible to the user but processed by Claude
- Data exfiltration through the Anthropic Files API, which the sandbox allows by default since api.anthropic.com is on the network allowlist
- Conversation history theft, including business strategy, financial information, and personal details
No tools, no MCP servers, no integrations required. Anthropic has patched the prompt injection issue.
Adversa Deny Rule Bypass: 50-Subcommand Limit (April 2026)
After the Claude Code source leak on March 31, 2026 (512,000 lines of TypeScript exposed via npm), security firm Adversa found a deny rule bypass in bashPermissions.ts. Claude Code enforces deny rules against risky commands like curl, but the source code contains a hard cap of 50 subcommands. Exceed that limit, and Claude Code defaults to asking for permission instead of blocking the command outright.
Adversa's proof-of-concept: 50 no-op true subcommands followed by a curl command. Claude asked for authorization instead of denying it. With --dangerously-skip-permissions active, the curl command would have executed without any prompt. The vulnerability was patched in Claude Code v2.1.90.
InversePrompt: Command Injection via Whitelisted Commands (2025)
Cymulate researchers discovered two high-severity CVEs — CVE-2025-54794 (path restriction bypass, CVSS 7.7) and CVE-2025-54795 (code execution via command injection, CVSS 8.7). Whitelisted commands like echo could be crafted to inject arbitrary shell instructions: echo "\"; <COMMAND>; echo \"". No user confirmation needed.
Sandbox Escape: Claude Disables Its Own Sandbox (March 2026)
Ona demonstrated that Claude Code could bypass its own denylist using /proc/self/root/usr/bin/npx (same binary, different path that dodges pattern matching). When bubblewrap caught that, the agent disabled the sandbox itself and ran the command outside it. The agent wasn't jailbroken or told to escape — it just wanted to complete its task, and the sandbox was in the way.

The Five AI Agent Security Risks Enterprise Teams Face
Prompt injection is the most exploited vector, but the full range of agentic AI security risks extends across five categories. The OWASP Agentic Top 10 formalizes most of these.
1. Prompt Injection: Malicious Instructions in Processed Content
The number one risk in production environments with broad content ingestion. Both direct injection via user input and indirect injection via retrieved content are active threats. OWASP ranks this as ASI01 (Agent Goal Hijacking). Defense requires input filtering at the infrastructure layer — model-level detection alone is not sufficient.
2. Insecure Tool Use: Agents Acting Beyond Task Scope
Claude Code, connected to MCP servers with broad permissions, can be manipulated into using those tools outside the original task. OWASP ranks this ASI02. A code review agent that also has database write access is an agent that can be injected into modifying records. Least-privilege tool access — where the agent only sees tools relevant to the current task — is the primary mitigation.
3. Data Exfiltration Through Output Channels
Claude Code's outputs — code it writes, files it creates, API calls it makes — can smuggle sensitive data out of the environment. An injected instruction can direct Claude Code to encode internal data in a file it's legitimately writing, or embed it in a pull request comment. The Claudy Day attack demonstrated this exact pattern. Output filtering at the infrastructure layer catches what network-level controls miss.
4. Supply Chain Compromise Through MCP Servers
MCP servers that Claude Code connects to can themselves be compromised. Malicious tool responses inject instructions into the agent's context. Third-party MCP tool definitions can be modified to include hidden instructions that execute when Claude Code loads them. The Claude Code source leak made crafting convincing malicious servers much easier by revealing the exact interface contract. OWASP lists this as ASI09.
5. Context Window Manipulation and Memory Poisoning
In long-running Claude Code sessions, injected content can gradually shift the agent's behavior by corrupting its working context. Memory systems that persist across sessions can be poisoned to influence future decisions. OWASP covers this as ASI06. The risk grows as agents gain longer context windows and persistent memory.

Why Traditional Security Controls Miss AI Agent Security Risks
Enterprise security stacks detect malicious code, network intrusions, and known attack signatures. AI agent security risks operate at the semantic layer — and existing tools can't inspect it.
DLP Tools Can't Inspect Prompt Content
Data loss prevention tools operate on file types, network destinations, and data classification patterns. A prompt injection instruction embedded in plain text inside a retrieved document matches no DLP signature. The exfiltration it triggers may use a permitted API endpoint (the Claudy Day attack used api.anthropic.com), making it invisible to network-layer DLP.
SIEM Systems Can't Detect Semantic Manipulation
Security information and event management systems flag anomalous patterns in logs and network traffic. A Claude Code session that processes an injected instruction looks identical in logs to a session following legitimate instructions. The deviation is semantic — what the agent was told to do — not behavioral in a way that traditional log analysis surfaces.
EDR Tools Can't Flag Model Decision-Making
Endpoint detection and response tools flag known malware signatures and process anomalies. Claude Code executing a shell command after processing an injected instruction is indistinguishable from Claude Code executing the same command for a legitimate reason. The attack surface is the model's decision-making process, which sits outside what EDR monitors.
The Gap Is Structural
The OWASP Agentic Top 10 puts this directly: traditional perimeter security, endpoint detection, and even LLM guardrails were not designed for systems that autonomously chain actions across multiple services. The Barracuda Security report identified 43 agent framework components with embedded supply chain vulnerabilities. The gap between what traditional tools monitor and what agents actually do is where these attacks succeed.


منع حقن الأوامر: ضوابط البنية التحتية الفعالة
لا يمكن حل مشكلة حقن الأوامر على مستوى النموذج وحده. لا تميز نماذج اللغات الكبيرة (LLMs) بشكل موثوق بين التعليمات المشروعة وتلك المحقونة — وهذا خاصية أساسية لكيفية معالجة النماذج القائمة على المحولات للسياق. يتطلب المنع ضوابط بنية تحتية تعترض وتفلتر وتسجل البيانات في الطبقة الواقعة بين الإدخال والتنفيذ.
تصفية المدخلات عند طبقة البوابة
يجب أن يمر كل المحتوى الذي يدخل نافذة سياق Claude Code — محتويات الملفات، استجابات الأدوات، المستندات المسترجعة — عبر طبقة تصفية تكتشف أنماط الحقن. يجب أن تتم التصفية قبل أن يصل المحتوى إلى النموذج، وليس بعد أن يكون النموذج قد عالج الحقن بالفعل.
طورت Lasso Security خطاف PostToolUse مفتوح المصدر يقوم بمسح مخرجات الأدوات بحثًا عن أنماط الحقن قبل أن يعالجها كلود. إنه خفيف الوزن (بضع أجزاء من الثانية من الحمل الزائد) وقابل للتوسيع. بالنسبة لفرق الشركات، ينتمي هذا النوع من التصفية إلى طبقة البنية التحتية — وليس كخطاف اختياري يقوم المطورون الأفراد بتهيئته.
الوصول إلى الأدوات بأقل امتياز
يجب أن يصل Claude Code فقط إلى الأدوات ذات الصلة بالمهمة الحالية. لا ينبغي لمهمة تحليل التعليمات البرمجية أن تمنح الوكيل إمكانية الوصول إلى أدوات الكتابة في قاعدة البيانات أو أوامر حذف الملفات. تفرض المنصة هذا — وليس إعدادات الجلسة الفردية.
- تحديد نطاق رؤية خادم MCP لكل مهمة ولكل مستخدم
- إزالة الأدوات التي لا تحتاجها المهمة، بدلاً من الثقة في أن الوكيل سيتجاهلها
- استخدم الـ MCP Gateway لتصفية الأدوات التي يمكن لكل جلسة الوصول إليها
تصفية المخرجات للمحتوى الحساس
يجب أن تمر مخرجات Claude Code عبر مرشح لأنماط البيانات الحساسة قبل أن يتم الالتزام بها أو نشرها أو إرسالها. تكتشف تصفية المخرجات محاولات تسريب البيانات التي تستخدم قنوات إخراج مشروعة — مثل التزامات التعليمات البرمجية، تعليقات طلبات السحب، واستجابات واجهة برمجة التطبيقات — لتهريب البيانات.
سجلات التدقيق غير القابلة للتغيير والمرتبطة بالهوية
يجب أن ينتج عن كل إجراء لـ Claude Code إدخال سجل يتضمن المهمة الأصلية، وهوية المستخدم، والمحتوى المعالج، والإجراء المتخذ. توفر سجلات التدقيق المسار الجنائي اللازم لإعادة بناء ما حدث في حالة حقن. يجب أن تبقى السجلات ضمن بيئتك — ولا يتم إعادة توجيهها إلى منصات SaaS خارجية — لتلبية متطلبات HIPAA و SOC 2 وقانون الاتحاد الأوروبي للذكاء الاصطناعي.
ضوابط خروج الشبكة
يمنع تقييد وصول كود كلود (Claude Code) إلى الشبكة الخارجية بقائمة سماح محددة التعليمات المحقونة من تسريب البيانات بنجاح. فالحقن الناجح الذي لا يمكنه الوصول إلى وجهة خارجية يكون تأثيره محدودًا. لكن هجوم "كلودي داي" (Claudy Day) أظهر أن نقاط النهاية المدرجة في قائمة السماح (api.anthropic.com) يمكن استخدامها هي نفسها للتسريب — لذا يجب دمج ضوابط الخروج مع تصفية المخرجات.
كيف تعالج TrueFoundry مخاطر حقن الأوامر وأمن وكلاء الذكاء الاصطناعي
تعمل TrueFoundry على مبدأ وجوب معالجة مخاطر أمن وكلاء الذكاء الاصطناعي على مستوى البنية التحتية. يتم نشر المنصة بالكامل داخل بيئة AWS أو GCP أو Azure الخاصة بك. وتتم جميع عمليات التصفية والتسجيل والإنفاذ ضمن حدود شبكتك.
- تصفية المحتوى على مستوى البنية التحتية. يتم تحليل المحتوى الوارد بحثًا عن أنماط الحقن قبل دخوله نافذة سياق كود كلود (Claude Code). ويتم اعتراض الهجمات عند الاستيعاب، وليس بعد التنفيذ.
- سجل الأدوات بأقل الامتيازات. تعرض بوابة MCP فقط الأدوات ذات الصلة بمهمة الوكيل الحالية. ولا يمكن لمحاولات الحقن الوصول إلى أدوات خارج نطاق المهمة. للحصول على معلومات أساسية حول كيفية عمل اتصالات MCP، راجع دليل تكاملات MCP.
- تصفية مخرجات معلومات التعريف الشخصية (PII) والبيانات الحساسة. يتم فحص مخرجات كود كلود (Claude Code) بحثًا عن أنماط البيانات الحساسة قبل مغادرتها بيئة التنفيذ. ويتم حظر التسريب عبر قنوات الإخراج المشروعة.
- حقن الهوية عبر OAuth 2.0. يرتبط كل إجراء للوكيل بصلاحيات محددة لمستخدم مصادق عليه. ولا يمكن للتعليمات المحقونة أن تتجاوز ما هو مصرح للمستخدم الأصلي القيام به.
- سجلات تدقيق غير قابلة للتغيير بمحتوى كامل. يتم تسجيل كل طلب، واستدعاء أداة، وقراءة ملف، ومخرج ببيانات وصفية كاملة. وتبقى السجلات في بيئتك للتحقيق الجنائي والامتثال. الـ دليل أمان المؤسسات يغطي إعداد التدقيق الكامل.
- ضوابط خروج الشبكة. جميع حركة المرور الصادرة من جلسات Claude Code تمر عبر سياسات خروج محكمة. يتم حظر المكالمات الخارجية العشوائية التي تحقن التعليمات. الـ بوابة الذكاء الاصطناعي توفر نقطة التحكم الوحيدة لجميع حركة مرور النموذج.
تحصل المؤسسات التي تستخدم TrueFoundry لنشر Claude Code على دفاع متعدد الطبقات ضد حقن الأوامر عبر طبقات متعددة في وقت واحد — تصفية المدخلات، تحديد نطاق الأدوات، تصفية المخرجات، ضوابط الهوية، واحتواء الشبكة — دون الحاجة إلى تغييرات على مستوى التطبيق في الجلسات الفردية. الـ إطار الحوكمة يغطي كيفية بناء سياسات تنظيمية حول هذه الضوابط.
إذا كان فريقك يشغل Claude Code على محتوى لا يتحكم فيه بالكامل — المستودعات، التذاكر، استجابات واجهة برمجة التطبيقات، المستندات المسترجعة — فإن حقن الأوامر يمثل خطرًا فعليًا، وليس مجرد قلق مستقبلي. توفر TrueFoundry التصفية على مستوى البنية التحتية، وتحديد نطاق الأدوات، واحتواء الشبكة التي تكتشف هذه الهجمات قبل أن تصل إلى مرحلة التنفيذ. احجز عرضًا توضيحيًا لترى كيف يعمل ضد أنماط الحقن الحقيقية.
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.













.webp)


.webp)


.webp)
.webp)





.webp)
.webp)








