Field Guide to AI, Security and Cybercrime

The Lethal Trifecta: When an AI Agent Can Read, Be Steered and Send

Why combining private data access, untrusted content exposure, and external communication guarantees exploitable agents.

Field Guide to AI, Security and Cybercrime·Abdolmadjid Masoomi·4 October 2026·8 min read

The lethal trifecta—private data access, untrusted content exposure, and exfiltration capability—is the architectural flaw behind every major agentic data breach. When all three converge, prompt injection becomes automatic theft. This essay explains the mechanism, the root cause, and the non-negotiable rule to break the chain.

The agent that reads, is steered, and sends

You deploy an AI agent to read your emails, summarise logs, and call APIs to act on insights. You assume it is a tool. In reality, you have built a perfect data exfiltration engine. When an agent can read untrusted content, be steered by hidden instructions, and send data externally, prompt injection becomes automatic theft. This is not a theoretical risk. It is the lethal trifecta: access to private data, exposure to untrusted content, and an exfiltration vector. Every major agentic breach—from GrafanaGhost to EchoLeak—follows this pattern. The flaw is not in the model; it is in the architecture that lets all three capabilities converge in a single session.

The LLM has no reliable way to distinguish between instructions and data. It processes system prompts, user queries, and retrieved content as one token stream. This architectural limitation means that any hidden instruction embedded in ingested data can hijack the agent’s goal. When the agent also has access to private data and a way to send it out, the hijack becomes a breach. You cannot fix the model’s reasoning. You must constrain its environment.

How the trifecta is exploited

The attack path is simple and repeatable. An attacker plants malicious instructions into content the agent will ingest—this is indirect prompt injection. The instructions may be hidden in a log entry using zero-width characters, embedded in a calendar invite description, or concealed in a shared document via off-screen HTML elements. The agent retrieves this poisoned context via RAG or direct search. Because the model cannot reliably separate instruction from data, it executes the hidden command as part of its reasoning.

Now the trifecta completes the crime. The agent uses its access to private data to gather the target information. It uses its exposure to untrusted content to justify the action as part of its normal task. And it uses its exfiltration vector—image rendering, email send, API call—to deliver the data to the attacker. In the GrafanaGhost attack, a hidden image URL in a log entry triggered the agent to fetch a payload that instructed it to exfiltrate subsequent logs. In EchoLeak, a poisoned calendar event caused Microsoft 365 Copilot to leak data via its email capabilities.

This pattern is amplified by the Model Context Protocol (MCP), which connects agents to external tools and data sources. A malicious MCP server can embed adversarial instructions in tool descriptions—the natural-language text the LLM reads to understand how to use the tool. This is tool poisoning. The agent thinks it is executing a legitimate function, but the hidden instructions redirect its behaviour.

The supply chain is now an instruction pipeline. Any third-party component—MCP servers, browser extensions, SaaS integrations—can become the poison source. The BragJack attack showed how a malicious browser extension can hijack the trusted channel between a Chromium-based AI assistant and its vendor origin, granting the attacker the agent’s elevated privileges. This is not a one-off flaw; it is a systemic risk inherent to agentic architectures that rely on dynamic tool loading.

The architectural rule: never all three

The defence is not incremental. You must enforce a strict architectural rule: no single agent session may possess all three properties of the lethal trifecta without a mandatory, human-approved checkpoint. This is the "Agents Rule of Two". An agent can have access to private data and process untrusted content, but it must have no direct exfiltration capability. It can browse the web and call APIs, but it must have no access to sensitive internal data. It can have data access and exfiltration tools, but only for pre-vetted, trusted content.

This rule is not optional. It is the only reliable way to break the prompt injection-to-data-theft chain. Since you cannot prevent injection, you must ensure that even if injection succeeds, the agent cannot complete the attack.

Controlling exfiltration vectors

Even with segmentation, you must treat every exfiltration vector as a potential data diode. Start with the most common: image rendering and markdown. Block all external image loading by default. Enforce a strict Content Security Policy that prevents img tags from fetching data from external domains. Sanitise markdown output to strip or disable all but the simplest formatting. Any feature that allows the agent to initiate a network request must be scrutinised.

For agents that legitimately need to call APIs, implement a stringent allow-list and proxy all calls through a security gateway. The gateway should validate the destination, inspect parameters for signs of data exfiltration, and enforce rate limits. It should also tokenise or strip any sensitive data that is not explicitly required for the external service’s function. Treat the agent’s output channel not as a benign display layer but as a potential data diode that must be filtered.

This control directly mitigates risks from the agentic supply chain. A malicious MCP server might try to use a legitimate tool for exfiltration. If the agent’s ability to call that tool is scoped to a specific allow-listed set of parameters and destinations, the attack is contained. The Cloud Security Alliance recommends capability-level permission scoping: an agent should not have a token with files:* access; it should have a token that only allows reading from a specific project folder. This limits the damage from both tool misuse and credential theft.

You must also enforce TLS 1.2+ for all remote connections, use short session token lifetimes, bind tokens to source IP or agent identity, and implement refresh token rotation. These measures prevent session hijacking and token theft, which are common vectors for exfiltration.

Poisoned tooling and the supply chain

The lethal trifecta is not just about poisoned data; it is increasingly about poisoned tools. The Model Context Protocol creates a new class of supply chain risk where the tools themselves become the injection vector. A malicious MCP server can embed adversarial instructions within its tool descriptions. The agent reads these descriptions to understand how to use the tool, and the hidden instructions are executed as part of its reasoning. This is a direct subversion of the trust you place in an integrated component.

This threat extends beyond MCP. Any plugin system or tool-calling framework where the LLM dynamically learns about capabilities from an external source is vulnerable. A malicious webpage can carry instructions that hijack the agent’s subsequent tool calls, redirecting them to attacker-controlled endpoints. The agent thinks it is using a legitimate tool, but the parameters have been altered to exfiltrate data. This shows that the boundary between "untrusted content" and "tool" is porous; a poisoned webpage can corrupt the agent’s understanding of its own tools.

Your defence must treat tool definitions as untrusted content. Do not allow agents with sensitive data access to dynamically load tools from unvetted sources. For approved tools, implement cryptographic verification of tool schemas and descriptions in a private registry. This can prevent rug pull attacks where a previously benign server is swapped for a malicious one. Furthermore, you should sandbox the execution of any tool call. The tool’s output should be treated as potentially malicious input and not given direct, high-privilege access to other systems.

Questions people ask

What is the single most important control to implement?

Segment your agent access so that no single agent session has all three properties of the lethal trifecta. This is a foundational architectural rule. An agent with access to sensitive data should not also be able to freely process untrusted web content and call external APIs. Enforcing this "Rule of Two" breaks the prompt injection-to-data-theft chain at the design level, making many sophisticated attacks impossible to complete.

How does the Model Context Protocol (MCP) change the risk?

MCP expands the attack surface by introducing many third-party tools and data sources. It creates new supply chain risks, such as malicious MCP servers, tool poisoning, and server spoofing. While MCP is powerful, you must treat every connected server as a potential source of untrusted content and a potential exfiltration vector. Strict allow-listing, cryptographic verification of tools, and network-level isolation for MCP connections are non-negotiable for any agent handling private data.

Can't we just fix prompt injection instead?

Prompt injection is a symptom of the LLM's fundamental architecture—it cannot reliably distinguish between instructions and data. While research continues, there is no general, reliable technical fix for this in the current generation of models. Therefore, you must build defences that assume injection will succeed. The lethal trifecta framework provides a clear, actionable model for those defences: limit the agent's capabilities so that a successful injection cannot lead to a serious breach.

What about agents that need all three properties to function?

If a business process genuinely requires an agent to have access to private data, process untrusted content, and communicate externally, you must insert a mandatory human approval checkpoint for any external communication action. The agent can prepare a summary or a draft, but a human must review and explicitly authorise the send, post, or API call. This procedural control breaks the automated exfiltration loop, turning a technical exploit into a social engineering attempt that a vigilant human can block.

Close

The lethal trifecta is not a complex theoretical model; it is a simple, brutal lens through which to evaluate the inherent danger of any AI agent deployment. If you see an agent with its hands in your data, its eyes on the open internet, and a voice that can shout to the outside world, you are looking at a data breach waiting to happen. The attacks against major enterprise copilots prove the pattern is being actively exploited. Your defence starts with recognising this combination as a cardinal sin in agent architecture.

The path forward requires a shift from feature-led integration to security-led design. You must map your agents' data accesses, content sources, and output channels. You must then deliberately break the trifecta through segmentation, stringent output controls, and rigorous supply chain management. This is the discipline required to harness the power of agentic AI without becoming its next victim. The work to secure shadow AI and the pipelines it touches begins with understanding and neutralising this core risk pattern. Build agents that are useful, but build them so that even when they are fooled, they cannot betray you.

Sources

  1. Cloud Security Alliance - Agentic MCP Security Best Practices (opens in a new tab)
  2. Cloud Security Alliance - Indirect Prompt Injection in the Wild (opens in a new tab)
  3. The Hacker News - Official MCP Python SDK Flaw (opens in a new tab)
  4. Tech Insider - BragJack Attack (opens in a new tab)
  5. Infosecurity Magazine - Infosec Europe Prompt Injection (opens in a new tab)
  6. SentinelOne - MCP Security (opens in a new tab)
  7. Human Security - OWASP Top 10 Agentic Applications (opens in a new tab)

Questions people ask

What is the single most important control to implement?

Segment your agent access so that no single agent session has all three properties of the lethal trifecta. This is a foundational architectural rule. An agent with access to sensitive data should not also be able to freely process untrusted web content and call external APIs. Enforcing this "Rule of Two" breaks the prompt injection-to-data-theft chain at the design level, making many sophisticated attacks impossible to complete.

How does the Model Context Protocol (MCP) change the risk?

MCP expands the attack surface by introducing many third-party tools and data sources. It creates new supply chain risks, such as malicious MCP servers, tool poisoning, and server spoofing. While MCP is powerful, you must treat every connected server as a potential source of untrusted content and a potential exfiltration vector. Strict allow-listing, cryptographic verification of tools, and network-level isolation for MCP connections are non-negotiable for any agent handling private data.

Can't we just fix prompt injection instead?

Prompt injection is a symptom of the LLM's fundamental architecture—it cannot reliably distinguish between instructions and data. While research continues, there is no general, reliable technical fix for this in the current generation of models. Therefore, you must build defences that assume injection will succeed. The lethal trifecta framework provides a clear, actionable model for those defences: limit the agent's capabilities so that a successful injection cannot lead to a serious breach.

What about agents that need all three properties to function?

If a business process genuinely requires an agent to have access to private data, process untrusted content, and communicate externally, you must insert a mandatory human approval checkpoint for any external communication action. The agent can prepare a summary or a draft, but a human must review and explicitly authorise the send, post, or API call. This procedural control breaks the automated exfiltration loop, turning a technical exploit into a social engineering attempt that a vigilant human can block.

Ask NEXUS about this article

Get an AI-powered summary, key points, or follow-up questions about The Lethal Trifecta: When an AI Agent Can Read, Be Steered and Send, grounded in the essay content and the broader corpus.

Field Guide to AI, Security and Cybercrime135 / 131