Why MCP trust gaps put AI agent networks at risk

Recent vulnerability acknowledgments from Google and four other organizations show how AI agents can pass harmful instructions across trusted internal systems. The issue centers on MCP, where agent-to-agent workflows can turn prompt injection into SSRF, data exposure, and broader protocol-level risk.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 0 ►

The story centers on AI agent trust failures enabling prompt injection, SSRF, data exposure, and harmful internal actions across networks.

Why MCP trust gaps put AI agent networks at risk

AI agents are being adopted across millions of organizations, and that shift is opening a new security problem: malicious instructions can move from one agent to another inside trusted systems. Recent findings around MCP show how a prompt aimed at one specialized agent can become a path to more powerful internal actions.

The risk inside agent-to-agent trust

In the past five months, Google and four other organizations have acknowledged vulnerabilities tied to AI agents passing harmful instructions across internal chains. The organizations had little in common beyond their use of AI agents, which makes the pattern important.

The attacks are a specialized form of prompt injection. Instead of targeting only the LLM, the malicious prompt targets a specific agent, such as one used for translation or data analysis. If that agent lacks strong guardrails, it may forward the instruction to another agent as part of an ordinary workflow.

The problem grows because the receiving agent may trust the first agent. That second agent can then treat the harmful instruction as a legitimate delegated task. In that chain, the dangerous action does not come from an obvious external command; it arrives through internal cooperation.

This matters because agent systems are designed to hand work to each other. When that delegation crosses a weak trust boundary, a prompt that might have been blocked by an LLM can still succeed elsewhere in the system.

How MCP changes the attack surface

MCP, short for Model Context Protocol, is one way AI apps and agents communicate with each other inside an internal network. Independent researcher Syed Anas Mohiuddin tested agents from organizations including Google, JP Morgan Chase, Weviate, Rapid7, the French government’s interministerial digital directorate, and the US federal government.

His proof-of-concept attacks focus on trust gaps in MCP. Many special-purpose agents do not have the guardrails that might reduce the consequences of prompt injection. MCP servers also store credentials for each agent, while agents are often built to trust other internal agents.

That combination can make a harmful prompt more effective than it first appears. A prompt crafted for the right agent may lead to server-side request forgery, a vulnerability that causes a web server to make unauthorized network requests.

Douglas McKee, director of vulnerability intelligence at Rapid7, described the issue as a chain where each component behaves as designed while the overall system becomes hard to defend. The key lesson is that internal delegation can hide the origin and risk of the instruction being passed along.

Google and Rapid7 show different severity levels

The Rapid7 issue found by Mohiuddin was tracked as CVE-2026-97228. It carried a severity rating of only 2.7 out of 10, and Rapid7 fixed it last month.

The Google vulnerability was more severe, with a rating of 8. It came from an MCP toolbox for databases, identified as googleapis/mcp-toolbox. The issue involved an HTTP client initialized without a CheckRedirect policy, which controls how a server handles errors or redirects to a different URL.

Google’s HTTP client also failed to validate target IP addresses. Mohiuddin explained that “A crafted path parameter could make the toolbox follow a redirect to an internal endpoint and send requests on the attacker’s behalf.”

Google’s fix used an allow-list of IP ranges and block lists. Mohiuddin said the fix rejects an unsafe base URL at startup rather than waiting until the first request, calling that “a real SSRF guard.”

Protocol pivoting or prompt injection?

Mohiuddin calls this class of attack “protocol pivoting.” The name reflects how an attacker can gain initial access through one protocol, exploit trust assumptions between protocols, and reach capabilities available through a different protocol.

In this model, MCP may be used to assign a task to an agent. That agent can then forward malicious instructions through another communication method, including Google’s Agent-to-Agent (A2A) protocol, used for inter-agent delegation, or emerging standards such as the Agent Network Protocol.

The central concern is that trust or authorization can get lost as the task moves between systems. A request that looks routine to one protocol may become more dangerous when another agent interprets it under a different trust model.

Not everyone agrees that the new label is necessary. Markus Vervier, a researcher at X41 D-Sec who has also devised AI attacks that exploit MCP, said the better term remains “prompt injection.” He described Mohiuddin’s technique as a simple subclass of that broader issue.

Vervier called it indirect prompt injection and said the cross-protocol path is not strictly required for such attacks to work. He also said the behavior is “unexpected and hard to mitigate in general.”

The zero trust lesson for agent systems

The broader warning is not limited to one vendor or one implementation. The pivoting technique worked across five organizations whose shared factor was MCP. The source article describes MCP as new and already widely present before it has been sufficiently tested and hardened.

The security issue is also architectural. Organizations building sprawling agentic systems may be moving away from zero trust, a principle that assumes one or more nodes may be infected. Under zero trust, nodes should require authorization before carrying out sensitive transactions with other nodes.

For AI agent networks, that means internal messages should not be treated as automatically safe. Inputs passed from an LLM to a tool can carry the same risk as input from outside the organization when prompt injection is involved.

The underlying bugs are not entirely new. The source points to older categories such as injection and SSRF. What changes with MCP and agent-to-agent workflows is the path those bugs can take through delegated tasks, trusted agents, stored credentials, and multiple protocols.

For defenders and standards bodies, the practical implication is clear: MCP servers, agent tools, and inter-agent protocols need security assumptions that match how these systems are actually used. When agents can act for each other, trust must be checked at every handoff, not only at the first door.