
The rapid spread of AI agents inside enterprises is exposing a security problem that may be more worrying than many organizations realize: Model Context Protocol, or MCP, can let malicious instructions move from one internal agent to another in ways that bypass the usual safeguards. Researchers say the issue is not a flaw in one company’s model so much as a structural weakness in the way agent-to-agent communication is being built and trusted.
MCP is becoming a bridge for attacks, not just a convenience layer
Over the past five months, Google and four other organizations with little in common beyond their use of AI agents have acknowledged vulnerabilities that allowed one agent inside a network to spread harmful instructions to others. The attack pattern is a form of prompt injection, but it is aimed at a specific agent rather than the large language model itself. Once a vulnerable agent accepts the instructions, it can relay them down the chain because other agents trust it as an internal peer.
That trust is the core issue. Special-purpose agents, such as tools for translation or data analysis, often have weak or absent guardrails. When they are connected through MCP, they may pass along instructions that would have been rejected if they were delivered directly to a model or reviewed by a human.
What the researcher found across Google, banks, governments, and security vendors
Independent researcher Syed Anas Mohiuddin tested agents from Google, JPMorgan Chase, Weviate, Rapid7, the French government’s interministerial digital directorate, and the US federal government. His proof-of-concept attacks focused on trust gaps in MCP, the Model Context Protocol used by AI apps and agents to communicate internally across an organization’s network.
According to the report, the weakness is amplified by the way MCP servers store credentials for each agent and by the assumption that internal agents can be trusted to hand work off safely. In some cases, a carefully crafted prompt can lead to server-side request forgery, or SSRF, a classic bug that makes a server issue unauthorized network requests.
Why the attack chain works
The problem is not that every agent is equally exposed to the original prompt. Instead, the attack often succeeds because one agent receives the malicious text, interprets it as a delegated task, and passes it to another agent that accepts it as legitimate. By the time the request reaches a more privileged service, the malicious content has traveled through multiple trusted layers.
As Rapid7 director of vulnerability intelligence Douglas McKee put it, “AI agents give attackers a fresh set of connections to walk across.” He said the chain works because “every piece in that chain did exactly what it was designed to do,” even though the overall result is unsafe.
One Google flaw was severe; another was rated low
The vulnerabilities were not all equal. Rapid7’s issue, tracked as CVE-2026-97228, carried a severity score of 2.7 out of 10 and was fixed last month. Google’s issue was more serious, with a severity rating of 8.
Syed said the Google flaw affected an MCP toolbox for databases, googleapis/mcp-toolbox. The code initialized its HTTP client without a CheckRedirect policy and also failed to validate target IP addresses. That meant a crafted path parameter could cause the toolbox to follow a redirect to an internal endpoint and send requests on the attacker’s behalf.
How Google fixed its implementation
Google’s remediation used an allow-list of IP ranges and block lists. Syed said the software now rejects an unsafe base URL at startup instead of waiting until the first request. He described that as the kind of SSRF defense he considers meaningful, but also noted that it requires more effort than many MCP servers have invested so far.
Why Syed calls it protocol pivoting
Syed is using the term “protocol pivoting” to describe attacks that begin in one communication system and end in another. In his view, the exploit starts when an app or server uses MCP to assign work to an agent, then the agent forwards malicious instructions through a different method, such as Google’s Agent-to-Agent, or A2A, protocol, or newer systems like the Agent Network Protocol.
He defines it as “a multi-step attack in which an adversary gains initial access through one protocol, exploits trust assumptions between protocols, and escalates to capabilities only accessible via a different protocol.” The key risk is that trust or authorization can be lost in translation as the request moves between systems.
Not everyone agrees on the label
Markus Vervier of X41 D-Sec, another researcher who has explored MCP-based AI attacks, said the better description is still indirect prompt injection. He argued that the fact that the malicious prompt arrives through a different protocol is not essential to the attack, even if it makes the behavior more surprising and harder to mitigate.
That disagreement is mostly about terminology. Both researchers agree that the underlying issue is a prompt or instruction arriving from an untrusted source and being accepted by a trusted agent later in the chain.
The bigger security lesson: zero trust still matters
The fact that Syed found the technique working across five organizations is what makes the report noteworthy. MCP is new, but it is already widely deployed before many of these agent chains have been thoroughly hardened or tested under realistic attack conditions.
Security experts say the rush to build sprawling agentic systems has led some teams to abandon zero trust principles. In a zero trust design, each component assumes another part of the network may be compromised and requires explicit authorization before sensitive actions are allowed.
- Any text handed from an LLM to a tool should be treated as untrusted input.
- Internal agents should not automatically trust instructions from other internal agents.
- Redirect handling, IP validation, and allow-lists remain essential defenses against SSRF.
- Agent handoffs should be authenticated and constrained, not assumed safe by default.
A familiar bug class in a new place
McKee said the bugs underneath are old ones, including injection and SSRF, and that the fixes have not changed in 20 years. What has changed is the environment: AI agents now create more pathways for attackers to move through trusted internal systems.
That makes MCP less of a niche integration detail and more of a security boundary. As agent deployments grow, the question is no longer just what an AI can answer, but what it can cause other systems to do once it has been given a task.
Source: Original report
Was this helpful?
Explore more: Application Audit & Review More Cybersecurity Tech News
Last Modified: October 6, 2026 at 10:31 pm
0 views

