The Model Context Protocol, or MCP, has become the standard way to connect AI agents to the outside world. An MCP server can give an agent access to your files, databases, ticketing system or a third-party API. That makes agents useful, and a target. One of the most effective attacks on MCP doesn't exploit a bug in the protocol. It exploits how models read.
How tool poisoning works
An MCP server describes each tool to the agent: a name, an explanation and its inputs. The model reads those descriptions to decide which tool to use and how. The person approving a tool usually sees a short summary. The model sees everything.
In April 2025, researchers at Invariant Labs showed what that gap allows. Their harmless-looking tool, add, added two numbers. Its description hid instructions for the model: read the user's MCP configuration file and private SSH key, and pass them along in an extra parameter called "sidenote." The user saw the correct sum. The model had sent their keys to the attacker.

Two variants make it worse. In a rug pull, a server changes a tool's description after you've approved it: clean on Monday, poisoned by Friday. In shadowing, a malicious server's description changes how the agent uses tools from trusted servers. In Invariant's example, a poisoned tool made the agent send every email to the attacker, whatever recipient the user chose.
Unlike ordinary prompt injection, which arrives in content read during a task, tool poisoning arrives in the tool's own metadata when the agent connects: present in every conversation, where few people look, with the authority of a tool the user chose.
Better models aren't the fix
MCPTox, a benchmark built on 45 live MCP servers and 353 real tools, ran 1,312 poisoning test cases against a range of models. The average attack success rate across all model settings was 36.5%, and the most vulnerable models were compromised more than 70% of the time. Even the model that refused most often did so in under 3% of cases. More capable models were often more susceptible, because the attack exploits what makes them good: following instructions closely. The controls have to sit around the model.
The controls that work
In May 2026 the US National Security Agency published guidance titled "Model Context Protocol (MCP): Security Design Considerations for AI-Driven Automation." It names the conditions tool poisoning exploits, among them uncontrolled automated actions, context poisoning, weak access controls and missing human oversight, and recommends maintained tools, human sign-off, minimum access, input screening, activity logs and separating systems by trust level.

Before you connect a server
- Allowlist servers your team has reviewed and trusts.
- Pin versions and hash tool definitions, and alert when one changes: the direct defense against rug pulls, as Invariant recommended.
- Review the whole description the model will see, and scan it for hidden instructions with tools such as mcp-scan.
While the agent runs
- Least privilege per tool. The MCP specification says servers must not accept tokens that weren't issued for them, and recommends minimal scopes.
- Sandbox local servers, as the specification recommends. A server that can't read your SSH directory can't leak it.
- Allowlist outbound destinations, so stolen data has nowhere to go.
- Separate trust levels. Keep untrusted servers out of sessions with high-privilege internal tools, where shadowing does its damage.
On sensitive actions
- Require human approval that shows every parameter. A prompt that hides the "sidenote" field protects nothing.
- Log every tool call: tool, agent, inputs and result, fed into your security monitoring.
Before release
- Red-team the agent with deliberately poisoned tools, and keep those cases in the evaluations run whenever a model, prompt or tool changes.
Treat tool descriptions as untrusted input
Ask about the agents you already run:
- Which MCP servers can they connect to, and who approved each one?
- Would we notice if a tool's description changed tomorrow?
- What could a compromised agent's credentials reach?
- Which actions run without a person seeing every parameter?
- Have we ever tested with a poisoned tool?
If any answer is "we don't know," start there. Anything an agent reads can steer what it does, so treat everything from outside your control like input on a public web form: untrusted until checked. Building these controls in early costs far less than discovering you needed them.
If you're building agents on MCP, our data and AI and custom software teams can design or review them with these controls.
