Skip to content

Tool poisoning: how attackers hijack AI agents through MCP, and the controls that stop them

AI agents read every word of a tool's description, including the parts users never see. How tool poisoning works, why capable models fall for it, and the controls that keep it out of production.

An engineer in glasses studying a screen in a dark room
The part of a tool description that matters most is often the part nobody reads.

By Mayrian

· 3 min read

Share

The Model Context Protocol, or MCP, has become the standard way to connect AI agents to the outside world. An MCP server can give an agent access to your files, databases, ticketing system or a third-party API. That makes agents useful, and a target. One of the most effective attacks on MCP doesn't exploit a bug in the protocol. It exploits how models read.

How tool poisoning works

An MCP server describes each tool to the agent: a name, an explanation and its inputs. The model reads those descriptions to decide which tool to use and how. The person approving a tool usually sees a short summary. The model sees everything.

In April 2025, researchers at Invariant Labs showed what that gap allows. Their harmless-looking tool, add, added two numbers. Its description hid instructions for the model: read the user's MCP configuration file and private SSH key, and pass them along in an extra parameter called "sidenote." The user saw the correct sum. The model had sent their keys to the attacker.

Diagram: what the user sees, 'add(a, b): adds two numbers', versus what the model reads, which includes hidden instructions to read ~/.ssh/id_rsa and pass it as a sidenote, followed by the four steps of the attack
The same tool, as the user sees it and as the model reads it. Based on Invariant Labs' proof of concept, simplified.

Two variants make it worse. In a rug pull, a server changes a tool's description after you've approved it: clean on Monday, poisoned by Friday. In shadowing, a malicious server's description changes how the agent uses tools from trusted servers. In Invariant's example, a poisoned tool made the agent send every email to the attacker, whatever recipient the user chose.

Unlike ordinary prompt injection, which arrives in content read during a task, tool poisoning arrives in the tool's own metadata when the agent connects: present in every conversation, where few people look, with the authority of a tool the user chose.

Better models aren't the fix

MCPTox, a benchmark built on 45 live MCP servers and 353 real tools, ran 1,312 poisoning test cases against a range of models. The average attack success rate across all model settings was 36.5%, and the most vulnerable models were compromised more than 70% of the time. Even the model that refused most often did so in under 3% of cases. More capable models were often more susceptible, because the attack exploits what makes them good: following instructions closely. The controls have to sit around the model.

The controls that work

In May 2026 the US National Security Agency published guidance titled "Model Context Protocol (MCP): Security Design Considerations for AI-Driven Automation." It names the conditions tool poisoning exploits, among them uncontrolled automated actions, context poisoning, weak access controls and missing human oversight, and recommends maintained tools, human sign-off, minimum access, input screening, activity logs and separating systems by trust level.

Diagram of layered controls: before you connect, while it runs, on sensitive actions, and before release
Controls drawn from the MCP specification's security best practices, NSA guidance and Invariant Labs' recommendations.

Before you connect a server

  • Allowlist servers your team has reviewed and trusts.
  • Pin versions and hash tool definitions, and alert when one changes: the direct defense against rug pulls, as Invariant recommended.
  • Review the whole description the model will see, and scan it for hidden instructions with tools such as mcp-scan.

While the agent runs

  • Least privilege per tool. The MCP specification says servers must not accept tokens that weren't issued for them, and recommends minimal scopes.
  • Sandbox local servers, as the specification recommends. A server that can't read your SSH directory can't leak it.
  • Allowlist outbound destinations, so stolen data has nowhere to go.
  • Separate trust levels. Keep untrusted servers out of sessions with high-privilege internal tools, where shadowing does its damage.

On sensitive actions

  • Require human approval that shows every parameter. A prompt that hides the "sidenote" field protects nothing.
  • Log every tool call: tool, agent, inputs and result, fed into your security monitoring.

Before release

  • Red-team the agent with deliberately poisoned tools, and keep those cases in the evaluations run whenever a model, prompt or tool changes.

Treat tool descriptions as untrusted input

Ask about the agents you already run:

  1. Which MCP servers can they connect to, and who approved each one?
  2. Would we notice if a tool's description changed tomorrow?
  3. What could a compromised agent's credentials reach?
  4. Which actions run without a person seeing every parameter?
  5. Have we ever tested with a poisoned tool?

If any answer is "we don't know," start there. Anything an agent reads can steer what it does, so treat everything from outside your control like input on a public web form: untrusted until checked. Building these controls in early costs far less than discovering you needed them.

If you're building agents on MCP, our data and AI and custom software teams can design or review them with these controls.

How we can help

Choose a service to see its capabilities. Point at one to see what it's used for and what you receive.

All services

Custom Software Development

Software Development

Web platforms and internal tools built around how your business works, released in small, tested increments.

Used for

  • Customer portals
  • Internal tools
  • SaaS products
  • Workflow and approvals
  • Marketplaces and booking platforms
  • Enterprise application extensions

What you receive

  • Source code in your repository, with a documented architecture
  • CI/CD pipeline and infrastructure as code
  • Automated test suite
  • Monitoring, logging and alerting
  • Security review against the OWASP Top 10:2025 and ASVS 5.0
  • Ownership and documentation handover
More on Software Development

Working on something like this?

Start with a free technical consultation: a plan covering the right tech stack, architecture, timeline and budget.

Start a project