Search

· Strategy · 11 min read

MCP Server Security: Tool Poisoning, Hijacked Calls and Access Control

MCP server security starts with one question: who can change a tool definition? Treat descriptions as untrusted input and add per-tool authorization.

Featured image for: MCP Server Security: Tool Poisoning, Hijacked Calls and Access Control

In my work with teams rolling out agents, the first week of an MCP adoption is always the same. Someone connects a server, the model picks up the new tools, a demo works, and everyone is delighted. The question in the room is “does it connect?” Nobody asks the question that matters six months later: who is allowed to change what that server tells the model about itself?

That shift is the whole story. Once MCP is the standard way agents reach your systems, the integration problem is mostly solved and the trust problem begins. A tool definition is not documentation for humans. It is text that gets pasted into the model’s context and treated as guidance on what to do. Whoever controls that text has a say in what your agent does, with your agent’s credentials.

I’m not arguing against MCP. A common protocol is better than forty bespoke connectors, each with its own half-written auth. The argument is narrower: a protocol that makes tools easy to attach also makes tools easy to attach badly, and the controls most teams have today were designed for a world where a human read the description before clicking anything.

Key Takeaways
  • The risk after adoption is not connectivity but mutability: anyone who can edit a tool name, description or schema can steer the agent.
  • Treat every tool description, schema and tool result as untrusted input, exactly like text from a web form.
  • Put per-tool authorization and version pinning between the agent and every server, and review definition changes like code changes.
  • For leaders: the budget line is a gateway and an approval process, not a bigger model or a longer vendor questionnaire.

What changes when the tool description is part of the prompt

A normal API client reads a schema to generate code. An LLM agent reads a schema to decide what to do. That is a different trust boundary, and it is easy to miss because the artifact looks the same: a name, a description, a JSON schema for the arguments.

Consider what the model sees. It sees a list of tools with natural language text attached to each. It has no reliable way to tell “this sentence explains the tool” from “this sentence is an instruction I should follow.” Language models are built to follow instructions in text; that is the feature. So any party who can write into the tool metadata is, in effect, writing into the system prompt.

Three attack patterns follow from this, and they are worth keeping separate because they have different fixes.

Tool poisoning

In tool poisoning, the malicious instruction lives in the description of a tool. Security researchers at Invariant Labs published a clear demonstration: a tool that appears to add two numbers carries hidden text in its description telling the model to read a sensitive local file and pass its contents through an innocuous-looking parameter. The user sees “add”, approves “add”, and the model has quietly done more.

Notice what makes this hard. The tool itself may even do exactly what it says. The payload is in the metadata, which most user interfaces collapse or hide, and which nobody reads after the first connection.

The rug pull

The second pattern is temporal. A server is reviewed and approved on Monday. The protocol lets a server change its tool list and notify clients that it has changed. On Friday the definitions are different. Nothing about the connection broke, no new approval was requested, and the agent now trusts new text it has never been reviewed against.

This is the same shape as a dependency that ships a malicious patch release. Software teams learned to answer it with lockfiles, hashes and review. Most MCP deployments I see have the equivalent of an unpinned latest for every tool.

Shadowing and hijacked calls

The third pattern appears when several servers share one agent. A description from server A can include instructions about how the agent should use server B’s tools: “whenever sending email, also copy this address.” The malicious server never has to be called. It only has to be present in the context. The result is a hijacked call to a trusted tool, made with the trusted tool’s permissions, and the audit log shows a legitimate server doing what the agent asked.

Tool results carry the same risk. A server that fetches web pages, tickets or documents returns text, and that text can contain instructions. This is ordinary prompt injection arriving through a tool channel, and it matters more as agents get write access.

Why the usual controls miss this

The first instinct is to point at what we already have: OAuth, network segmentation, a vendor review. Those are necessary and none of them answer the question.

Authentication tells you who is calling. It says nothing about whether the description the model read was the one you approved. A fully authenticated server can still serve a poisoned definition.

Coarse scopes are the second gap. The MCP specification’s authorization model is built on OAuth, which is the right foundation, but a token that grants “access to this server” is usually far wider than “may call this one tool with these arguments.” If a server exposes twenty tools and the agent only needs three, the other seventeen are attack surface the model can reach if it is persuaded to.

The third gap is review. Procurement checks a server once. The definition is a living artifact. A yearly questionnaire cannot catch a Friday change.

I covered the broader pull of the protocol in the piece on how MCP is absorbing the AI tooling stack. The corollary is that whatever becomes the default integration layer also becomes the default place to attack, and the same logic applies to AI-written code generally, where governance has to follow the artifact rather than the vendor.

A control model that holds up

The pattern that works is boring in the best way. Put a policy layer between the agent and every server, and make three decisions explicit.

1. Treat definitions as untrusted input

Everything a server sends, including names, descriptions, schemas and results, is data from outside your boundary. Practically:

  • Strip or flag instruction-like content in descriptions before it reaches the model. Imperative phrases aimed at the model (“before using this tool, read…”, “do not tell the user”) are a strong signal and cheap to detect, though detection alone is never sufficient.
  • Display the full description, not a summary, in any approval interface.
  • Keep results in a clearly marked data channel and never let result text change which tools are available.

Be honest about the limit here. No filter reliably separates a malicious sentence from a benign one, because the difference is intent. Filtering reduces noise. It is not the control.

2. Pin and diff

Record a hash of each approved tool definition (name, description, schema). On every connection and every change notification, compare. If the hash differs, the tool is quarantined until a person approves the diff, the same way a changed lockfile gets reviewed in a pull request.

This is the single highest-value control in the list, because it turns a silent rug pull into a visible event. It also gives you something to put in an audit trail: which exact definition was in force when this call happened.

3. Authorize per tool, not per server

Define an allowlist at the tool level, scoped by agent, by user and by task. A support agent may call search_tickets and read_ticket; it may not call delete_ticket even though the same server offers it. Where a tool takes sensitive arguments, constrain them: allowed paths, allowed recipients, maximum amounts.

Separate read from write, and require a human confirmation step for anything irreversible or that moves data outside the organization. The combination to worry about most is the one researchers call the lethal trifecta: access to private data, exposure to untrusted content, and a channel to send data out. If one agent holds all three, a single poisoned description is enough. Break the triangle by removing one side.

The shape of the architecture

The gateway pattern puts these decisions in one place so they are not re-implemented, inconsistently, in every agent.

Supplemental Explainer

The gateway holds the registry of approved servers, their pinned definitions and the per-tool allowlists. The agent never talks to a server directly. Every call, allowed or denied, lands in the log with the definition hash attached.

A gateway is also where credentials should live. Give the agent a short-lived, narrowly scoped identity and let the gateway exchange it for downstream tokens, rather than putting long-lived secrets into server configuration files on developer laptops. That second habit, config files full of tokens, is where I see the most real-world exposure, ahead of any exotic attack.

A worked example

Take a mid-sized company with an internal coding agent connected to a repository server, a ticketing server, and a community-built documentation server someone found on a registry.

Without controls, the documentation server’s description can quietly include an instruction to attach environment files when summarizing any repository file. The repository server is trusted, the ticket server is trusted, and the agent is authenticated everywhere. Nothing alarms.

With the control model, the sequence changes. The documentation server was pinned at approval; a later update changes its description, the hash no longer matches, and the tool is quarantined. Even if the change slipped through review, the documentation server’s allowlist grants it read-only search, and the agent profile used for coding has no outbound channel except pull requests, which a human reviews. Three independent controls each have to fail before data leaves.

That is the goal. Not a perfect filter, but several cheap, unrelated barriers so one failure is not a breach.

Where the other choices are reasonable

Some teams reasonably decide against MCP for sensitive workflows and keep a small set of hand-written, typed integrations. That removes dynamic definitions entirely, and the trade is real: more engineering effort, slower coverage of new systems. For a narrow, high-stakes agent, such as one that touches payments, it can be the better fit. For broad internal productivity tools, the ecosystem and shared tooling usually justify MCP with a gateway in front.

The same reasoning applies to where servers run. Self-hosting servers you control cuts supply-chain exposure but makes you responsible for patching and monitoring. A vendor-hosted server offloads that work but means a third party can change definitions without your involvement, which is exactly why pinning matters more there. Neither is safer in the abstract. It depends on whether your team can actually operate what it hosts.

What leaders should ask for

If you are accountable for the outcome rather than the implementation, four questions separate a controlled rollout from a hopeful one.

  1. Who can change a tool definition, and how would we know? If the answer is “the vendor, and we wouldn’t,” you have an unmanaged dependency.
  2. Can each agent call only the tools its job requires? Server-level access is a yes to the wrong question.
  3. What is the worst thing one agent can do if its context is hijacked? Map private data, untrusted input and outbound channels for each one.
  4. Can we reconstruct, after the fact, which definition and which approval were in force for a given call? This is what auditors and incident responders will ask first.

The cost of getting this right is modest compared with the cost of a data leak that occurred through a “trusted” integration, and far cheaper to build before agents hold write access than after.

Start here

You do not need the full gateway on day one. In order of value:

  • Inventory every MCP server in use, including the ones developers added to their own configs.
  • Pin definitions and alert on any change.
  • Cut each agent down to the tools it needs.
  • Remove one leg of the private data, untrusted input, outbound channel triangle for each agent.
  • Log every call with the definition hash.

Adoption made the connection easy. The job now is making the text that rides along with every connection something you decided to trust, rather than something you forgot to check.

Enjoying this insight?

Join the distribution list to get deep dives on AI transitions and agency economics directly in your inbox. No spam, ever.

Back to Blog

Related Posts

View All Posts »