← All posts Engineering

MCP Tool Allowlist: Stop Tool Drift Before Agents Use It

<article> <p><strong>For platform engineers and security engineers responsible for MCP governance, the answer is direct: an <em>MCP tool allowlist</em> should approve specific servers and specific tool contracts, then fail closed when a tool appears or changes without review.</strong> Treat the MCP tool list as a supply-chain surface, not a developer preference. Snapshot <code>tools/list</code>, hash descriptions and schemas, review drift in CI, and enforce policy outside the model before <code>tools/call</code>.</p>

<p>The rollout problem usually starts innocently. A few teams connect coding agents to GitHub, Jira, or a filesystem server. The demos work. Developers ask for more tools. Security asks who approved them. Platform asks which repos can use them. Nobody has a precise answer because the control was installed at the server level, while the risk lives at the tool level.</p>

<p>MCP makes that distinction important. The protocol lets servers expose model-controlled tools with a <code>name</code>, natural-language <code>description</code>, <code>inputSchema</code>, optional <code>outputSchema</code>, and annotations. The same tool metadata becomes part of the agent's operating context. If a server adds a new write tool, changes a schema, or modifies a description with hidden instructions, the agent may receive a different authority surface than the one your team reviewed.</p>

<p>This is the same trust boundary covered in <a href="/guides/mcp-security-for-ai-agents">MCP security for AI agents</a>: tool metadata and tool output are part of the agent attack surface.</p>

<h2>Why an MCP Tool Allowlist Is More Than Server Approval</h2>

<p>A server allowlist answers one question: can this MCP server be connected? An MCP tool allowlist answers the question platform teams actually need for production: which exact tools, with which exact contracts, may be visible and callable in this environment?</p>

<table> <thead> <tr> <th>Control</th> <th>What it covers</th> <th>What it misses</th> </tr> </thead> <tbody> <tr> <td>Server allowlist</td> <td>Approved server command, URL, namespace, or product integration</td> <td>New tools, changed descriptions, changed schemas, and excessive scopes inside an approved server</td> </tr> <tr> <td>Tool allowlist</td> <td>Approved tool names and expected tool contracts for a server</td> <td>Handler-only behavior drift and malicious output from tools that still match the contract</td> </tr> <tr> <td>Runtime authorization</td> <td>Whether a specific user, repo, data class, and workflow may call the tool now</td> <td>Unreviewed metadata if discovery and CI controls are absent</td> </tr> </tbody> </table>

<p>The trade-off is friction. A strict allowlist slows down the first week of adoption because every useful tool needs a review path. The alternative is quieter and riskier: a trusted server can silently widen the model's options after rollout. For platform teams, that is the wrong default. MCP's own specification describes tools as arbitrary code execution, and the Tools spec says applications should show exposed tools, show invocations, allow denial, and prompt for confirmation.</p>

<p>Those user-facing controls matter, but they are not sufficient as the only governance layer. GitHub Copilot repository MCP docs warn that configured MCP tools may be used autonomously without approval and recommend allowlisting specific read-only tools. Cursor Enterprise supports MCP server allowlists and per-server tool allowlists, with an empty tool allowlist meaning all tools from that server. Anthropic's MCP connector supports enabling all tools, allowlisting specific tools, or denylisting unwanted tools at the request/API layer. The common pattern is clear: product control planes are moving toward explicit tool admission.</p>

<h2>The Rollout Story: From Useful Demo to Governed Platform</h2>

<p>Consider a platform team rolling MCP into engineering workspaces. In the pilot, the team allows one repository assistant to connect to a small set of MCP servers. The first decision is whether to optimize for developer speed or governance consistency.</p>

<p>Scenario A is open discovery. Developers can connect arbitrary servers and see whatever tools the server advertises. It is fast, but it creates an inventory problem on day one. OWASP's MCP Tool Poisoning guidance is blunt on this point: do not let users connect to arbitrary servers.</p>

<p>Scenario B is deny-by-default. Developers request a server, platform approves the server identity, security reviews the exposed tools, and CI records the approved tool contracts. It takes longer to onboard the first integrations, but it gives the team a stable base for drift detection, audit, and emergency revocation.</p>

<table> <thead> <tr> <th>Rollout choice</th> <th>Advantage</th> <th>Trade-off</th> </tr> </thead> <tbody> <tr> <td>Allow all tools from approved servers</td> <td>Fast onboarding and fewer support tickets during pilots</td> <td>Server approval becomes implicit approval for future tools and changed contracts</td> </tr> <tr> <td>Allow specific tools from approved servers</td> <td>Clear least-privilege boundary for each repo and workflow</td> <td>Requires review workflow, metadata snapshots, and exception handling</td> </tr> <tr> <td>Gateway or client policy before calls</td> <td>Central place for telemetry, denial, and kill switches</td> <td>Needs integration across clients whose MCP controls differ</td> </tr> </tbody> </table>

<p>The practical compromise is to make read-only tools easy to approve and privileged tools deliberately harder. A read-only repository tool and a privileged write tool should not share the same approval path. NSA guidance recommends tool segregation by data classification, sandboxing, egress and DLP controls, invocation validation, and logging. That maps naturally to a tiered rollout: read, write, privileged write, external egress, and destructive operations.</p>

<p>Privileged tool tiers should inherit the same containment assumptions as <a href="/guides/ai-agent-sandboxing">AI agent sandboxing</a>, especially for filesystem, shell, and network access.</p>

<h2>Tool Drift Is the Failure Mode to Design Around</h2>

<p>MCP tool lists can change dynamically. Servers advertise <code>tools.listChanged</code>, clients call <code>tools/list</code>, and servers may send <code>notifications/tools/list_changed</code> when available tools change. That is useful protocol machinery, but it does not decide whether the new or changed tool is safe for your environment.</p>

<p>For governance, drift means any unapproved change in the tool contract or capability surface. That includes a new tool name, removed tool, changed description, changed <code>inputSchema</code>, changed <code>outputSchema</code>, changed annotations, changed server identity, changed package digest, or changed protocol version. Practitioner and vendor writeups in the brief converge on the same snapshot and fingerprint pattern: compare tool names, schemas, descriptions, output schemas, and capabilities over time.</p>

<p>Do not treat drift as downtime. A drifted server may be healthy, reachable, and returning valid JSON-RPC responses. The failure is more subtle: the agent is now reasoning over a contract your team did not approve.</p>

<h2>Build an MCP Lockfile for Tool Contracts</h2>

<p>There is no universal MCP lockfile format in the brief, so do not wait for one. Create a local policy artifact that records the reviewed state of each approved server and tool. The exact file format matters less than making it versioned, reviewable, and enforced.</p>

<p>A useful lockfile record should include:</p>

<ul> <li>Server identifier: URL or command, plus registry namespace where applicable.</li> <li>Server package or artifact digest where your deployment model supports it.</li> <li>MCP protocol version used during approval.</li> <li>Tool name scoped to that server.</li> <li>Tool description hash.</li> <li><code>inputSchema</code> hash.</li> <li><code>outputSchema</code> hash when present.</li> <li>Annotations hash, treated as untrusted unless the server is trusted.</li> <li>Approved risk tier, owner, review date, and allowed environments.</li> </ul>

<p>The identity problem is real. MCP tool names are unique within a server, not globally. A strong local identity usually needs the server URL or command, namespace, package digest, protocol version, tool name, schema hash, and description hash. That is heavier than a simple list of tool names, but it avoids the most dangerous ambiguity: two tools with the same name and different authority.</p>

<p>The review process should be boring by design. A developer proposes a new MCP server or tool change. CI runs <code>tools/list</code> against the configured server. The generated snapshot is compared with policy. Unknown tools, changed schemas, or changed descriptions fail the check. A platform or security owner reviews the diff and approves only the needed tools.</p>

<p>Treat approved MCP servers like dependencies: the same review habits from <a href="/guides/coding-agent-dependency-security">coding agent dependency security</a> apply to tool contracts and package identity.</p>

<h2>Detect Tool Poisoning Before It Reaches the Model</h2>

<p>OWASP's MCP Top 10 lists Tool Poisoning as MCP03, including rug pulls, schema poisoning, and tool shadowing. The mechanism is different from a traditional package exploit. A malicious or compromised MCP server can influence the agent through the text and schemas the agent sees.</p>

<table> <thead> <tr> <th>Risk</th> <th>What changes</th> <th>Control</th> </tr> </thead> <tbody> <tr> <td>Rug pull</td> <td>A previously acceptable tool changes behavior or contract after approval</td> <td>Snapshot approved contracts, hash metadata, and fail closed on drift</td> </tr> <tr> <td>Schema poisoning</td> <td>Input or output schema steers the agent toward unsafe arguments or assumptions</td> <td>Review schemas as security-relevant code, not documentation</td> </tr> <tr> <td>Tool shadowing</td> <td>A tool imitates or influences another trusted capability</td> <td>Scope identity to server plus tool contract, and isolate privileged tools</td> </tr> <tr> <td>Malicious results</td> <td>Tool output contains instructions or data intended to steer later agent behavior</td> <td>Treat tool results as untrusted and enforce policy outside the model</td> </tr> </tbody> </table>

<p>Description scanning can help, but the brief is clear about the gap: scanning will have false negatives. Use scanners as one signal, not as the trust boundary. The stronger control is to prevent unreviewed metadata from entering the agent context and to validate every sensitive invocation before execution.</p>

<h2>Enforce in CI and at Runtime</h2>

<p>CI gives you change control. Runtime policy gives you blast-radius control. You need both because tool drift can arrive through configuration changes, server updates, dynamic discovery, or live list-change notifications.</p>

<p>In CI, check MCP configuration changes the same way you check dependency and infrastructure changes. The check should fail when a repository adds an unapproved server, uses malformed allowlist JSON, exposes a tool not present in the approved policy, or changes a tool hash without review. GitHub Enterprise managed settings provide a useful product example: <code>allowedMcpServers</code> and <code>deniedMcpServers</code> evaluate deny before allow, block non-matching servers when an allowlist exists, and fail closed on malformed allowlist JSON except built-in defaults.</p>

<p>At runtime, enforce before <code>tools/call</code>. The policy decision should include server identity, tool identity, user, repo, environment, data classification, requested arguments, and operation risk. For sensitive operations, require explicit confirmation. For privileged tools, use isolation, sandboxing, egress controls, DLP checks, output limits, and logging. NSA's May 2026 guidance specifically warns that dynamic tool discovery should be treated cautiously unless paired with origin verification or authorization checks.</p>

<p>Runtime denials and list-change events should feed the same operating loop as <a href="/guides/mcp-server-monitoring">MCP server monitoring</a>.</p>

<p>A gateway can centralize those decisions across clients. Client-side controls still matter because developers live in GitHub, Cursor, Claude Code, Codex, and other hosts. But product control planes differ, and cross-tool policy normalization remains an open engineering problem. A gateway gives platform teams one place for telemetry, denial reasons, kill switches, and SIEM events.</p>

<h2>What to Log for Audit and Drift Response</h2>

<p>Logging should let an operator answer three questions without reading the full agent transcript: what tools were exposed, what changed, and what was called. Microsoft's field report frames MCP governance as an inventory plus gateway problem with reviews, drift checks at connection, telemetry, and future policy-as-code. That is the right operating model for larger organizations.</p>

<p>At minimum, log server identity, protocol version, tool list snapshot hash, tool name, approved contract hash, runtime contract hash, user, repo, environment, arguments after redaction, approval result, policy decision, output summary, error state, and whether the operation changed data or reached external systems.</p>

<p>For drift incidents, keep the response path short. Disable the changed tool, revoke or deny the server if needed, preserve the observed snapshot, identify which sessions saw the changed tool list, and review any calls made after drift first appeared. <code>notifications/tools/list_changed</code> helps live sessions notice changes, but it does not catch handler-only behavior drift, compromised package releases, or clients that cache tool lists incorrectly.</p>

<h2>A Practical MCP Tool Allowlist Checklist</h2>

<ul> <li>Start deny-by-default for MCP servers and tools.</li> <li>Approve servers by URL, command, namespace, and artifact identity where possible.</li> <li>Approve tools by server-scoped name plus hashed description and schemas.</li> <li>Snapshot <code>tools/list</code> during review and on connection.</li> <li>Fail CI on unknown servers, unknown tools, or changed tool contracts.</li> <li>Separate read-only, write, privileged, external egress, and destructive tools.</li> <li>Require explicit confirmation for sensitive operations.</li> <li>Treat tool metadata, annotations, and results as untrusted unless policy says otherwise.</li> <li>Enforce authorization outside the model before <code>tools/call</code>.</li> <li>Send tool exposure, drift, calls, denials, and errors to audit logs or SIEM.</li> <li>Maintain tool owners, review dates, approved environments, and kill switches.</li> </ul>

<h2>The Bottom Line</h2>

<p>An MCP rollout fails quietly when server approval turns into unlimited tool approval. The better model is narrower: admit servers deliberately, expose only reviewed tools, record the exact contracts, and treat drift as a security event.</p>

<p>That creates some operational cost. Teams need a review queue, a lockfile or policy artifact, CI checks, runtime enforcement, and audit plumbing. The payoff is concrete: when an MCP server adds a new capability or changes the instructions the model sees, your platform catches it before the agent can act on it.</p> </article>

Get started

Deploy your fleet.

Put a fleet of sandboxed agents to work on your own infrastructure, provisioned in seconds and watched live from one console.

Get started →

Admin-provisioned · Self-host in one command · Your data never leaves your VM