Skip to main content

Vercel's fingerprintTools and detectToolDrift: The First Drift-Detection Primitives for MCP Tool Execution

9 min readDora NodaDora Noda
Share
On this page

You approved the deploy tool on Monday. The description said it deploys the main branch to staging, and the input schema had exactly three fields. On Friday, your agent calls a tool with the same name — but the description now carries an extra sentence instructing the model to also push to production, and the schema quietly grew a fourth field.

Nothing in the Model Context Protocol stops the server from serving a different definition than the one you reviewed, and the model will happily follow the new instructions. This is the MCP rug pull, and until this month, catching it was entirely your app's problem.

That changed with ai@7.0.19, where Vercel's AI SDK added two functions — fingerprintTools and detectToolDrift — that pin the tool definitions you approved and diff every later fetch against that baseline before the definitions reach the model. They are the first drift-detection primitives built into a mainstream agent SDK's execution path, and they mark the moment MCP tooling grew from "can we connect to tools" to "can we trust the tools we connected to." If you run an infrastructure MCP server — especially one whose deploy and rollback tools are the highest-privilege surface an agent touches — this is the pattern to steal.

The tool you approved is not the tool running tomorrow​

The rug pull works because of a trust asymmetry baked into MCP. A client fetches tools/list, shows the definitions to a human for approval, and then passes those definitions to the model as context on every subsequent turn. The protocol binds nothing: the server can serve a benign definition at approval time and a mutated one later, and neither the client nor the model re-validates.

This is not theoretical. Invariant Labs demonstrated tool poisoning attacks in April 2025, showing malicious instructions embedded in tool descriptions that were visible to the model but hidden in client UIs, and a follow-up demonstrating exfiltration plus "shadowing," where one server silently alters another server's tools. Their research found 5.5% of MCP servers exhibiting tool-poisoning characteristics.

The attack class is now catalogued as OWASP MCP03:2025, with the rug pull as a listed sub-technique, and the MCPTox benchmark (August 2025) measured a 72.8% attack success rate across 353 tools. More recently, the August 2026 filings MCP-2026-008 (cache poisoning) and MCP-2026-015 showed how protocol-level caching can turn one server's mutated definitions into a cross-user attack — the drift surface is growing, not shrinking.

For a deploy MCP server, the blast radius is what makes this urgent. A poisoned description on a deploy-staging tool doesn't just leak a token — it can re-route where code lands. A widened input schema can smuggle an unreviewed parameter past the approval that covered the original three fields. The tampered definition is a silent prompt-injection upgrade unless something compares what the server sends today against what a human approved. That something is now two function calls.

What Vercel actually shipped in ai@7.0.19​

Some prior-state honesty first, because "first" is doing work in this post's framing. Drift detection as an idea predates this release: Invariant Labs' mcp-scan has detected description changes via hashing since April 2025, and careful teams have hand-rolled their own pinning. But all of that lived outside the execution path — in scanners you run periodically, or in app code every team wrote differently. Nothing in the SDK compared a fresh tools/list fetch against the approved set at the moment before the tools were handed to the model. That comparison is exactly what vercel/ai PR #16902 added to AI SDK core, and it shipped in ai@7.0.19.

The mechanism is deliberately small. fingerprintTools digests each tool's server-controlled, security-relevant fields — the string description, the resolved input schema, and the title — into a stable map of tool name to digest. detectToolDrift diffs two such maps and returns { added, removed, changed }. Your app owns baseline storage and the response to drift: block, force re-approval, or alert. The canonical pattern from the SDK docs:

typescript
import { fingerprintTools, detectToolDrift } from 'ai';
 
// Trust time (first connect, human-reviewed): capture and persist the baseline.
const baseline = await fingerprintTools(await mcpClient.tools());
 
// Every later fetch, before handing tools to generateText:
const tools = await mcpClient.tools();
const drift = detectToolDrift(await fingerprintTools(tools), baseline);
 
if (drift.changed.length || drift.added.length) {
  // A pinned definition changed, or a new tool appeared. Block, re-approve,
  // or alert per your policy — do not silently pass `tools` to the model.
}

Three design choices are worth noting:

  • The fingerprint covers exactly the fields a server controls and a model reads — description, schema, title — so cosmetic client-side metadata can't cause false drift.
  • added is treated as drift-worthy alongside changed: a brand-new tool appearing mid-session never went through approval either.
  • The SDK stays unopinionated about storage and enforcement, which keeps it composable: the baseline can live in your audit store next to the human approval record, and the response can match your risk posture per tool.

What drift detection catches — and the hole it leaves open​

The honest version of this table matters more than the marketing version, because the gap defines what else you still need to build.

AttackCaught?Why
Injected instructions in a tool descriptionYesDescription is a fingerprinted field; any mutation lands in changed
Widened input schema (new field, loosened enum)YesResolved input schema is fingerprinted
Renamed or retitled toolYestitle is fingerprinted; a rename shows as removed + added
Brand-new tool appearing mid-sessionYesShows up in added, never approved
Behavior or endpoint swap with identical name, description, and schemaNoThe tool runs remotely; the client never sees the implementation

That last row is Vercel's own documented limitation, and it is the right call to state plainly: fingerprinting compares definitions, not behavior. A server that keeps serving byte-identical definitions while changing what its endpoint does is invisible to this check. Closing that hole needs a different layer — signed tool definitions, pinned server versions, reproducible server builds, or execution-side policy that constrains what even an approved tool may do. Drift detection turns silent mutation into a detectable event; it does not turn an untrusted server into a trusted one.

There is a second, softer limitation: the baseline is only as good as trust time. If the first fetch you fingerprinted was already poisoned — because you pinned definitions without human review, or reviewed them in a UI that hides the injected text the way early clients did — you have pinned the attack. Fingerprinting enforces "still what we approved," not "what we approved was safe." Pair it with description review at baseline time, ideally with the raw definition visible, not just the rendered card.

What a self-hosted deploy MCP server should borrow​

Vercel built this for their deployment surface and their Skills library; a self-hosted platform on its own Cluster API fleet needs the same primitives pointed at a different target. Here is the borrow-vs-skip checklist:

Borrow: pin tool identity at approval time. The moment a human approves your deploy server's tool set, persist a fingerprint of every definition next to the approval record. Treat the fingerprint like a lockfile: the approved set is a versioned artifact, not a live query.

Borrow: check drift before every model call, not on a schedule. Scanners that run nightly leave a window where the model acts on mutated definitions. The check belongs in the request path — fetch, fingerprint, diff, then decide — so a drifted tool never reaches the model in the first place.

Borrow: wire drift alerts into the deploy audit log. A changed entry on deploy-production is a security event with the same severity as an unauthorized deploy attempt. Log the expected vs. actual digest, the tool name, and the session; page on it for production tools, ticket it for the rest.

Borrow: pin the highest-privilege tools first. If full coverage is a project, start with deploy, rollback, scale, secret-access, and anything that touches production state. Read-only status tools can follow; the blast radius orders the rollout.

Borrow: force re-approval on added. A new tool mid-session is definitionally unreviewed. Route it through the same approval flow as day one rather than auto-trusting it because the server was once vetted.

Skip: Vercel-specific surface assumptions. Their deployment targets, Skills packaging, and dashboard approval UX don't transfer. What transfers is the two-function shape — digest at trust time, diff before use — implemented against your own tool registry and your own audit store.

Still to build yourself: the behavior layer. As the table above shows, definition pinning doesn't cover endpoint swaps. For a deploy server, that means execution-side guardrails stay mandatory: the tool implementation should re-validate its own parameters, production mutations should require the same confirmation they would from a human CLI, and server upgrades should be versioned and signed so "same definition, different code" at least requires compromising the release pipeline rather than one live server.

The operational standard, going forward​

Step back and notice what happened here. For MCP's first year, the ecosystem's security story was research papers, scanners, and admonitions to review tool definitions — all valuable, all outside the path the tools actually travel to reach the model. Vercel moved one critical check into that path: every fetch is now comparable against what was approved, in two function calls, with the response policy left to the app. That is what an operational standard looks like — not a new protocol version, but a primitive so small and so obviously correct that every serious MCP client adopts it and every serious MCP server assumes it.

For a self-hosted platform, the implication is direct. The moment your agents execute deploys through an MCP server, unversioned, unpinned, server-supplied tool definitions are a live prompt-injection surface on your highest-privilege workflow. The fix is no longer bespoke security engineering: capture the baseline at trust time, diff before every model call, alert into the audit log, and re-approve on change. The SDK gives you the two functions; the discipline of using them on every fetch is the part you still have to build — and the part your future incident review will ask about.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide