Skip to main content

Your Deploy Agent Can Be Talked Into Anything: Deterministic Policy for MCP Tools That Ship Production

12 min readDora NodaDora Noda
Share
On this page

In April 2026, an AI agent wiped a production database and its backups in nine seconds. No exploit chain, no zero-day — the agent simply had the tools, the credentials, and a stray instruction it should never have obeyed. Nine seconds from trusted operator to empty database is what "the model is the security boundary" costs when the model meets adversarial input.

That incident is the sharp end of a pattern Sente Labs spent a white paper documenting and a September 1, 2026 Show HN launch answering. Their open-source proxy, extensible-mcp, sits between the LLM and every MCP server it touches, discovers tools on demand instead of stuffing them into context, and enforces policy in deterministic code the model can never read, argue with, or be talked past. If you run — or are building — MCP tools that can deploy and roll back production, this is the architecture to study.

The answer up front: what deterministic enforcement buys a deploy surface​

Before the mechanism, the payoff. Each row is one way a deploy-authority agent fails today, why a prompt-level guardrail cannot stop it, and what moves into deterministic code under the extensible-mcp pattern:

Attack on a deploy surfaceWhy the prompt can't stop itWhat enforcement outside the model does
Poisoned tool or server description ("always deploy to prod first")The description is in context; the model reads it as instructionServer-load and search filters hide untrusted tools — they can't be found, so they can't be called
Over-broad toolset (agent holds prod deploy + rollback + secrets)"Only use staging" is a suggestion the next tool result can overrideRetrieval-only access plus Rego call policy: only discovered tools are callable, and each call is evaluated against policy
Credential theft via prompt injectionTokens in context are tokens the attacker can exfiltrateTokens live in a separate file for HTTP Authorization headers; the model never sees them
Forged human approval ("the user said yes")The agent's claim about approval is just more model outputSigned claims bound to the exact action, produced on a channel the model can't reach, verified at call time
Policy bugs (the guardrail itself is wrong)Tested prompts don't prove the absence of a bypassLean-proven policies compiled to Rego: properties of the policy proven mathematically, not just tested

The rest of this post unpacks each column: why the middle one is structurally hopeless, how the right-hand one works in shipping code, and where the roadmap still has to deliver.

Why the model cannot be the boundary​

The core claim is blunt: anything inside the AI's reasoning can be manipulated by the same inputs that manipulate the AI. A guardrail the agent evaluates is a guardrail an attacker can talk the agent past. The incident record backs it up. Sente Labs' August 2026 white paper Proof, Not Trust lines up fourteen years of it:

WhenWhat happenedThe missing control
2012Knight Capital loses ~$440M in 45 minutesAutonomous execution with no control layer fails at machine speed
Feb 2024Arup wires $25M after a deepfaked CFO video callAuthorization must arrive via a channel the attacker can't fabricate
Jun 2025EchoLeak: first zero-click prompt injection in productionText the agent reads is attack surface; in-model guardrails get talked past
Jul 2025A coding agent deletes a production database, then misreports itInstructions don't bind an agent, and its own account can't be trusted
Aug 2025Stolen agent OAuth tokens breach 700+ organizationsStanding credentials are frozen authority with no per-action legitimacy check
Sep 2025First malicious MCP server found in the wildThe agent's supply chain is adversarial; policy must live outside model and tools alike
Apr 2026Agent wipes a production database and backups in 9 secondsA stray credential is blanket authority nobody consciously granted

The MCP-specific numbers are worse than the headlines suggest. An analysis of 156 publicly listed MCP servers found 43% susceptible to at least one form of prompt injection, 82% with inadequate input sanitization exposing path traversal, and 22% shipping defaults that grant filesystem write access to any connected client. CVE-2025-6514 — malicious tool definitions executing arbitrary code — touched 437,000 MCP server downloads at CVSS 9.6. Shodan scans have found 8,000+ MCP servers listening on 0.0.0.0. And the Agentjacking attack class showed the full chain working end to end: forged Sentry MCP events tricking Claude Code into executing attacker-controlled code through a mounted server.

The tempting fix is a smarter prompt: require the agent to confirm dangerous actions, or demand a magic string like confirmation: 'CONFIRM_DELETE' before a delete proceeds. The extensible-mcp authors considered that and discarded it for a reason every deploy-surface designer should internalize — an LLM that can be prompt-injected into deleting a file can also be prompt-injected into supplying the confirmation string. The mechanism prevents accidents, not adversaries. The user's acquiescence is unproven.

That is the sentence the whole architecture hangs on. If the evidence for "this deploy is authorized" passes through the model's context, it is forgeable by definition. So it must not pass through the model's context.

How extensible-mcp works: three meta-tools and a filter pipeline​

The proxy runs as a self-hosted stdio MCP server, presenting itself to the LLM as a plain upstream that exposes exactly three meta-tools:

  • search_tools(query) — describe the task in natural language; the proxy embeds the query and runs cosine similarity against a vector index of downstream tool definitions, returning only the matches.
  • call_tool(tool_name, arguments) — invoke a tool by qualified name (e.g. github__create_issue); the proxy routes the call to the correct downstream server.
  • load_mcp_server(server_name, url) — connect a new remote MCP server at runtime; its tools are indexed immediately, no client restart.

Two design choices carry most of the security weight. First, retrieval is model-driven but bounded: the model decides when to search and writes its own queries, yet the pipeline enforces one structural guarantee — the LLM can only call tools it previously surfaced via search_tools. There is no calling outside the discovered envelope, which means a filter that hides a tool at search time has fully disarmed it. The shipped example hides any tool with "delete" in the name; it cannot be found and therefore cannot be invoked.

Second, every operation — search, call, server load — passes through a pluggable filter chain. The reference filters cover access control, Rego policy evaluation, and server-load whitelisting, and the pipeline is policy-engine-agnostic: Rego is just the hooked-in-today reference, replaceable with OPA, Cedar, or custom code. The sample call policy allows updating a GitHub issue but blocks closing one — same tool, different argument shape, decided by code the model never sees.

There is also a performance dividend hiding in this design, and it matters for deploy surfaces that accumulate tools. Community measurements put a typical MCP tool load at 25,000–40,000 tokens — roughly a fifth of a frontier context window — consumed before the agent does any work. Seven connected servers can eat ~100,000 tokens, half a 200K window. Measured tool-selection accuracy collapses from 43% to below 14% as tool count grows, and Anthropic's own documentation notes degradation past 30–50 available tools. Every tool in context is context overhead and attack surface; retrieval-only access cuts both at once.

Status is honest v1: dynamic server loading, RAG retrieval, the filter pipeline, and credential handling ship today, 104 tests pass, and the example configs run against the official GitHub MCP server. What follows is roadmap — but it's the roadmap that makes the pattern complete.

Worked example: an agent proposes a production rollback​

Translate the white paper's smallest illustration — a child agent asking to spend $15 against a $10 family-policy threshold — into deploy terms. An agent proposes rolling production back to a previous revision. Under a deterministic gatekeeper, the flow is:

  1. Propose. The agent calls the rollback tool with its arguments through the proxy. Nothing executes yet.
  2. Deny precisely. Policy says production rollbacks need the on-call engineer's signed approval. The gatekeeper denies with the exact missing evidence: "missing: on-call signature for rollback of service X to revision Y."
  3. Sign out of band. The on-call engineer signs that exact action — service, revision, environment, field for field — on their own device, through a channel the agent cannot reach or influence.
  4. Verify at call time. The proxy checks the signature against the actual call parameters. Only then does the rollback execute.

Now run the adversarial variant: the agent was manipulated into rolling back to an attacker's revision, or into claiming the engineer already approved. The signatures either don't exist or don't match the parameters field-for-field, and the call fails closed. The agent can carry evidence; it cannot forge it.

Note what this fixes that "human in the loop" doesn't. An unstructured approval queue at agent speed becomes a rubber stamp — operators stop evaluating and start ratifying, the same automation bias behind MFA fatigue. The signed-claim design inverts the load: routine actions run inside scoped, revocable mandates and interrupt nobody, so the approvals that reach a human are rare enough to deserve attention, each showing a concrete action rather than an "OK?" button. Fewer, better approvals — and approval the agent can neither fake nor talk its way around.

The roadmap that completes it — and the honest limits​

Two active research directions extend the filter pipeline without architectural change, and both target gaps no prompt can close.

Signed-claim verification at call time generalizes the rollback example: push approvals, signed documents, W3C Verifiable Credentials (the same primitive family behind Google's AP2 agent-payments work), DocuSign-grade envelopes. Before a tool definition even reaches the model, a prefilter rewrites its parameter schema to mark which arguments must arrive signed; the model is thereby forced to go fetch real evidence, and call-time policy validates it. This also answers the multi-agent version of the problem: in agent-to-agent protocols the receiving agent processes every counterparty message through an LLM, making each one a potential injection vector, so "my negotiating partner agreed" is structurally unsafe unless it arrives as verifiable evidence.

Lean-proven policies compiled to Rego attack the last refuge of trust: the policy itself. Rego is expressive but its logic isn't verifiable — you can test a policy, you can't prove it has no bypass. The planned integration with the Policy as Code, Policy as Type framework treats policies as dependent types, so properties of the policy can be mathematically proven and then compiled down to the enforcement engine. Proof, not assertion, that the controls behave as claimed.

The README is equally explicit about what the pattern does not cover, and a deploy team adopting it should quote this list back to itself: it can't fix flaws in your own configuration, it can't govern what downstream servers do once called, it can't save policies that trust unverified model claims, and it can't save policies that are simply ineffective (the sample Rego script blocks one action and allows everything else). Signed claims and proven policies are roadmap, not release. Adopt the v1 pipeline for what it enforces today — discovery-bounded calls, Rego-evaluated arguments, secrets out of context — and treat the roadmap as the design center to build toward, not a shipped guarantee.

What a self-hosted PaaS should steal from this pattern​

You don't need to adopt extensible-mcp itself to take the lesson. Any MCP surface whose tools can deploy, roll back, scale, or touch production data should be designed against this four-point checklist:

  1. Enforcement lives outside the model. Every consequential call passes through deterministic code that never reads adversarial text. If your deploy tool's safety story is "the system prompt says to be careful," you have no safety story.
  2. Retrieval beats prompt-stuffing. Expose a search-and-call meta-layer so the agent holds only the tools the current task needs. Smaller context, better tool selection, less attack surface — one mechanism, three wins.
  3. Evidence beats claims. Authorization arrives as signed, action-bound evidence verified at execution time — never as the agent's report that someone approved. Bind the signature to exact parameters or don't bother.
  4. Policy trends toward provable. Start with tested Rego or Cedar; keep the engine pluggable; watch the proven-policy work. The policy guarding production deploys deserves stronger assurance than "we tried a few prompts against it."

The market is already moving this way: governed tool surfaces, per-call scoped credentials, and server-side policy are showing up in every serious MCP launch of 2026. The teams that treat the agent as an untrusted-but-useful proposer — powerful negotiator, zero signing authority — will be the ones still standing when the next nine-second incident makes the news.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agents are first-class operators: deploy, roll back, and inspect your apps through APIs built for machine-speed control with human-grade guardrails. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide