Skip to main content

Your MCP Deploy Tool's dry_run Flag Is a Suggestion, Not a Lock

10 min readDora NodaDora Noda
Share
On this page

On August 6, 2026, a discussion in modelcontextprotocol/modelcontextprotocol documented a real incident and then generalized it into something every infrastructure-tool author should sit with: MCP tool annotations — readOnlyHint, destructiveHint, and friends — are advisory metadata for the client, not enforcement. Nothing in the protocol stops a caller from supplying its own value for a parameter the server considers its own. The annotation said read-only. The client sent its own value anyway. The server, having no rule against it, obeyed.

Here is the thesis before the pattern: if your deploy tool's safety depends on the caller respecting a hint, you have documentation, not a control. This post gives you the control — a three-line server-side authority pattern with a worked sketch — plus why "the annotation said it was read-only" is not a defense you can put in an audit log for infrastructure someone else owns.

The three-line rule: derive, reject, log​

For an infrastructure MCP server, four parameters are the whole ballgame: dry_run, target_environment, tenant_id, and confirm. These are exactly the parameters an agent must not be able to set for itself — and exactly the ones a client can supply, because the input schema is the client's to fill. The fix is three rules, applied on the server, in this order:

  1. Derive the sensitive fields from the authenticated session, never from the tool arguments.
  2. Reject rather than merge client-supplied values for those fields — a conflicting value is an error, not an input.
  3. Log the attempted override with the session identity attached, so the audit trail shows the attempt, not just the outcome.

Concretely, each of the four dangerous parameters has exactly one legitimate source:

ParameterLegitimate sourceWhat to do with a client-supplied value
dry_runServer policy: write-class tools default to dry-run unless the session holds an execution grantReject the call — the caller doesn't get to vote itself out of the sandbox
target_environmentAuthenticated session scope (token claims, deployment binding)Reject on mismatch — production in the args means nothing if the session is scoped to staging
tenant_idToken claims (tenant_id claim, organization scope) — implicit from auth, never a paramReject — no header, URL segment, or argument may select a tenant
confirmA server-issued confirmation token from a prior preview call, bound to the sessionReject unknown or replayed tokens; a bare confirm: true from the client is not confirmation

The sketch below is the whole pattern in about twenty lines. Note what it pointedly does not do: it never reads the sensitive fields from args for any purpose other than detecting the override.

python
async def deploy_service(args, session):
    # 1. DERIVE: sensitive fields come from the session, never the client.
    tenant_id = session.claims["tenant_id"]
    environment = session.scope.environment
    execution_granted = session.grants.contains("deploy:execute")
 
    # 2. REJECT: a client-supplied value for a server-owned field is an error.
    for field in ("dry_run", "target_environment", "tenant_id", "confirm"):
        if field in args:
            # 3. LOG: record the attempt with the session identity attached.
            audit.warn("override_attempt", field=field,
                       supplied=args[field], session=session.id)
            raise ToolError(f"'{field}' is server-authoritative and "
                            f"must not be supplied by the client.")
 
    if not execution_granted:
        return preview_plan(tenant_id, environment, args)  # dry-run by default
    return await execute_plan(tenant_id, environment, args,
                              receipt_to=session.id)

Three properties make this a control rather than a wish. First, the override is impossible by construction for the honest-but-confused client: there is no code path where a client value flows into the authoritative field. Second, the dishonest or compromised client gets an error, not a merge — silent precedence rules ("server wins on conflict") hide attacks, while rejection surfaces them. Third, the log entry binds the attempt to the session, which is what turns the incident from "the model did something" into an attributable event with an identity, a timestamp, and a rejected value.


Why annotations can't do this job​

To be fair to the protocol, annotations were never supposed to do this job. The MCP tools specification defines ToolAnnotations — readOnlyHint, destructiveHint, idempotentHint, openWorldHint — as self-reported hints the server advertises so the client can make UX decisions: show a confirmation prompt, gate auto-approval, label the tool in a picker. They are, as one server author put it, "advisory hints for MCP clients, not security enforcement." Another put the corollary in its README: the hints are advisory, so the server-side confirmation gate and spend caps apply regardless of what any client does with them.

Three details make the gap concrete rather than theoretical:

The spec itself distrusts them. The tools specification says clients must consider tool annotations untrusted unless they come from trusted servers. That single sentence is the protocol telling you the threat model out loud: any client talking to any server it hasn't allowlisted is expected to treat readOnlyHint: true as a claim from a stranger. Building a safety case on a claim the spec instructs the other side to distrust is building on sand.

Mislabeled tools are already endemic. A 2026 survey of 508 MCP servers surfaced 8,286 findings, including 612 "privilege annotation mismatches" — tools labeled readOnlyHint: true that can produce side effects when chained. The failure mode is worse than a missing label: clients use the read-only hint to gate auto-approval, so a mislabeled tool may execute without confirmation on a trusted server. The hint doesn't merely fail to protect; it actively downgrades the friction the client would otherwise apply. Every one of those 612 findings is a tool whose author presumably believed the annotation was doing something.

The dangerous direction is client-to-server, and annotations only speak server-to-client. Annotations flow outward in tools/list: the server describing itself to the client. But the dry_run override flows inward in tools/call: the client supplying arguments to the server. No metadata the server publishes about itself can constrain what the next tools/call contains. That asymmetry is why the August discussion generalizes beyond one incident — it names a structural fact about the protocol's information flow, not a bug in any single implementation.

The companion thread that same week asked the structural version of the question outright: where should deterministic host-authority decisions and receipts fit in MCP? Annotations answer "how should the client present this tool." Nobody had specified where the host's decisions — this session may execute, that one may only preview — live, or what receipt proves the decision happened. Which brings us to the audit log.


The structural question: decisions need a home, and receipts need a ledger​

The August threads converge on a framing worth adopting verbatim: MCP is a projection of authority your system already has; it creates no new authority. Your deploy API, your RBAC, your tenant isolation — those exist before MCP enters the picture, and the MCP server is a new doorway into the same rooms. Every gate on the old doorways (admin auth, tenant isolation, dry-run defaults, approval-gated commands, ledger receipts) must be preserved on the new one. "Agent-initiated sends are propose-only; a human approves" is a property of the underlying system that the MCP surface inherits — or, if nobody wires it through, silently drops.

That framing settles two design arguments that otherwise eat weeks:

Elicitation is UX, not enforcement. MCP's elicitation/create lets the server ask the client — and thereby the human — a question mid-call. It is the right tool for "which region did you mean?" and the wrong tool for "are you authorized to touch production?" The spec's own security considerations assign confirmation prompting to clients (clients SHOULD prompt; servers SHOULD NOT) and name destructiveHint as the canonical signal — which routes us straight back to the advisory metadata we just established can't carry weight. Elicitation asks the requester to confirm the requester's own request. Enforcement asks the session what the policy allows. Those are different questions with different answerers, and only the second one belongs in an audit log.

Receipts are pre-action authority records, not post-hoc logs. The host-authority thread's lasting contribution is the idea that the interesting record is created before the action: this principal, holding these grants, was authorized for this operation in this scope, at this time — with the tool result and any revocation pointing back to it. A post-hoc "deploy happened" log entry answers what. A pre-action authority record answers why it was allowed, which is the question an auditor, an incident review, and a customer whose infrastructure you operate all actually ask. When your MCP server derives tenant_id from token claims and binds the confirmation token to the session, the receipt writes itself: the decision inputs are already server-side facts.

The punchline for operators: if your infrastructure MCP server can't produce, for any given tool call, the answer to "which session, which grants, which scope — and what did the client try to supply that you rejected," you don't have an MCP security posture. You have annotations and hope.


The pattern in the wild​

This isn't theoretical hygiene. The derive/reject/log shape is already what careful operators converge on, in three independently arrived-at forms:

Dry-run by default. Several production servers make write-class tools default to no-op: the tool returns a preview of what would happen unless the session already holds an execution grant. One Kubernetes MCP server splits the surface explicitly — a _preview tool that returns the dry-run diff and a separate apply-gated tool that refuses to run without operator approval, with the standing instruction that the agent must never claim a change is applied without the approval receipt. Another requires an explicit server-side confirmation token rather than a client confirm: true, so "but the agent said confirm" is never the reason something executed. The common thread: the default path is safe, and the unsafe path requires something the client cannot mint for itself.

Two-stage confirmation tokens. The stronger variant binds confirmation to a specific preview: the preview call returns a token encoding what was previewed, and the execute call must present that token back. A bare confirm: true in the arguments is meaningless — rejected, per the table above — because confirmation isn't a boolean the client asserts; it's a capability the server issued for one specific action. Replay the token against different arguments and it fails closed. This is derive/reject/log with the derivation step made cryptographic.

Tenant-from-token, structurally. The cleanest implementations don't merely validate a client-supplied tenant — they never accept one at all. Tenant is implicit from the token: the organization scope is read from the authenticated identity, and no header, URL parameter, or tool argument can select a different tenant. As one operator's docs put it, this makes "never trust model-supplied tenant IDs" structural rather than procedural — there is simply no code path where a model-supplied tenant flows anywhere. The Rust MCP security guidance shows the same shape: extract tenant_id from token claims, fail the call if the claim is missing, and never consult the arguments for it.

Notice what all three share with the August incident: each closes exactly the inward tools/call channel that annotations, facing outward, cannot touch. The protocol gives you the doorway; the authority stays yours to enforce.


What changes for operators​

If you run an MCP server that touches infrastructure — deploys, DNS, secrets, tenants — the August 2026 threads hand you a short punch list. Enumerate every parameter your tools accept that a client must not control, and check each against the three-line rule: derived from the session, rejected when client-supplied, logged on attempt. Audit your annotations separately, as UX copy: fix the readOnlyHint labels that lie (the survey averages more than one mismatch per server), but stop treating the fix as a security control.

And make sure every execution leaves a pre-action authority record — session, grants, scope — that the result points back to, because "the annotation said it was read-only" will not survive first contact with an incident review, let alone a customer audit.

The protocol is honest about what annotations are. The incident happened because a server believed they were something else. Derive, reject, log — and let the hints go back to being hints.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agents are first-class operators there, which is exactly why server-side authority matters: every deploy decision lands in an audit log tied to a session, not a hint. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide