A coding agent is debugging a failing CI pipeline at 2 AM. It reads the logs, finds a plausible fix, and — because its token can — pushes the fix straight through to production. Total model cost: eighty cents. Business impact: a production outage nobody approved (the constructed incident is fictional; the failure shape is not).
The uncomfortable part is that nothing malfunctioned. The agent did exactly what its credentials permitted. A broadly scoped deploy token held by an agent is standing permission to touch production, and every production action the agent takes spends down trust that was granted once, in bulk, at setup time. The July 2026 MCP specification release — revision 2026-07-28, the largest since authorization landed — gives infrastructure tools a protocol-native way to stop spending that trust blindly: Multi Round-Trip Requests (MRTR), which let a server pause mid-call and ask a human to confirm the concrete cost, target, or consequence before acting.
The gated flow this enables is four steps, and the rest of this post is the receipt for each one:
- The agent calls a deploy tool with an intent; the server builds a rollout plan and answers
input_required, showing the exact diff, target, and blast radius. - A human sees an approval card bound to that plan's digest — not a blank "are you sure?"
- The client retries the same call carrying the approval plus the server's opaque
requestState. - The server verifies the digest, the single-use binding, and the expiry — then executes exactly once and returns an auditable receipt.
No standing permission is exercised anywhere in that loop. The token gets the agent to the gate; only a human opens it, and only for the plan they actually read.
What the July 2026 MCP release changed
To see why this is new, compare it with what elicitation looked like before. Elicitation — a server asking the user for input mid-tool-call — arrived in the 2025-11-25 protocol version as a server-initiated elicitation/create request pushed back down an open stream to the client. That works when client and server share a long-lived session. It falls apart for remote servers behind gateways and load balancers, where there is no stable stream to push down and no session for the server to hold state in. One team's blunt postmortem: interactive confirmation "proved unable to round-trip through real agent stacks" and was removed.
Revision 2026-07-28 (changelog) redesigns the protocol core as stateless: the initialize handshake and the Mcp-Session-Id header are gone, every request self-describes its protocol version and capabilities in _meta, and servers expose a mandatory server/discover method. Cross-call state, where it exists, travels as server-minted handles passed back as ordinary arguments. Elicitation, sampling, and roots — every flow that previously required the server to call back into the client — were replaced by a single pattern: the server returns an interim result asking for what it needs, and the client retries the original request with the answers attached.
Concretely, when a tools/call handler needs input it now returns resultType: "input_required" carrying typed inputRequests and an opaque requestState. The client gathers the answers and retries the identical call with inputResponses plus the echoed requestState. No session, no held-open stream, no affinity between the two requests — the retry can land on any worker behind the gateway. Sampling and roots were not even folded into the new pattern; they were deprecated outright, with elicitation-via-MRTR as the surviving way for a server to ask for something mid-call. As one migration guide puts it, a handler that needs input throws InputRequired, the call answers input_required, and the client retries the same request.
That inversion — the server never calls back, it just asks and waits for a retry — is what makes approval gates deployable on real remote infrastructure. The gate no longer depends on transport behavior the operator cannot control.
The gate, step by step
Here is the deploy/rollback flow, with each step saying what crosses the wire and what binds it.
Step 1 — the agent prepares a plan, the server prices it. The agent calls something like deploy with an intent: service, revision, environment. The server does not execute. It resolves the intent into a concrete rollout plan — the image digest, the diff against what is running, the target cluster and namespace, the expected blast radius (how many replicas restart, whether the migration runs before or after the cutover), and the estimated cost or consequence. It returns input_required with that plan rendered as an elicitation form and a digest of the plan sealed inside requestState.
The agent cannot edit the plan after this point; it can only present it.
Step 2 — a human approves the exact plan. The client surfaces an approval card: what will change, where, and what it costs. The human's "approve" is captured as the elicitation answer. Critically, the approval names the plan digest, not the tool. "Yes, deploy api at sha:9f3c… to prod, 12 replicas rolling, migration 014 first" is a fundamentally different object from "yes, the agent may deploy things." The first is a single-use capability; the second is standing permission wearing a confirmation dialog as a costume.
Step 3 — the client retries with the approval attached. The client re-issues the original tools/call with inputResponses (the approval) and the echoed requestState. From the protocol's view this is just a retried request; from the security view it is the only request that can execute, because it is the only one carrying proof a human saw the plan.
Step 4 — the server verifies, executes once, and receipts. Verification is a checklist, and every item matters:
- Digest match. Recompute the plan hash from the sealed
requestStateand confirm the approval references it. If the plan changed between elicitation and retry — a new commit landed, the migration list moved — the digest mismatches and the call fails closed. The agent must start over with a fresh plan. - Single use. The
requestStatehandle is consumed on first execution. A replayed retry, a double-clicked approval card, or a retried webhook delivers "already consumed," not a second deploy. - Expiry. The handle carries a short TTL — minutes, not hours. An approval granted against a plan written at 2 AM should not still be spendable after the morning's three follow-up commits.
- Audit. Execution returns a receipt: who approved, which digest, when the handle was minted and consumed. The receipt is the artifact your incident review reads, not the agent's chat transcript.
This shape is already appearing in the wild. The RailCall 2026 workflow entry for a production deploy approval gate runs exactly this spine — merged PR plus green CI, then a dry-run blast-radius plan, then a human approval, then the deploy trigger with a signed, offline-verifiable receipt. DEVA, a local-first DevOps agent, puts the human in exactly one place — Approve and Merge — and never grants itself write access to the target. Zscaler's MCP server adopted the 2026-07-28 revision so that delete confirmation asks a person through elicitation, with anything other than an explicit delete aborting before the API is touched. Different tools, same invariant: the dangerous call cannot complete on token authority alone.
Four ways teams get this wrong
These are the failure modes to design against. Each one turns the gate back into decoration.
1. Approving without a digest. The most common mistake is eliciting a bare yes/no — "deploy to production?" — with no plan attached. Between approval and execution the world moves: the agent rebases, a teammate pushes, the migration set changes. Without a digest binding the approval to the exact plan, the gate suffers a time-of-check/time-of-use gap and the human approved something different from what ran. The fix is structural, not procedural: the elicitation payload must include the plan, and the server must re-verify the digest at execution time. If your approval card could be screenshotted and reused against a different plan, it is not bound.
2. No expiry or replay protection. An approval that never expires and can be replayed is standing permission with extra steps. This is where the requestState design earns its keep: because the handle is server-minted and opaque to the client, the server can seal the digest, a nonce, and a timestamp inside it (signed, so clients cannot mint their own), then enforce single-use and TTL on retry. Statelessness does not mean memorylessness — it means the memory travels with the request instead of living in a session. Teams that treat requestState as an echo field and skip the verification checklist get the latency of two round trips with the security of zero gates.
3. Using elicitation to collect secrets. Form elicitation must never ask for passwords, API keys, or tokens — the spec and every serious implementation guide say so plainly, because elicitation answers flow through the agent, the client, and logs you do not control. If your deploy flow needs a credential to proceed, that is an OAuth flow (the 2026-07-28 revision also hardened authorization with RFC 9207 issuer validation and Client ID Metadata Documents), not a form field. A gate that asks "paste the production database password to approve" has converted your approval UX into a credential-harvesting pipeline.
4. No fallback when the client cannot elicit. Not every client supports form elicitation — support varies across harnesses, and headless runners may have no human to ask at all. A server that returns input_required into a void either hangs the automation or, worse, tempts the operator to add a "skip approval in CI" flag that becomes the default. Backblaze's b2-mcp shows the disciplined alternative: clients with compatible elicitation get the human prompt, and everything else falls back to an explicit confirm: true retry — a deliberate, auditable second call, not a silent bypass. Design the fallback before you need it, and make the fallback louder than the gate, not quieter.
And keep the enforcement server-side regardless: client-side approval UX is convenience, but the server refusing an unapproved write is the actual control — defense in depth, with the tool as the last line, not the first.
There is a broader pattern underneath all four: the gate must live where the authority lives. An agent-side "are you sure?" prompt can be prompt-injected away; a token-scope check at the gateway cannot re-ask about consequences. The server holding the production credential is the only party that can both see the concrete plan and refuse to run it, which is why MRTR puts the question-asking power there.
What this means for self-hosted platforms
Step back from the protocol details and the operational picture is simple: approval UX just became portable infrastructure. Before MRTR, a human gate on agent actions meant bespoke per-client plumbing — a Slack bot here, a CLI prompt there, a web console modal somewhere else — each wired to a session the server had to babysit. After MRTR, the gate is a property of the tool: any client that speaks the retry flow gets the same approval card, the same digest binding, the same expiry, because the server defines all of it. Build the gate once, in the MCP server that fronts your deploy pipeline, and every present and future client inherits it.
That consolidation lands at a moment when the numbers say gates are the norm, not the exception. In the first large production-agents study, 68% of deployed agents execute ten or fewer steps before a human intervenes — autonomy in production is treated as a risk surface to bound, not a goal to maximize. The platforms absorbing that lesson fastest are the ones making the bound structural: propose in chat, approve on a card, execute once, keep the receipt. ChatOps grew up doing exactly this for human operators; MRTR extends the same discipline to agents, with the protocol — not each team's Slack bot — carrying the handshake.
For a self-hosted PaaS, the implication is concrete. Your deploy pipeline already has the plan (the diff, the target, the migration list); your MCP server is the natural place to seal it into requestState and ask. The clients your tenants use — chat harnesses, IDE agents, headless runners with explicit-confirm fallback — all speak the same retry. And because the gate is stateless, it survives the realities of self-hosted operations: gateway restarts, rolling worker upgrades, and approval cards answered minutes later on a phone, long after any session would have died.
The industry spent 2025 handing agents the keys and learning, incident by incident, what standing permission costs. The 2026-07-28 revision is the protocol catching up to that lesson: authority should be granted per plan, per human, per minute — not per token, per quarter, per hope. Build the gate into the tool, bind the approval to the digest, expire it fast, and keep the receipt. Your 2 AM self will thank you.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with agent-operable deploys behind a Render-compatible API. Star the repo on GitHub or deploy your first app today.



