The scariest sentence in AI infrastructure this year is five words long: "the agent can deploy now." Somewhere between the demo where a chatbot restarts a staging service and the incident where it restarts the wrong one, a team discovers that giving an agent a deploy tool is really giving it an identity problem, a scoping problem, and an audit problem wearing a trench coat. The good news is that the Model Context Protocol's July 28, 2026 specification — its largest revision since launch — finally gives you the pieces to solve all three properly.
Here is the production design in one box. The rest of this post is the evidence for why each line is there.
- Discover, don't configure: the deploy MCP server publishes OAuth protected-resource metadata so clients find the correct authorization server automatically — or fail closed.
- Audience-bind every token: a token minted for staging is rejected by production, even when the signature is valid.
- Scope narrowly, enforce twice: staging can execute, production can only propose until a human approves — checked at the transport layer and again at each tool.
- Approve destructive tools out of band: deploys and rollbacks to production return "input required" and wait for a human decision with a TTL. Silence means deny.
- Audit the chain, not the secrets: every decision records the human, the agent, the tool call, and the approval — never the token.
This is a design for a remote MCP server allowed to deploy or roll back apps: not a local stdio helper with your shell's permissions, but a network service an agent reaches over Streamable HTTP — the shape that actually needs authorization, and the shape the 2026-07-28 spec was rewritten around.
What the 2026-07-28 spec actually changed
The July revision deleted the protocol's session state: no more initialize handshake, no Mcp-Session-Id header, list endpoints no longer varying per connection. Every request carries its protocol version and client capabilities in _meta, servers must implement server/discover, and version mismatches return UnsupportedProtocolVersionError. Each request is now independently routable — Streamable HTTP POSTs carry required Mcp-Method and Mcp-Name headers so gateways route without parsing bodies.
Authorization got hardening rather than a rewrite, and that distinction matters. Two of the most important controls — publishing protected-resource metadata under RFC 9728 so clients discover the right authorization server, and binding tokens to one resource server with RFC 8707 resource indicators — have been mandatory since the 2025-06-18 revision, over a year earlier. As one enterprise implementation writeup put it, those two requirements already retired most of the confused-deputy patchwork teams had written in application code. What July adds on top of that foundation:
| Area | Before 2026-07-28 | After |
|---|---|---|
| Protocol core | initialize handshake + Mcp-Session-Id sessions | Stateless; per-request _meta, mandatory server/discover |
| Request routing | Body parsing by intermediaries | Required Mcp-Method / Mcp-Name headers |
| Client registration | Dynamic Client Registration (DCR) | DCR formally deprecated in favor of Client ID Metadata Documents (CIMD) |
| Authorization responses | No issuer binding | Authorization servers SHOULD send RFC 9207 iss; clients MUST validate it when present |
| Stored credentials | Reused across servers in practice | MUST be keyed by issuer, MUST NOT be reused with a different authorization server |
| Long-running work | Experimental tasks in core | Tasks moved to an official extension with tasks/get polling and tasks/update input |
| Server-initiated requests | sampling/createMessage, elicitation/create, etc. | Replaced by the Multi Round-Trip Request pattern (input_required results) |
| Unchanged foundation | — | RFC 9728 + RFC 8707 mandatory since 2025-06-18; OAuth 2.1 with mandatory PKCE |
Two rows in that table carry the whole deploy-from-chat design. CIMD replacing DCR means a client identifies itself with an HTTPS URL serving its own metadata — your fleet's authorization server decides which agent clients it trusts before any token exists. And the Multi Round-Trip Request pattern is the spec-native shape of an approval gate: a server that needs something returns input_required, and the client retries with the answer. A deploy tool that needs a human is exactly that shape.
One caveat: authorization is still formally OPTIONAL in MCP — the spec only says that once you protect an HTTP transport, OAuth 2.1 is how. Skipping that step is how you land in scans like MCPInspect's, which found 833 vulnerable MCP servers in the wild, 18 with suspicious or misleading tool descriptions suggestive of deliberate behavior-poisoning. This design treats "protected" as non-negotiable.
Discovery: find the right authorization server or fail closed
Every secure session with a remote MCP server starts the same way: the client POSTs without a token, the server answers 401 with a WWW-Authenticate challenge pointing at its protected-resource metadata, the client fetches that metadata, reads the authorization_servers list, then fetches the authorization server's own metadata to find the token endpoint. Only then does the PKCE-protected authorization flow run, with the resource parameter set to the MCP server's canonical URI so the resulting token is audience-bound to exactly that server.
For a deploy server, get three things right here:
- Publish the metadata. The server's
/.well-known/oauth-protected-resourcedocument names the authorization server. Clients that cannot discover it must fail closed — never fall back to a token endpoint from config, a chat message, or (worst) the agent's own suggestion. Metadata URLs are also an SSRF vector, so clients should fetch them over HTTPS only and refuse private or link-local addresses. - Bind the audience per environment. Staging and production are different protected resources with different canonical URIs, even if one binary serves both. A staging token presented to production is rejected even though the signature verifies — that one check turns "one leaked token" from a fleet-wide incident into an environment-scoped one.
- Validate the issuer on the way back. The RFC 9207
issparameter must match the recorded issuer before the code is redeemed. Combined with issuer-keyed client credentials, this closes the trick of redeeming one authorization server's code at another.
None of this is MCP-specific cleverness; it is ordinary OAuth hygiene applied where most agent demos skip it. The NSA's 2026 MCP guidance makes the gap explicit: servers rely on Bearer tokens as defined in OAuth 2.1, but the core protocol does not mandate token lifecycle management — expiration, rotation, and reuse control are your architecture's job, not the spec's gift.
Scopes: a token for staging must not deploy production
Once tokens reach the right server, scopes decide what they can do. The rule for a deploy server is that scopes name the environment and the authority separately, and that the sensitive half of each pair requires a human:
| Scope | Allows | Human approval? |
|---|---|---|
deploy:read | Status, history, logs, pending proposals | No |
deploy:staging:execute | Deploy and roll back staging directly | No |
deploy:staging:rollback:execute | Roll back staging directly | No |
deploy:production:propose | Create a production deploy plan, dry-run it | No |
deploy:production:approve | Approve one pending proposal within its TTL | Yes — this scope IS the human |
deploy:production:rollback:propose | Create a production rollback plan | No |
deploy:production:rollback:breakglass | Execute a production rollback without prior approval | Counted as approval after the fact; short TTL, paged review |
Three properties make this table work. First, there is no deploy:production:execute scope — the agent cannot hold a credential that deploys to production unilaterally, because no such credential exists. Second, the break-glass rollback scope exists because waiting for a human can be worse than acting: a failed deploy at 3 AM should be reversible in seconds, but the token expires in minutes and its use pages for retroactive review. Third, each environment's scopes ride on that environment's audience-bound token, so binding and scoping multiply rather than overlap.
Enforce scopes twice. The transport layer checks that the token is valid, current, audience-bound, and carries the baseline scope. Each tool then re-checks the specific scope its action needs — rollback_production demands break-glass or an approved proposal regardless of what the transport accepted. Per-call checks are shipping practice in mid-2026, which leaves no excuse for "one valid token opens every tool."
The approval gate: destructive tools wait for a human
Here is where the 2026-07-28 spec pays for itself. Before July, asking a human mid-flow meant the deprecated elicitation request or an out-of-band hack. Now the deploy tool follows the Multi Round-Trip Request pattern: the agent calls deploy_production, the server validates everything it can, and instead of deploying returns input_required naming the approval it needs. The client surfaces that to the human, the human approves (minting a short-lived approval bound to the exact proposal hash), and the client retries carrying the approval. No session state, no handshake — the approval is just data on the retried request.
The gate needs four rules to be trustworthy:
- Approve the plan, not the intent. The approval binds to a hash of the exact deploy plan — image digest, target environment, migration list. If the agent regenerates the plan after approval, the hash changes and the approval no longer matches. "Yes, deploy" without a pinned artifact is how prompt-injected scope creep ships.
- Silence is deny. Every approval carries a TTL in minutes, and expiry denies by default. This is the load-bearing rule for unattended agents: an approval queue nobody is watching must converge to "nothing happened," never to "it went ahead."
- Pre-authorize the rollback path, not the deploy. The healthy default is that approving a production deploy also pre-authorizes its automatic rollback on failed health checks within a bounded window — the rollback is the safer action and should never wait on a human who just approved the thing being undone. Unplanned rollbacks outside that window use the break-glass scope and get reviewed after.
- Track long deploys as tasks, not connections. A production rollout outlives any HTTP request. The Tasks extension exists for exactly this: the server returns a task handle, the client polls
tasks/get, and further input flows throughtasks/update. Combined with the stateless core, a deploy can survive client restarts, gateway failovers, and load-balancer re-routes that would have killed a session-pinned flow in the old protocol.
The autonomy ladder behind these rules: read-only tools run free, staging executes free, production proposes free, and anything touching production state waits for a named human or a break-glass token with a TTL. Operators converge on the same boundary — move an agent to read-only or approval-required mode before debating anything else — because live permissions change the risk immediately while policy debates do not.
Audit logs: the human/agent/tool-call chain
When something ships at 3 AM, the question is never "did we log." It is "can we reconstruct exactly who decided what, in which order, with which credential." A deploy MCP server should emit one structured audit event per authorization decision — allowed or denied — with at least these fields:
{
"event": "deploy.authorization_decision",
"decision": "allowed",
"timestamp": "2026-10-02T03:14:00Z",
"human": "oncall@example.com",
"agent": {"client_id": "https://agents.example.com/deploy-bot", "model": "example-model-2026-09"},
"tool": "deploy_production",
"tool_args_digest": "sha256:9f2c…",
"proposal_digest": "sha256:41ab…",
"approval_id": "apr_8f3e…",
"token_issuer": "https://auth.example.com",
"token_audience": "https://mcp.example.com/production",
"scope_presented": "deploy:production:approve",
"trace_id": "4bf92f…"
}Present: the full chain from human through agent and model to tool call and approval, plus the token's issuer, audience, and scope so a reviewer can verify the authorization without trusting it. The trace_id rides the spec's OpenTelemetry _meta propagation, joining the audit event to the chat turn that started it. Absent: the token, any secret, and full prompts or arguments — digests prove what was approved without turning the audit log into a credential store. Denied decisions log with the same shape: a denied deploy_production from an unexpected agent is the incident you want to find in the morning.
Route these events somewhere append-only with months of retention, and alert on what matters: break-glass usage, repeated denials from one client, approvals expiring faster than humans can review them. Forensic visibility you fail to record during the deploy is evidence you cannot reconstruct after it.
Walkthrough: Friday deploy from chat
A safe deploy-from-chat session under this design:
- "Ship the new checkout flow." The agent calls
server/discover, lists tools with its staging-audience token, and drafts a plan: build image, migrate staging database, deploy to staging. - Staging goes ahead. The staging token carries
deploy:staging:execute, so the agent deploys and smoke-tests without bothering anyone. Each call emits an audit event naming the agent and the staging audience. - Production becomes a proposal. The agent calls
deploy_productionwith the staging token; the server rejects it on audience alone. The agent re-authenticates against the production resource, gets a token with onlydeploy:production:propose, and submits the plan for a proposal ID and digest. - The human decides. The tool returns
input_required. The on-call engineer sees the image digest, migration list, and staging results, and approves — bound to the proposal digest with a 15-minute TTL. - The deploy runs as a task. The client retries with the approval; the server starts the rollout and returns a task handle. The agent polls
tasks/getwhile the human watches the same handle in a dashboard. Mid-rollout health checks fail. - Rollback needs no permission slip. The deploy approval pre-authorized automatic rollback, so the server rolls back immediately and logs the chain: approval ID, health-check failure, rollback execution. The engineer reviews it over coffee.
Every step uses a mechanism already introduced above. Nothing requires trusting the agent's judgment about authority — only its competence at the mechanics, with a human holding every production decision not strictly safer than waiting.
Ship checklist
Before your deploy MCP server touches a real fleet, verify these ten:
- Protected-resource metadata is published; undiscoverable servers fail closed.
- Staging and production are separate resources; cross-audience tokens are rejected.
- Authorization responses validate
iss; client credentials are keyed by issuer. - No
deploy:production:executescope exists; production mutation requires approval or break-glass. - Every tool re-checks its required scope per call; transport checks alone are insufficient.
- Approvals bind to proposal digests with TTLs; silence denies.
- Deploy approvals pre-authorize bounded automatic rollback; unplanned rollbacks use break-glass with paged review.
- Long operations return task handles; no deploy depends on a live connection.
- Every decision emits a chain-complete audit event with digests, never secrets.
- Denials and break-glass usage alert a human; retention covers your incident-review window.
The through-line of the 2026-07-28 revision: MCP grew from a local-tool protocol into a network protocol — stateless requests, routable headers, explicit task state, authorization that assumes hostile networks. A deploy tool is the sharpest thing most teams will hand an agent. Build the five lines in the box at the top, and "the agent can deploy now" stops being the scariest sentence in your infrastructure.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



