Your deploy bot's memory is about to be deleted — by the protocol, on purpose. On July 28, 2026, the Model Context Protocol shipped its largest revision since Anthropic open-sourced it: the initialize handshake is gone, the Mcp-Session-Id header is gone, and every request must now carry enough context to land on any server replica behind a plain round-robin load balancer. If your MCP server holds a multi-step deploy in session memory, that design just became a bug.
This is not a deprecation notice. It is the H2 2026 roadmap arriving all at once: stateless servers, machine-readable Server Cards for discovery, and agent-to-agent coordination through A2A. With 97 million monthly SDK downloads and more than 10,000 public servers in production, MCP is no longer a chat-plugin protocol — it is infrastructure. And infrastructure has different rules. This post works through what the three changes demand from one concrete server: a deploy-from-chat MCP server whose tools push code, tail logs, and roll back production.
The short version, for the impatient: thread explicit handles through your tools instead of session memory, declare your blast radius in a Server Card before strangers call you, and let A2A do the coordinating while MCP does the tool-calling. The rest of this post is the receipt for each clause.
What stateless actually takes away
The 2026-07-28 spec makes MCP stateless at the protocol layer. The handshake (initialize / notifications/initialized) is removed; each request self-describes with its protocol version, client capabilities, and optional client identity in _meta. Protocol-level sessions are removed with it, so there is no session to pin, migrate, or lose on deploy. Server-initiated requests — the mechanism behind elicitation prompts and sampling callbacks — are replaced by the Multi Round-Trip Requests pattern, where the client retries with more information instead of the server calling back.
Lists gain cacheable ttlMs hints, and a new mandatory server/discover RPC lets clients interrogate capabilities without ceremony.
For a deploy server, the before/after contract looks like this:
| The old assumption | What replaces it |
|---|---|
| Session holds the build-debug context across calls | Server-minted handles passed as ordinary tool arguments; durable state lives outside the connection |
| Server asks the user a follow-up mid-tool-call (elicitation) | Multi Round-Trip Requests: the tool returns "need X," the client retries with X supplied |
| Sticky routing keeps one client on one replica | Per-request _meta makes every request self-contained; any replica serves any call |
| Clients learn tools by connecting and listing | server/discover plus Server Cards advertise capabilities before first contact |
The scalability win is real — no session affinity, no 404-when-the-replica-dies, no live sessions breaking mid-conversation on every deploy. The .NET SDK already flips stateless to the default, and warns on stateful-only holdouts. But notice what the table's first row quietly demands: you now own the state the protocol used to hold for you. That is where deploy servers break.
The build-debug loop, redesigned
Consider the canonical deploy-from-chat loop: the agent calls deploy, the build fails, the agent calls get_logs, reads the error, calls fix_and_redeploy. On the old protocol, the server could stash the attempt history — which commit, which build id, how many retries — in the session and pick it up on each call. The tool signatures stayed clean because the session was the hidden argument.
Stateless operation deletes the hidden argument. The fix is to make it explicit: deploy returns a deployment handle, and every subsequent call threads it back in — get_logs(handle), redeploy(handle, fix). The handle is a server-minted, opaque token; the durable state (attempt count, build artifacts, rollback point) lives in your own store keyed by that handle, not in transport memory. Any replica can serve get_logs(handle) because the handle, not the connection, carries the continuity.
This is strictly better engineering — explicit state survives deploys, scales horizontally, and can be audited — but it has one gotcha the session model never forced you to think about: handles outlive conversations. A session died when the chat ended; a handle persists until you expire it. Your design now needs an expiry and revocation story: handles that time out after the deploy settles, handles a human can revoke, and a get_status(handle) that says "unknown or expired" instead of resurrecting a week-old rollback point. If your handle has no TTL, you have not removed session state — you have moved it to a database with no eviction policy.
There is a second, subtler casualty: mid-call human approval. The old elicitation flow let a server pause a promote_to_production call and ask the user "are you sure?" inside the same tool invocation. Under Multi Round-Trip Requests, that becomes a structured round trip — the tool returns a machine-readable "approval required" response, and the client retries with the approval attached. Gateways are already standardizing this as requireApproval tool metadata with human-in-the-loop pauses. The approval gate still exists; it just moves from an ambient server callback into the tool's declared contract, where clients, audit logs, and policy engines can all see it.
What your Server Card must confess
Statelessness changes how clients call you. Server Cards change who calls you: SEP-1649 proposes machine-readable manifests at .well-known endpoints so registries, crawlers, and unfamiliar agents can discover what your server does, which transports it speaks, and what auth it requires — without ever establishing a connection. The Server Card working group (Anthropic and GitHub leads, chartered March 2026) is standardizing the format, and the July spec's mandatory server/discover RPC is the in-band half of the same idea: capabilities you can read before you commit.
Discovery-without-connection is wonderful for adoption and terrifying for a server whose tools touch production. The moment a strange agent can find your deploy server from a Card, the Card's declarations stop being documentation and start being a security boundary. Before your deploy tools are discoverable, your Card needs to confess at least four things:
- OAuth scopes, per tool. Not "this server uses OAuth 2.1" but which scope each tool demands —
deploy:stagingversusdeploy:productionare different trust decisions, and the July spec's tightened OAuth profile (RFC 9207 issuer validation, Client ID Metadata) exists precisely so clients can verify them. - Blast radius. Which environments, clusters, and tenants this server can touch. An agent deciding between two deploy servers needs "can roll back production in eu-west" in the Card, not buried in a README.
- Human-approval gates. Which tools pause for a human and which do not.
get_logsshould never require approval;promote_to_productionalways must. If the gate is only enforced server-side at call time, the discovering agent cannot plan around it. - Tool allowlist posture. Whether unknown clients get the full tool list or a reduced one — the gateway pattern of allowlisting which tools reach the agent belongs in the discovery layer, not as a runtime surprise.
None of this is hypothetical hygiene. Ten thousand public servers plus crawler-driven discovery equals agents calling tools their operators have never seen. The servers that survive that world are the ones whose Cards let a stranger answer "should I trust this with my production deploy?" before the first tool call, not after the first incident.
MCP calls tools; A2A coordinates agents
The third roadmap pillar is the easiest to misunderstand, because it looks like a competitor. Google's Agent-to-Agent protocol (A2A) standardizes how independent agents discover each other, delegate work, and exchange results — Agent Cards, task delegation, multi-step coordination across vendors. MCP standardizes how one agent calls tools. These are different layers, and production multi-agent systems already run both: MCP is the agent-to-system interface, A2A is the agent-to-agent one.
For deploy infrastructure, the layering rule is concrete. Picture a release coordinator agent that plans the rollout talking over A2A to a deploy agent that actually holds your MCP deploy tools. The coordinator delegates ("roll out v2.4 to staging, then production") without needing your credentials; the deploy agent executes tool calls against your server, where the OAuth scopes, blast radius, and approval gates from the previous section are enforced. Coordination state — who approved what, which step is next — lives in the A2A task exchange, not in your MCP session (which no longer exists) and not in any single chat transcript.
The category error to avoid is treating MCP as the coordination layer: chaining tool calls through conversation context and hoping the session holds the plan together. That was always fragile; under a stateless spec it is simply unavailable. The plan belongs in the agent layer, expressed in A2A tasks one agent can hand to another; your server's job is to make each tool call independently verifiable — explicit handles, declared approvals, durable state. Both protocols are converging under Linux Foundation stewardship (MCP was donated in December 2025; IBM's rival agent protocol merged into A2A in early 2026), which means this layering is the bet the ecosystem is standardizing on, not one vendor's opinion.
The deploy-server checklist
Everything above compresses into one checklist. Each row has a pass/fail test — run them against your server this week:
| # | Requirement | Pass test |
|---|---|---|
| 1 | No tool reads session memory | Kill any replica mid-loop; retry the same call with the same handle on another replica and get the same answer |
| 2 | Multi-step flows thread explicit handles | get_logs and redeploy take a handle argument; there is no "current deploy" ambient state |
| 3 | Handles expire and revoke | Every handle has a TTL; an expired handle returns "unknown or expired," never stale state |
| 4 | Approvals are declared, not ambient | promote_to_production advertises its human gate; the MRTR round trip is machine-readable and auditable |
| 5 | Server Card declares scopes, blast radius, gates | A stranger can answer "may I call this against production?" from the Card alone |
| 6 | Coordination lives outside MCP | The rollout plan survives in A2A task state (or your own store) with no MCP connection open |
Items 1–3 are the stateless migration. Items 4–5 are the discovery hardening. Item 6 is the architecture decision that makes the other five stick.
From chat plugin to infrastructure
Step back and the direction is unmistakable. Stateless requests, crawler-readable Cards, and a separate coordination protocol are what a system looks like when it stops being a clever ChatGPT add-on and starts being load-balanced, discovered, and delegated-to infrastructure — the same trajectory REST walked fifteen years ago. The servers that treat this as a migration chore will resent every row of the checklist. The servers that treat it as a promotion will notice what it buys: horizontal scale for free, discovery by strangers as a feature, and multi-agent rollouts without bespoke glue.
The window is open now because the spec is fresh — July 2026 — and most of the 10,000 public servers have not migrated yet. Redesign the tool contract while your deploy agent is still the only caller, and you inherit the ecosystem's scale and discovery instead of bolting them on after the first incident.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. AI agents are first-class operators: they deploy and manage apps through a Render-compatible API, which is exactly the kind of tool contract this post argues for. Star the repo on GitHub or deploy your first app today.



