Only 4% of live MCP endpoints speak the new stateless wire. A wire census published October 2, 2026 probed 186 public MCP endpoints twice on the same cohort: 7 answered the 2026-07-28 revision, 65 answered older stateful revisions, 86 sat behind auth, and 28 errored out (Pennyforge). Of the 72 endpoints that returned any protocol version at all, just 9.7% spoke the new wire — and three of the seven shared a single operator.
The local side is even further behind. The same research probed ten prominent npm MCP servers and found every answering one — including the official server-filesystem and server-everything reference packages — still speaking only the 2025-06-18 wire and rejecting the new server/discover call with "Method not found" (Pennyforge). The flagship packages are, in the author's words, "the oldest wire in the ecosystem."
So the MCP ecosystem is currently speaking five revisions at once, and most of them are invisible to each other without an error. If you operate MCP servers — especially a deploy-from-chat surface where an agent drives deploys and rollbacks — you don't get to live in the clean stateless future yet. You get to live in the transition window. This post is the runbook for it: how to route each traffic class today, how to migrate tools that genuinely need cross-call state, and why anything you build new should ship stateless-first.
What the 2026-07-28 spec actually removed
The July 28, 2026 revision is the largest protocol break since MCP launched: the initialize handshake is gone (spec proposal SEP-2575), the Mcp-Session-Id header and protocol-level sessions are gone (SEP-2567), and every request now carries its own protocol version and client capabilities in _meta instead of negotiating them once per conversation (MCP Stateless Revolution). The release candidate locked May 21, 2026 and the final revision shipped July 28. September's release coverage summed up the payoff in one line: "Your MCP client can speak to a load balancer that connects with any server" (September 2026 release coverage).
The rewrite goes further than sessions. Capability probing moves to a new server/discover RPC, version mismatches get an explicit UnsupportedProtocolVersionError, a single subscriptions/listen POST stream replaces the old HTTP GET plus resources/subscribe pair, ping, logging/setLevel, and roots/list_changed are removed, long-lived Tasks graduate to the official io.modelcontextprotocol/tasks extension, and multi-request/multi-response (MRTR) replaces server-initiated requests (Pennyforge changelog summary).
Two companion changes matter most for operators. First, SEP-2243 requires Streamable HTTP requests to carry Mcp-Method and Mcp-Name headers, so a gateway can route, throttle, and meter MCP traffic on headers without parsing JSON-RPC bodies (MCP Stateless Revolution). Second, authentication is now per request: there is no session in which to establish identity once and trust it for the rest of the conversation, so a revoked credential takes effect on the very next call — and your gateway performs an identity check on every call, a real latency cost to budget when sizing it.
The transition routing table: three traffic classes, three rules
Here is the core runbook. Classify every MCP conversation hitting your ingress into one of three buckets and route accordingly:
| Traffic class | How to recognize it | Routing rule |
|---|---|---|
| Spec-current (2026-07-28) | Version + capabilities in _meta; Mcp-Method header present; no Mcp-Session-Id | Plain round-robin load balancer. No stickiness, no session store. Route tasks/* methods to a differently-sized pool by header if you run long-lived work. |
| Legacy stateful (2025-11-25 and older) | initialize handshake; Mcp-Session-Id header on follow-ups | Keep pinned/sticky routing alive on the old path. The ALB stickiness annotations or ingress affinity rules you already run stay exactly where they are. |
| Mixed-version clients | Same client identity alternating between handshake and _meta-style requests across retries | Version-sniff at ingress: presence of Mcp-Method/_meta version selects the stateless pool, handshake traffic stays pinned. Never let one pool serve the other class's wire. |
The realistic migration posture is running both transport paths side by side behind the same ingress until every client you support has moved (MCP Stateless Revolution). The sticky-routing configuration doesn't disappear the day you upgrade your servers — it stays live serving the old path while new traffic routes to the stateless pool through header-based rules.
Note the asymmetry in what was removed versus deprecated. The spec's deprecation policy gives phased-out features a twelve-month compatibility runway, but the session header and the initialize handshake were removed outright, with no grace period (MCP Stateless Revolution). That means there is no "temporarily translate old sessions into new handles at the gateway" middle path blessed by the protocol — you serve two wires or you break old clients.
Migrating cross-call state to explicit handles
Stateless protocol does not mean stateless application. If a server needs to remember something between tool calls, it mints an explicit handle on the first call and the client passes that handle back as an ordinary tool argument — the same pattern any REST API has always used. What disappears is the requirement that the same pod be the one to look it up.
Take a deploy/rollback toolset as the worked example, since deploy-from-chat is exactly the surface this transition hits hardest. The stateful shape looks like this: the agent calls deploy_plan, the server stashes the plan in session memory, the agent calls deploy_apply, and the pinned server instance reads back what it stored. Under sticky routing this works until a rolling deploy kills the pod mid-conversation — the client either gets a reset or silently talks to a session that no longer exists.
The stateless shape moves the state into the handle:
deploy_planreturns{ "deployment_id": "dep_7f3a…", "plan": { … } }. The plan itself is persisted to your backing store (Postgres, object storage, whatever the platform already runs), keyed by that id.deploy_applytakesdeployment_idas a required argument. Any replica can serve it: look up the id, validate it, execute.deploy_verifyanddeploy_rollbacktake the same handle. Rollback needs no memory of the conversation — it needs the id and the stored pre-deploy snapshot the id points to.
The migration rule of thumb: anything the tool needs on call N+1 that it didn't receive as an argument must either travel inside an opaque handle or live in a backing store the handle keys into. Session memory is the one place it may no longer live. Audit each tool with that sentence and the migration list writes itself.
What stateless does not fix
Three honest caveats before you delete anything. First, per-request auth is a real cost. The gateway now verifies identity on every call rather than once per conversation, so model that added latency explicitly when you size the gateway layer — it lands on every tool call your agents make (MCP Stateless Revolution).
Second, revocation semantics get sharper. Under sessions, a revoked credential often kept working until the session expired; now it bites on the next call. That is better security, but it changes incident behavior: rotating a leaked key mid-incident will immediately break in-flight agent conversations that were mid-deploy, so your runbook should say what the agent is expected to do when call N+1 returns unauthorized after call N succeeded.
Third, the ecosystem fragmentation above means your shiny stateless server still has to interoperate with a world that is 90% old wire. Of endpoints that answered any version in the October census, roughly nine in ten negotiated a pre-2026 revision (Pennyforge). If your deploy-from-chat agent calls third-party MCP servers — and it will — your client side needs the same dual-wire tolerance as your server side for as long as the census looks like that.
Ship stateless-first, and know your decommission criterion
If you are building a new self-hosted MCP server today — a deploy/rollback surface, a preview-environment provisioner, any tool an agent calls against your platform — do not build the session store. The stateful path is a migration you would be scheduling on day one: sticky routing config, a Redis or equivalent to survive rolling restarts, and a later handle-passing rewrite of every tool, all to arrive where the protocol already is.
Instead, ship stateless-first: per-request _meta, explicit handles for cross-call state, header-based routing from the first commit. Serve legacy clients only if you concretely have them, on the pinned old path from the routing table above — not as your architecture.
And write down the decommission criterion now, while the pain is fresh: the sticky-routing config and any session store get deleted when the last legacy client moves, measured by wire-class traffic at ingress, not by calendar date. The census says that day is not soon — 4% adoption two months after finalization is a long tail by any measure. But when your own ingress shows zero handshake traffic for a full client-upgrade cycle, that is the signal. Delete the affinity rules, decommission the store, and enjoy the ordinary load balancer the new spec promised.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agents are first-class operators: machine-readable infrastructure state your deploy-from-chat tooling can call today. Star the repo on GitHub or deploy your first app today.



