The most important MCP announcement of 2026 is a deletion. The protocol's July 2026 specification removes the initialize handshake and protocol-level sessions outright — no more Mcp-Session-Id, no more pinning a client to the server instance that accepted its connection. An MCP request can now land on any replica behind an ordinary load balancer, the same way any boring HTTP request can. For anyone running an MCP server as real infrastructure instead of a single-user local process, that one deletion changes the entire deployment shape.
TL;DR — the before/after:
| Stateful MCP (before) | Stateless MCP (2026-07-28) | |
|---|---|---|
| Request routing | Client pinned to the replica that issued its session id; other replicas reject its requests | Every POST carries its own routing headers; any replica serves any request |
| Scaling | One instance, sticky sessions, or a shared session store (Redis) | Ordinary round-robin load balancer, zero session infrastructure |
| Rolling restarts | In-flight sessions die; clients must re-handshake | No sessions to kill; restarts are invisible |
| Deploy/rollback tools | Long calls hold a stateful connection open to one process | Long calls still need an answer — that is the Tasks primitive, still landing |
The rest of this post unpacks that table: what the 2026 roadmap actually prioritizes, what has already shipped, why a deploy/rollback authority is the workload where stateful hurts most, and — the part that matters if you operate one — which pieces are safe to build today versus genuinely blocked on transport work still in flight.
What the roadmap actually prioritizes
In December 2025, Anthropic donated MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded by Anthropic, Block, and OpenAI. The donation moved the protocol out from under a single vendor's roadmap. The 2026 roadmap, published under that neutral governance, then made a striking choice: fix the foundation before adding features.
The roadmap organizes the year's work into four areas. First, and explicitly first, is Transport Evolution and Scalability. Then come agent communication (the Tasks primitive for long-running work), enterprise readiness (audit trails, SSO-integrated auth, gateway behavior), and security (a tool-validation framework). The roadmap's own diagnosis of the transport problem is blunt: stateful sessions fight with load balancers, horizontal scaling requires workarounds, and there is no standard way for a registry or crawler to learn what a server does without connecting to it.
That ordering is the bet. MCP spent 2025 winning the client footprint — every major coding agent and assistant learned to speak it. But the servers those clients talked to were overwhelmingly local processes over stdio: one user, one machine, one process. The moment a server graduates to shared infrastructure — one deploy authority serving a whole fleet of agents — the stateful session model becomes the bottleneck, and the roadmap puts that bottleneck ahead of every new capability.
What already shipped: the stateless core
The roadmap's transport work has already landed its centerpiece. The 2026-07-28 specification release — developed with the MCP Transports Working Group, co-founded by Google and Hugging Face — removes transport-level session management entirely and gives the protocol a stateless core.
Concretely, three things changed on the wire. The initialize handshake is gone. Protocol-level sessions, and with them the Mcp-Session-Id header, are gone. And every POST now carries the routing information a load balancer or gateway needs inline: Mcp-Method and MCP-Protocol-Version on each request, plus Mcp-Name on the calls that name a target — tools/call, resources/read, and prompts/get. A gateway can route on headers alone without ever holding connection state.
This did not come out of nowhere. The March 2025 specification had already replaced the old SSE-only remote transport with Streamable HTTP and added a stateless operation mode that early production deployments adopted as the scaling path — FastMCP's stateless_http mode, for example, serves each request on an ephemeral transport with no session to keep alive. What 2026-07-28 does is make stateless the protocol's center of gravity instead of an optional mode: servers no longer need to remember connections, and clients no longer need to stay pinned.
The same release tightens the auth story that shared infrastructure depends on. It aligns the spec more explicitly with OAuth 2.0 and OpenID Connect practices, adds RFC 9207 issuer validation, and binds client credentials to the authorization server that minted them — closing the credential-reuse gaps that appear the moment one MCP server accepts tokens in a multi-tenant setting.
Why a deploy authority is the worst case for stateful
Not every MCP server feels the stateful bottleneck equally. A read-only docs server serving short queries can live behind sticky sessions indefinitely. A deploy/rollback authority — the server whose tools are deploy, rollback, promote, scale — is close to the worst case, for three compounding reasons.
First, its calls are long. A deploy is not a millisecond lookup; it is minutes of build, push, rollout, and health-check verification. On a stateful transport, that means a long-lived session pinned to one process for the entire operation. Rolling-restart the server mid-deploy and the session dies with it — the client gets "session not found" and must re-handshake against whatever state the deploy left behind.
Second, its callers are concurrent. A deploy authority does not serve one user; it serves every agent in the org, plus CI jobs, plus the occasional human-driven client. Concurrent long-lived sessions on one pinned process is precisely the load pattern that forces the choice between vertical scale-up (a bigger single box) and session-store surgery (Redis in front of every replica, with all the invalidation semantics that implies).
Third, its operations are stateful in the application sense even when the protocol goes stateless. Knowing that replica 3 can serve any request does not answer "where did my deploy get to, and how do I pick up the answer?" That question is what the Tasks primitive (SEP-1686) exists to answer: durable, pollable state machines — working, input_required, completed, failed, cancelled — that wrap long-running tool calls in a call-now, fetch-later pattern. The client fires the deploy, goes away, polls tasks/get for status, and fetches the result with tasks/result when it is ready.
Tasks shipped as experimental in the 2025-11-25 specification, which means the shape of the answer is visible but the stable contract is still being settled — exactly the "safe to prototype, premature to hard-depend on" status the next section is about.
Build today vs genuinely blocked
Here is the verdict the roadmap implies for a team running its own deploy-authority MCP server, split into what you can ship now and what is still landing.
Safe to build today:
- Stateless Streamable HTTP behind a round-robin load balancer. Run one server per capability as an ordinary container, N replicas, routed on
Mcp-Method. No sticky sessions, no shared session store. This is the default production shape now, not an experiment. - Per-request OAuth 2.1 auth against an external authorization server. Audience-scoped tokens, validated per request, no session-bound credentials. The
2026-07-28issuer binding makes this cleaner, but the pattern works on the existing spec. - Unauthenticated
/healthprobes. Liveness endpoints for the load balancer, container health checks, and orchestrator probes — the unglamorous prerequisite every multi-replica deployment needs. - Audit logging of every tool call. MCP's standardized request-response shape means each invocation is loggable, traceable, and attributable through one layer. For a server whose tools mutate production, this is table stakes, and it needs no protocol upgrade.
Genuinely blocked or experimental:
- Tasks (SEP-1686) for long deploys. The primitive exists and is the right answer for call-now, fetch-later deploys — but it is experimental, and its stable semantics are part of this year's roadmap work. Prototype against it; keep a polling fallback of your own until it settles.
- Standard server discovery. The roadmap wants registries and crawlers to learn what a server offers without opening a connection. Until that standard exists, discovery stays bespoke: docs, config, out-of-band negotiation.
- Gateway behavior and tool validation. Enterprise gateway semantics and the security framework for validating tool definitions are roadmap items, not shipped specs. Multi-tenant deployments still assemble this layer themselves.
The practical upshot: the transport fix unblocks the infrastructure shape (scale-out, rolling restarts, ordinary load balancing) today, while the workflow conveniences on top (durable tasks, discovery, gateway policy) remain build-it-yourself or wait. A deploy authority needs the first category desperately and the second eventually — which is a kind word for the roadmap's ordering.
What self-hosting it looks like
Put the shippable pieces together and the self-hosted shape is reassuringly boring — which is the point. A deploy-authority MCP server becomes an ordinary web service: containerized Streamable HTTP servers, one per capability, behind a round-robin load balancer routing on headers, with an external OAuth authorization server and per-request validation. Health probes keep the load balancer honest; audit logs keep the humans honest.
Two migration notes for anyone already running stateful servers. First, the Redis session store does not disappear on day one: gateways still need it for older stateful clients speaking pre-stateless spec revisions, with stateless header-based routing alongside for current clients. Run both during the transition, then retire the store when your client fleet has moved. Second, keep legacy-session support as the compatibility shim it is — new integrations should speak the stateless core from the start, so the shim has a shrinking user base instead of a growing one.
There is a larger point hiding in that boring shape. Every capability MCP's roadmap deprioritized in favor of transport work — tasks, discovery, gateways — is easier to build, operate, and reason about once the layer underneath is a stateless HTTP service instead of a session-pinned singleton. Stateful-to-stateless is the kind of fix that compounds: it unblocks the scaling story directly and simplifies everything stacked on top of it.
Agents as operators need boring infrastructure
Step back and the roadmap reads as a maturity statement. MCP won the client war; 2026 is the year it grows up into infrastructure. The protocol's stewards looked at the gap between "my agent talks to a local tool process" and "a fleet of agents operates production through shared services" and decided the highest-leverage work was making the second shape as boring as possible: no sessions, no affinity, no special snowflake routing — just HTTP semantics every platform team already knows how to run.
For a self-hosted platform, that is the best possible outcome. A deploy authority you can run as N identical containers behind a load balancer you already operate, authenticated the way your other services authenticate, is a deploy authority you can actually own. Build the stateless shape now; track Tasks toward stability; and let the roadmap's feature work arrive on a foundation that is already horizontal.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agents that deploy through boring, self-hosted infrastructure are the whole thesis: star the repo on GitHub or deploy your first app today.



