For twenty months, running a Model Context Protocol server in production meant running something that looked nothing like the rest of your infrastructure. Every client opened a session, every session had an ID, and every request after the handshake had to find its way back to the server instance that owned that session — sticky routing, shared session stores, and a whole class of scaling problems that your stateless web services had solved a decade ago. September 2026 coverage is calling the latest MCP revision the protocol's biggest update ever, pitched explicitly at massive enterprise production deployments. Here is the verdict before the why: the enterprise release is secretly a self-hosting release. An MCP server is now an ordinary HTTP service, and serving one on hardware you own just got radically simpler.
The before/after fits in one table, and it is the whole post in miniature:
| Before (2025-11-25 and earlier) | After (2026-07-28, crowned in September) | |
|---|---|---|
| Protocol shape | Stateful: initialize handshake, Mcp-Session-Id on every request | Stateless: every request self-describing, any request lands on any replica |
| Single-node serving | One process must own the session; restart it and clients re-handshake | One binary or container behind a plain reverse proxy; restarts are invisible |
| Fleet serving | Sticky routing or a Redis-backed session store in the gateway | Plain load balancer or Gateway API route, HPA replicas, zero MCP-specific routing |
| Tool listings | Re-negotiated per session | Cacheable list results with TTL, served from cache or CDN |
| Auth | Bring-your-own-token tolerated | Formal OAuth 2.1 resource server, mandatory issuer validation |
If you stopped reading here, you would already know the operational punchline: delete the session affinity, keep the load balancer. The rest of this post justifies the "biggest ever" label, walks through what changed on the wire, and splits the deploy-from-chat roadmap into what gets simpler and what you can safely leave to someone else.
July shipped the wire, September named it the biggest
The revision at the center of this story is the 2026-07-28 specification, finalized on July 28, 2026 after a ten-week release-candidate freeze. Coverage in September — VentureBeat calling it MCP's biggest update ever, "for massive enterprise production deployments," and the Agentic AI Foundation calling it the largest revision since launch — is the ecosystem absorbing what shipped and rendering a verdict. That layering matters: July changed the bytes, September changed the planning assumptions.
What July changed on the wire is a clean break from everything before it. The initialize handshake is gone. The Mcp-Session-Id header is gone. In their place, every Streamable HTTP request carries routing headers (Mcp-Method, Mcp-Name) plus the protocol version and client capabilities in the request's _meta, and servers identify themselves in result _meta with a server/discover method replacing the old capability dance. As one migration guide put it, if you learned MCP transport from a pre-August 2026 tutorial, most of what it taught is now wrong: every request is self-describing, so any request can land on any server instance. Multi-round-trip requests give clients and servers a stateless way to do what used to need server-initiated elicitation, and list results became cacheable with explicit TTL and scope — tools/list can finally sit behind a cache instead of being re-negotiated per session.
Bundled into the same revision is the enterprise package that earned the superlative. Authorization moved to formal OAuth 2.1: mandatory issuer validation under RFC 9207, protected-resource metadata (RFC 9728), resource indicators (RFC 8707), and Client ID Metadata Documents replacing Dynamic Client Registration. Every remote MCP server is now a proper OAuth resource server — no more skipping auth because a server only talks to internal agents.
Two headline capabilities graduated into a new formal extensions framework: MCP Apps (interactive server-rendered interfaces) and the redesigned Tasks extension for long-running async work. And a formal deprecation policy guarantees a minimum twelve-month window between deprecation and removal, so enterprise teams can build against the spec without fearing silent breakage. Roots, Sampling, Logging, and HTTP+SSE entered that twelve-month window, and tool definitions moved to JSON Schema 2020-12.
Why "biggest since Anthropic's November 2024 launch"? No previous revision touched all four layers at once: the transport core went stateless, the auth model was rewritten against real-world OAuth, two major capabilities graduated, and governance got a written stability contract — all under the Agentic AI Foundation's stewardship as a Linux Foundation directed fund. Earlier releases added features to a stateful core. This one replaced the core and wrote down the rules for everything built on top. Twenty months of accumulated scaling pain, resolved in one breaking revision with SDK updates (TypeScript, Python, Go, C#) landing alongside it.
Serving an MCP server on owned hardware: before and after
The old shape was described bluntly by gateway authors: through 2025-11-25, session affinity was a genuine MCP gateway problem. Mcp-Session-Id needed sticky routing, DELETE terminated a session, and Last-Event-ID resumability meant a load balancer had to return a client to the same replica. The Kuadrant MCP gateway's scaling guide still documents the split personality this forced on operators: stateless 2026-07-28 traffic routes header-based with no shared state, while legacy traffic needs an external Redis session store so in-memory client-to-backend mappings survive replica churn. If you self-hosted an infrastructure MCP server — deploy tools, fleet state, runbook automation exposed to agents — you either pinned clients to replicas or ran session infrastructure for a workload whose actual capabilities were stateless. The protocol forced statefulness on stateless work.
The new shape is boring, and boring is the point. A 2026-07-28 server is a process that answers independent HTTP requests. On a single node, that means one binary or container behind Caddy, nginx, or whatever reverse proxy already terminates your TLS — no affinity config, no session store, restarts invisible to clients. On a fleet, it means the same Gateway API route and HorizontalPodAutoscaler pattern as every tenant app: replicas scale on CPU, requests, or queue depth, and the load balancer needs no MCP awareness at all. Google's own guide to running a remote MCP server on GKE makes exactly this pitch — stateless servers, horizontal scaling, Gateway API plus certificates for encrypted traffic — and every line of it applies to a Cluster API fleet on Hetzner just as well as to Google's managed control plane. Your MCP serving layer stops being a special snowflake and becomes one more Deployment behind one more route.
One honest caveat, well put by an early migration note: the protocol is stateless, but the work is still stateful. A deploy triggered from chat still has progress, logs, and a terminal state; stateless transport does not erase application state, it just stops forcing the transport to carry it. That is precisely what the Tasks extension is for — long-running work gets a first-class async lifecycle instead of a held-open session — and it is why the "what gets simpler" list below is longer than the "what breaks" list.
What gets simpler on the deploy-from-chat roadmap
Start with scaling, because it is now free. An infrastructure MCP server that exposes deploy, rollback, log-tail, and fleet-status tools sees spiky traffic: idle for hours, then ten agents hammering it during an incident or a deploy window. Under sessions, absorbing a spike meant pre-provisioned replicas with affinity intact; under stateless requests, it means an HPA doing its job. Multi-round-trip requests extend the win to approval flows — a deploy tool that asks "confirm production rollout?" no longer needs a live session to hold the question open; the round trips are ordinary requests any replica can serve.
Cached tool listings are the quiet performance win. Agents enumerate tools constantly, and tools/list with TTL and cache scope turns that chatter into cache hits — served from memory, a sidecar, or a CDN edge instead of re-negotiated per session. For a deploy-from-chat server whose tool surface changes only when you ship new tools, this is close to free elimination of the most frequent call pattern. And the Tasks extension finally gives long-running operations — a build, a migration, a rolling restart — a protocol-blessed shape: start the task, poll or subscribe, collect the result, with no session to babysit across the minutes or hours a real deploy takes.
There is a second-order simplification worth naming: your MCP server and your tenant apps now share a serving contract. Same Gateway API layer, same TLS story, same autoscaling, same observability (standard HTTP metrics, no session-reassembly in the access logs). The team that operates the platform no longer context-switches between "how we run apps" and "how we run the agent tool layer." For a small self-hosted operation, that consolidation is worth more than any single feature — it is one runbook instead of two.
What stays someone else's problem
The enterprise-production framing brings conformance demands, and a self-hosted deploy-from-chat operator should know exactly which ones to defer. Three are safe to park.
First, Enterprise Managed Authorization and the surrounding governance paperwork. Centralized policy over hundreds of MCP servers across business units is a Fortune 500 problem: many teams, many servers, auditors asking who approved what. A single-team box or a small fleet with one deploy MCP server gets its authorization from the spec's hardened OAuth core — validate the issuer, scope the tokens, log the calls — without the managed-authorization superstructure. Defer it until you have more MCP servers than you can name.
Second, formal audit-trail infrastructure. The 2026 roadmap chatter puts audit trails at the top of the enterprise gap list, and regulated multi-tenant deployments will need tamper-evident, retention-policied tool-call journals. Your deploy-from-chat server should absolutely log every tool call with identity and parameters — that is just good operations — but a compliance-grade audit pipeline with legal-hold retention is not the price of admission for a team deploying its own apps from chat. Structured logs plus your existing log retention gets you there.
Third, multi-vendor governance conformance. The AI-conformance-style test suites and cross-vendor certification bars are written for platforms selling to enterprises that demand proof. If your MCP server's clients are your own agents and your own team, passing someone else's conformance suite buys you nothing; interop with the SDKs you actually use is the bar that matters. Track the governance work — it keeps the ecosystem honest — but do not staff it.
The pattern across all three: the spec's core (stateless transport, hardened auth, deprecation guarantees) is yours to adopt on day one; the enterprise superstructure (managed authorization, compliance audit, certification) binds only when your deployment looks like an enterprise's. A self-hoster adopting the core while deferring the superstructure is not cutting corners — it is matching cost to risk.
What breaks if you already serve MCP
This is a breaking revision, and hand-rolled servers feel it most. If your MCP server was built directly against the wire instead of on an official SDK, the checklist is concrete: remove the initialize handshake and all Mcp-Session-Id handling, require and route on the new Mcp-Method/Mcp-Name headers, move capabilities into _meta, make list results cacheable, and treat auth as OAuth 2.1 with mandatory issuer validation. One widely shared migration note warned operators they had essentially a weekend between the release candidate and final to absorb this — sessions gone, handshake gone, auth rewritten — and the final spec kept every breaking bit.
SDK users have it easier: TypeScript, Python, Go, and C# SDKs shipped matching updates with migration notes, and most of the wire changes are SDK-mediated. Either way, note the deprecation ledger: Roots, Sampling, and Logging moved out of core into extensions with the twelve-month removal window ticking, and Dynamic Client Registration is deprecated in favor of CIMD. If your server registers OAuth clients dynamically, that migration — not the stateless core — is likely your biggest real work item. Tool definitions moving to JSON Schema 2020-12 deserves a validation pass too, since stricter schemas can reject tool shapes that previously slipped through.
The pragmatic migration order for a self-hosted deploy server:
- Upgrade the SDK and flip the transport to stateless.
- Delete the session store and verify behind your existing route — no gateway changes needed.
- Schedule the auth migration (CIMD, issuer validation) as the deliberate second step.
The stateless flip is the rare breaking change that deletes infrastructure instead of adding it.
The enterprise release that deleted your session store
Step back and notice the shape of this story. The largest MCP revision since launch was pitched at massive enterprise production deployments — stateless scale, hardened auth, written stability guarantees — and every one of those enterprise demands cashes out as something a self-hoster wants: no sticky routing, no Redis for sessions, ordinary load balancers, cacheable listings, a deprecation policy that lets a two-person team plan a year ahead. The enterprise bar and the self-hosting bar turned out to be the same bar, viewed from opposite sides. Enterprises needed MCP to behave like the cloud-native infrastructure they already run; self-hosters needed MCP to behave like the ordinary services they already know how to operate. Statelessness satisfied both at once.
Expect the ecosystem to reorganize around that fact over the next year. Gateways will shed their session stores, tutorials will stop teaching handshakes, and "serve an MCP server" will become a five-line container spec instead of a stateful-service design exercise. If you have been waiting to expose your deploy pipeline to agents, the waiting rationale just expired: the protocol finally meets your infrastructure where it already is.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Deploy-from-chat is on the roadmap: an infrastructure MCP server your agents can operate, served from the same fleet as your apps. Star the repo on GitHub or deploy your first app today.



