Skip to main content

MCP Went Stateless: What the 2026-07-28 Spec Means for Long-Running Deploy Tools

13 min readDora NodaDora Noda
Share
On this page

Your deploy-from-chat endpoint just lost its memory, its held-open socket, and its right to ask questions unprompted — all in one spec revision.

On July 28, 2026, the Model Context Protocol published revision 2026-07-28, the largest change since Anthropic open-sourced MCP in late 2024. The headline is architectural: MCP is no longer a stateful, bidirectional session protocol. It is now a stateless request/response protocol, the same shape as every HTTP API you already run.

The initialize handshake is gone. The Mcp-Session-Id header is gone. Every request is self-contained, carrying its own protocol version, client identity, and capabilities — which means any request can land on any server instance behind a plain round-robin load balancer.

That is genuinely good news for scaling. It is also a breaking change for exactly one category of tool: the kind that takes minutes. If your MCP server exposes deploy, rollback, or promote tools that run for longer than a single HTTP round trip — pausing mid-run for a human to confirm "yes, ship to production" — the three mechanisms that made that work under the old protocol no longer exist. This post is the migration map: what broke, and the two stateless patterns that replace it.

The 30-second version: three things that break, three replacements​

If you run an infra MCP server and want the whole migration on one screen, here it is. The rest of the post is the worked detail behind each row.

Stateful pattern (pre-2026-07-28)Why it breaksStateless replacement
Hold the stream open while a deploy runs, streaming progress backNo long-lived bidirectional stream; each request is an independent round tripReturn a Task handle immediately; the client polls tasks/get for status
Server sends elicitation/create mid-run to ask "confirm production deploy?"Servers MUST NOT initiate JSON-RPC requests; there is no channel to ask onMulti Round-Trip Requests: return input_required, the client retries with answers
Remember the pending deploy in session memory between callsNo sessions; the next request may land on a different instanceServer-minted handles passed as ordinary arguments, backed by a shared task store

Your deploy logic doesn't change. What changes is where the conversation state lives: out of the connection and into explicit, addressable handles.

What actually shipped in 2026-07-28​

The revision had a long runway, which tells you how settled the new shape is. Maintainers had been flagging stateful connections as the thing that made MCP hostile to serverless since December 2024 — a protocol that needs a persistent, sticky connection per client needs a warm server per client, and one early-2025 analysis put the cost of keeping those connections alive at over $100 a month before doing any useful work.

In December 2025 the MCP Transports Working Group, co-founded by Google and Hugging Face, laid out the formal case: sticky routing blocks clean autoscaling, session affinity defeats ordinary load balancers, and every gateway in front of an MCP server had to become MCP-aware to route by session. Google's Kurtis Van Gent sponsored SEP-2575, the proposal that made statelessness happen. The release candidate froze on May 21, giving SDK maintainers a ten-week validation window, and the final revision shipped July 28.

Concretely, the new protocol removes:

  • The initialize/initialized handshake (SEP-2575, SEP-2567). No capability negotiation round trip before the first real call.
  • The Mcp-Session-Id header. No transport-level session exists to identify.
  • Server-initiated JSON-RPC requests. A server can no longer open a request of its own to the client — no unsolicited elicitation/create, sampling/createMessage, or roots/list.

And it adds:

  • Self-describing requests. Protocol version, client identity, and client capabilities ride inline in _meta on every call.
  • Mcp-Method and Mcp-Name routing headers. A gateway can route and authorize without parsing the JSON body.
  • A server/discover RPC. Discovery-first negotiation replaces the handshake: ask what the server supports instead of negotiating a session for it.

Stateless servers create no transport sessions, expose no standalone SSE endpoints, and cannot send unsolicited server-to-client requests. SDKs are already there — Microsoft's MCP C# SDK v2.0 implements the new revision with stateless as the default, warning loudly if you opt back into stateful mode. For short-lived tools (search, fetch, transform), the migration is close to free: the same tool, minus the handshake.

Why minute-long tools are the ones that feel it​

A deploy tool that shells out, builds a container, pushes it, rolls it across machines, and health-checks the result does not finish in milliseconds. Under the stateful protocol, three conveniences made that shape easy, and all three are gone.

First, the held-open socket. The old pattern was to keep the response stream open for the duration of the run, pushing progress events ("build done," "3 of 8 machines updated") down the same stream the client was waiting on. In a stateless world there is no stream to hold. Each HTTP request gets one response; a deploy that outlasts the request's patience has nowhere to report to. Timeouts that used to be a client-side courtesy become a hard ceiling: if your gateway kills idle responses after 60 seconds and the migration step alone takes four minutes, the client learns nothing unless you redesign the reporting path.

Second, the mid-run question. The deploy reaches the point of no return — production traffic is about to shift — and the tool needs a human to confirm. Under the old protocol the server just asked: a server-initiated elicitation/create traveled down the open channel, the human answered in chat, the run continued.

The new rule, SEP-2260, is absolute: servers MUST NOT initiate JSON-RPC requests at all. There is no channel on which to ask an unprompted question. A deploy that pauses for confirmation and waits on a held socket now pauses forever.

Third, session memory. The old server could stash "deploy #4821 is at step 4 of 7, awaiting confirmation" in memory keyed by session, confident the follow-up would arrive on the same connection to the same process. Stateless requests carry no session, and consecutive requests from one client may land on different instances behind the load balancer. Anything the server "remembers" between calls that isn't encoded in an explicit handle is a bug waiting for the second replica.

Notice what these three have in common: each one treated the connection as the unit of work. The new protocol treats the request as the unit of work. Long-running operations have to become addressable things in their own right — which is exactly what the Tasks extension is for.

Rebuild 1: long runs become Tasks​

The Tasks extension (io.modelcontextprotocol/tasks, SEP-2663) is the formal answer to "this tool takes minutes." It had existed as experimental core in the 2025-11-25 revision via _meta.task; 2026-07-28 promotes it to a full extension with three methods — tasks/get, tasks/update, tasks/cancel — plus a polymorphic-result discriminator and a Task shape carrying status, in-progress requests, and the final result or error.

The flow for a deploy tool looks like this:

  1. The client calls tools/call for deploy, advertising the Tasks extension in its per-request capabilities.
  2. The server starts the deploy in the background and immediately returns a result with resultType: "task" — a handle, not a completion. Task creation is server-directed: the client signals support, and the server decides per request whether to materialize a task.
  3. The client polls tasks/get with the handle. Each poll is an ordinary stateless request; progress ("step 4 of 7") comes back as task status.
  4. If the deploy needs input mid-run, that arrives through the task too — more on that in the next section — and the client feeds answers via tasks/update.
  5. The human types "abort" in chat; the client sends tasks/cancel; the server stops the run and the task resolves to a cancelled error.

Two properties of this design matter more than the method names. First, polling inverts who holds the waiting: the client drives, the server never blocks a connection. Ten-minute deploys, hour-long migrations, and multi-step rollbacks all fit the same shape because no HTTP request outlives its own round trip.

Second, the handle is just data — a string the client passes back as an ordinary argument. That is what makes it survive load balancing: tasks/get for handle t-4821 can land on any instance, as long as every instance can resolve the handle. Which brings us to the part self-hosters own: the task store behind the handle cannot be process memory anymore. One instance's in-memory map is invisible to the replica the next poll lands on. Task state needs a shared home — Postgres, Redis, your platform's existing store — with a TTL policy, because a task handle for a deploy that finished Tuesday should not resolve forever. Several early implementations default to a 30-minute in-memory TTL for single-instance setups and require a shared store the moment you scale past one replica. Treat that as the rule, not the exception.

Rebuild 2: mid-run confirmation becomes a round trip​

Tasks solve duration. Confirmations need a second pattern, because "pause and ask the human" is a different shape from "run long": it needs a question, an answer, and a way to resume exactly where the run stopped.

That pattern is Multi Round-Trip Requests (MRTR, SEP-2322/SEP-2260), and its governing sentence is worth memorizing: a server may only ask while it is answering. Since unsolicited server-to-client requests no longer exist, the server folds its question into a response. Instead of sending elicitation/create down an open channel, the server returns an InputRequiredResult — resultType: "input_required" — carrying an inputRequests map describing what it needs ("confirm shifting production traffic to build 4821: approve / reject") plus an opaque requestState string that captures where the run stopped.

The client presents the question to the human, collects the answer, and retries the request carrying the response and the untouched requestState. The server validates the state token, resumes the run past the confirmation point, and either completes or asks the next question the same way. Sampling — the old "server asks the client's model to generate text mid-run" primitive — is deprecated in favor of this same retry shape, and roots/list moves with it. Every former server-initiated request becomes a response that invites a retry.

For the deploy-from-chat case, the practical consequence is that confirmations become resumable by construction. Under the old protocol, a dropped connection during the confirmation pause meant ambiguous state: did the human answer, did the run proceed? Under MRTR the pause is just a stored requestState awaiting a retry, and the retry is idempotent by design — the client can re-send the answer after a network blip without risking a double deploy, because resuming from the same state token twice converges on the same step. The tradeoff is latency: each confirmation costs a full client round trip through whatever chat surface fronts the agent, so a deploy flow with five confirmations now has five human-speed round trips in its critical path. That is a reason to consolidate confirmations (one "approve this plan" up front beats three mid-run "are you sure?" pauses), not a reason to avoid the pattern.

Combined, the two extensions cover the whole old surface: Tasks carry the run across time, MRTR carries the questions across the stateless gap. A deploy tool that returns a task handle, reports progress through polls, pauses via input_required, resumes via retry, and aborts via tasks/cancel is the direct translation of the old held-socket flow — with better failure semantics, at the cost of explicit state you now operate.

What self-hosting changes about the bill​

Here is where the self-hosted endpoint diverges from the vendor serverless endpoint the spec was largely designed for, and the honest version has two sides.

The win is real and immediate. A stateless MCP server is, from the infrastructure's point of view, just an HTTP service. It sits behind a plain round-robin load balancer with no session affinity, no sticky routing, no MCP-aware gateway. It autoscales on the same signals as everything else — request rate, queue depth, CPU — because any instance can serve any request.

It passes through a standard WAF, an API gateway doing header-based auth on Mcp-Method/Mcp-Name, and a CDN-style edge without protocol-specific plugins. On owned hardware, where every sticky-session workaround was operational surface you maintained yourself, deleting that surface is a genuine reduction in what can break at 3 a.m.

The cost is the task store, and on owned hardware it is yours. The vendor endpoint gets task persistence as a managed primitive. Your self-hosted endpoint holds handles, TTLs, and retry state in whatever you provision: a Postgres table with a swept TTL column, a Redis namespace with key expiry, your platform's existing job store with a new row type.

That store is now load-bearing for every long-running tool: lose it and every in-flight deploy becomes an unresolvable handle. It needs the boring operational virtues — backups you test, expiry you monitor, capacity you plan. Size it honestly: task rows are small but write-heavy, so give it headroom for write amplification instead of sharing the disk your build cache is already saturating.

A quieter cost: observability moves. When the connection was the unit of work, one trace span covered the whole deploy. When the work is a task polled across dozens of requests, the handle must be the correlation key on every poll, update, cancel, and retry — or debugging a stuck deploy means grepping fifty disconnected request logs. Emit it as a structured field from day one.

Net: statelessness removes infrastructure special-casing and adds one stateful component you operate. For a team that already runs Postgres and Redis, that is a trade worth making — but it is a trade, not a free upgrade.

The migration checklist​

If you operate a deploy-from-chat MCP server today, here is the ordered path from stateful to 2026-07-28:

  1. Audit every tool for session dependence. Grep for Mcp-Session-Id handling, in-memory per-client maps, and any code that assumes two calls from one client reach the same process. Each hit is a migration item.
  2. Move every tool slower than seconds onto Tasks. Return resultType: "task" from deploy, rollback, promote, and migrations; implement tasks/get, tasks/update, and tasks/cancel before you need them in production.
  3. Convert every mid-run question to MRTR. Replace elicitation/create sends with InputRequiredResult responses carrying inputRequests and requestState; make retries idempotent on the state token.
  4. Externalize task state with a TTL. Shared store, swept expiry, sized for write-heavy progress updates — before the second replica, not after.
  5. Correlate on the handle. Structured task-handle fields on every log line and trace span touching a task.
  6. Verify against mixed-version clients. Older clients that never advertise the Tasks extension still need a graceful answer — a clear error naming the required extension beats a silent hang.

The deeper shift: the protocol stopped pretending the network is a conversation and started treating it as what it is — independent requests that sometimes cooperate. Long-running work didn't get harder; it got explicit. Handles instead of held sockets, polls instead of pushed progress, retries instead of interruptions. Every pending operation is now a row you can inspect, a handle you can cancel, and a state token you can resume — no matter which instance picks up the next request.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agent-ops friendly: every deploy is an API call your agents can drive. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide