Four hundred million SDK downloads a month, up roughly fourfold this year — that is the number Anthropic attached to the MCP 2026-07-28 release, and it marks the moment the Model Context Protocol crossed from laptop experiment to default integration surface. The "USB-C for AI tools" tagline has been aspirational since MCP launched in November 2024; at this scale, it is a supply-chain fact. Every SaaS vendor shipping an MCP server multiplies the tool catalog your tenants expect to wire into their agents — including the deploy-from-chat agents pointed at your platform.
Here is the growth snapshot in one table, because the title's claim deserves receipts before analysis:
| Signal | Number | Source |
|---|---|---|
| Monthly SDK downloads | 400M+, ~4x year-over-year | Anthropic, at the July 28, 2026 release |
| Public MCP servers | 17,000+ (Q1 2026); 13,000+ on npm/GitHub (May 2026) | Ecosystem census; 400% YoY growth in new server registrations |
| Servers in Claude's connector directory | 950+ | Anthropic, July 2026 |
| Fortune 500 with MCP in AI workflows | 28% | Zuplo State of MCP, via March 2026 adoption survey |
| Monthly SDK downloads, alternate count | 97M, 3x in six months | Python + TypeScript packages, mid-2026 |
Two counts of downloads deserve a note, not a shrug. The 400M figure is Anthropic's all-SDK number; the 97M figure tracks the Python and TypeScript packages specifically. Both agree on the shape: triple-to-quadruple growth inside a year, from roughly 100,000 downloads at launch to nine figures a month. Whichever denominator you prefer, the curve is the same — and curves like this are what turn "should we support MCP servers?" into "which thousand do we support first?"
The spec change that makes the number hostable
Growth alone would be a burden if every server still needed a babysat connection. For most of its life, MCP was stateful: clients performed an initialize handshake, servers tracked sessions behind Mcp-Session-Id, and anything behind a load balancer had to manage connection lifecycle. That fought the stateless, horizontally scaled infrastructure every platform team already runs.
The July 2026 specification answered the most technically grounded criticism of the protocol by removing the mechanism entirely: no handshake, no protocol-level sessions, every request carrying its own version and capabilities, and a server/discover RPC replacing handshake probing. The 2026-07-28 changelog reads like a deletion log — sessions, server-initiated requests, and the old subscription endpoints all gone — a revision Anthropic's own staff called the most substantial change to the protocol since authorization was added.
The operational consequence is the sentence this post's hosting half rests on: a remote MCP server is now just another stateless container behind an ordinary load balancer. No sticky sessions, no connection draining semantics beyond what your ingress already does, nothing stopping you from running it on serverless or edge runtimes. AWS shipped same-week AgentCore Gateway support for the new spec, which tells you how eagerly the managed side wanted this property. For a self-hosted PaaS on owned Hetzner capacity, the translation is simpler: each MCP server your tenants want becomes one more deployable unit on infrastructure you already operate, with the same TLS termination, autoscaling, and log pipeline as any other workload.
One qualification before anyone draws a straight line from "stateless protocol" to "stateless system." Servers that need cross-call state now pass explicit, server-minted handles as ordinary arguments — the state moved into the application layer, not into nonexistence. And plenty of servers in the wild still speak older revisions, so a platform hosting third-party servers inherits a version matrix, not a flag day. Stateless at the protocol layer means your load balancer stops caring; it does not mean your audit log can.
What 4x multiplies for a PaaS
Start with demand. When 28% of the Fortune 500 already runs MCP servers in production AI workflows and Claude's own connector directory lists over 950 of them, "bring any MCP server" stops being a power-user request and becomes table stakes for any agent surface — deploy-from-chat included. Your tenants will arrive with a GitHub server, a Linear server, a database server, and an internal-docs server, and they will expect all four to work against the deployment target you operate.
The catalog they expect is not the 950 in one directory; it is the 17,000 in the wild, growing 400% year-over-year. No platform team curates that by hand. You either build a registry with a trust decision per entry, or your tenants build one for you out of unvetted installs.
Then the counter-pressure, because growth this fast always ships its own limit. Every MCP connection starts with tool discovery: the agent reads a manifest of every tool, its description, parameters, and schema, every session, before doing anything useful. A database server exposing 106 tools consumed 54,600 tokens just to initialize. The MCPGauge framework measured context retrieval inflating input-token budgets by up to 236x while frequently degrading accuracy.
Scale from one server to ten or twenty and tool descriptions alone can eat 40% of the context window. This is the physics behind the year's loudest backlash — Eric Holmes's "MCP is dead, long live the CLI" hitting the top of Hacker News, Thoughtworks parking naive API-to-MCP conversion in the Hold ring — and the backlash has a point: converting every REST endpoint into a tool and dumping all of them into one agent session does not scale.
But notice what the backlash actually argues for. The critics are not saying tool protocols are useless; they are saying unfiltered tool protocols are unusable at scale — which is a curation argument wearing a rejection costume. The ecosystem's own verdict points the same way: MCP is narrowing toward multi-user, team, and enterprise workflows that need access controls and vendor-maintained integrations, while solo developers reach for CLIs and agent skills for the simple cases. For a PaaS, that lands as a design constraint, not a reason to sit out: your agent-tool surface needs scoped, discoverable subsets per task, not the union of everything installed. The platform that wins this is the one that answers "which of my 17,000 options should this deploy agent see?" with something better than "all of them."
The governance gap: a four-item operator checklist
Hosting is the easy half — stateless containers are a solved problem. Governing what those containers can do to your tenants is the gap 4x growth opens, and it has four concrete items. Each one names the spec machinery or the missing piece, so this doubles as a build list.
1. Which servers are trusted. At 17,000 public servers and 400% annual registration growth, "install from anywhere" is a supply-chain incident with extra steps. The enterprise pattern that emerged in the first half of 2026 is a private registry with RBAC and approval workflows in front of it, and the registry authorization spec now treats the registry itself as an OAuth 2.1 resource server with dedicated read and write scopes.
A self-hosted PaaS needs the same shape at its own scale: a curated catalog where each entry carries provenance (who published it, which version, what changed), a default-deny posture toward anything outside the catalog, and an approval path for tenants to nominate additions. Pinterest's registry-first deployment — covered in our September enterprise-readiness post — is the production reference for how this looks at scale.
2. Who scopes their permissions. Where a server implements auth, the July 2026 spec makes the requirements concrete: OAuth 2.1 with the MCP server as resource server, mandatory Protected Resource Metadata (RFC 9728) for authorization-server discovery, and RFC 8707 resource indicators binding each token to the exact server it was issued for — so a token minted for the docs server cannot be redeemed at the deploy server.
On top of that, the enterprise-managed authorization extension (stabilized June 2026, generally available in Claude that August) replaces per-employee-per-server OAuth clicks with an Identity Assertion grant the organization mints centrally. The operator takeaway: enforce the resource-indicator binding on every server you host, prefer centrally managed grants over per-user consent the moment more than one tenant is involved, and treat any server that cannot do OAuth 2.1 as untrusted input, not as a trusted tool.
3. Where the audit trail lives. Here is the honest gap: the audit-trail extension is the piece the roadmap deferred, which means today there is no standard envelope for "which agent called which tool with which arguments under whose grant." Until one ships, the platform has to manufacture the join itself: log every tool call with tenant identity, server identity, tool name, argument summary, grant reference, and outcome — at the gateway or ingress you already control, not inside each server. That join is also what makes the token-tax problem diagnosable: per-tool invocation counts plus per-session discovery costs tell you exactly which servers are worth their context-window rent. Build the logging now against your own schema; adopt the standard envelope when it lands.
4. What the tools can do to the agent. Tool descriptions are prompt-injection surface — a server advertising a tool whose description hides malicious instructions is attacking the calling agent, not the user — and tool results carry indirect injection from whatever content the tool legitimately retrieved. Anthropic's answer to the April 2026 design disclosure was to call the behavior expected and put sanitization on developers, so do not wait for a protocol fix. The CVE record through 2026 is mostly command injection and path traversal in local servers that trusted agent- or tool-supplied input without checking it, which sets the floor: validate every subprocess call and path your hosted servers touch, run untrusted servers sandboxed with no ambient credentials, and treat server-authored text — descriptions, instructions, discovery metadata — as untrusted input to the agent's prompt, never as trusted system content.
What to do Monday morning
Three actions, in priority order, for a platform team that just learned the ecosystem quadrupled:
- Stand up the curated registry first. Before you host one more server, decide the trust shape: default-deny catalog, provenance per entry, tenant nomination flow. Everything else on this list keys off "which server is this," and without a registry that question has no authoritative answer.
- Host stateless servers as plain containers. The 2026-07-28 spec removed the last reason to treat MCP servers as special snowflakes in your scheduler. Put them behind your existing ingress, autoscale them like any other workload, and spend the saved complexity budget on item 3.
- Log the tool-call join now. Tenant × server × tool × arguments × grant × outcome, on every call, in your own schema. It is your audit trail, your token-tax ledger, and your incident timeline in one pipeline — and it is the only item on this list no future spec release will build for you retroactively.
Growth this steep ends the debate about whether MCP matters; it starts the debate about who runs it well. The protocol did its half by going stateless. The other half — curation, scoping, audit — is platform work, and it is buildable today on infrastructure you already own.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



