Skip to main content

CNCF Says MCP Gateways Belong in the Platform. Here's What That Actually Requires

12 min readDora NodaDora Noda
Share
On this page

In July 2026, the CNCF published its Platform Engineering 2.0 framing, and one line in it should change how your team talks about AI agents. The new AI-Native Platform pillar says the platform must provide "first-class support for GPU/TPU allocation, model serving, MCP gateways, and agentic guardrails" — with AI systems treated as platform consumers that "require governance, access controls, and operational guardrails similar to those applied to human users." (CNCF, July 6, 2026)

That is a promotion with teeth. An MCP gateway is no longer a nice-to-have sidecar your most enthusiastic team stands up. It is platform infrastructure, next to identity and ingress. But most teams nodding along have never pinned down what a gateway does that a single MCP server doesn't — or what "baking it into the platform" concretely obligates them to build. This post is that breakdown, with the auth, rate-limiting, and tool-scoping specifics, plus an honest accounting of how much of it a self-hosted PaaS control plane already absorbs.

Here is the core of it up front:

ResponsibilityAd-hoc MCP server (per team)MCP gateway (platform layer)
AuthOften none; API key if you're luckyOAuth 2.1 + PKCE, per-tool scopes, verified caller identity
Rate limitingNone, or coarse per-IP at the edgePer caller, per target, and per caller-per-target
Tool scopingEvery connected client sees every toolDeny-by-default, per-role tool policies
Tool provenanceWhatever got installedNamespace-verified registry + definition-drift detection
AuditServer logs, maybeComplete per-call trail: who, which agent, which tool, which args
Exposed surfaceN endpoints for N serversOne /mcp endpoint, one federated tool list

If that table surprises you, the rest of this post is for you. If it doesn't, skip to the threat-model section — the 2026 numbers are worse than you think.

A gateway is a front door, not a bigger server​

Start with the definition, because "gateway" gets stretched to cover everything from a config file to a SaaS. An MCP gateway sits in front of one or more upstream MCP servers and presents them to clients as a single endpoint. Clients see one /mcp URL and one merged tool list. The gateway handles everything behind that: routing each tools/call to the correct backend, enforcing auth, and applying policy before the call ever reaches a server.

Concretely, that means protocol-level federation, not just a reverse proxy. Session-aware gateways keep per-client state across the stateful Streamable HTTP transport and route to the right backend session. Envoy-based designs inspect tools/call payloads with an external processor and fan out to the owning server. The point is the same in every design: the client never learns how many servers exist, where they run, or what credentials each one needs.

The landscape has already stratified by scale. On a laptop, 1MCP and the Docker MCP Gateway aggregate local servers behind one endpoint, with Docker adding signed, container-isolated server packaging. In Kubernetes, Microsoft's MCP gateway does session-aware stateful routing and lifecycle management, while agentgateway positions itself as the policy-driven data plane for agent and MCP traffic. At enterprise scope, IBM's ContextForge federates MCP, A2A, and REST behind one governed front door. Same pattern at every tier; what grows with scale is the policy surface, not the concept.

The per-team alternative is the thing most shops run today without naming it: each team stands up the MCP servers its agents need, each with its own auth story (or none), its own network exposure, and its own logging. It works until the second team, and then you have N front doors with N policies and no unified answer to "which agent called which tool with what arguments last Tuesday."

Auth: the protocol finally grew up​

The single biggest change enabling gateways as platform infrastructure landed in the protocol itself. The July 28, 2026 revision of the MCP authorization specification mandates OAuth 2.1 with PKCE (S256) for any remote, HTTP-based MCP server, aligning with OpenID Connect and retiring the older session model. (spec summary) API keys still linger in local-server contexts like IDE plugins, but for remote servers they are no longer spec-compliant — the enterprise guidance is explicit that OAuth 2.1 plus Dynamic Client Registration (RFC 7591) and Authorization Server Metadata (RFC 8414) is the bar in 2026.

This matters for gateways because it gives them something standard to enforce. The pattern the ecosystem has converged on is per-tool scopes on the token: a token scoped mcp:tool:filesystem:read grants the filesystem server's read tools and nothing else — not the billing database server, not the deploy server. The gateway validates the token signature, checks the scope claims, and only then routes. A developer's token for one server never implicitly becomes a token for all of them.

Note the second persona shift hiding here, which CNCF's Multi-Persona Experience pillar names outright: AI agents are now non-human platform consumers "with their own access, scope, and governance needs." A calling agent is not the principal — it acts with delegated, least-privilege scopes, and every invocation records which agent made the call. That is non-human identity as a first-class platform concern, and it is why the 12-month AI-native readiness checklists now list "non-human identity" and "MCP gateway exposure" side by side. (Platform Engineering 2.0 readiness)

What the gateway enforces before routing, then, is a short list with no shortcuts: valid token, audience-bound to this gateway, scopes covering the exact tool being called, and an identity (human or agent) worth writing into the audit trail.

Rate limits and tool scoping are per-caller-per-target​

Once auth is settled, the next two responsibilities separate gateways from servers. Start with rate limiting, because per-IP limiting at the edge — the thing most teams already have — is the wrong shape for agent traffic. Tenants share IPs; one runaway agent behind NAT shouldn't starve the nine well-behaved ones next to it.

Gateway-grade rate limiting works on three axes: the caller (user or agent identity, by group membership), the target (MCP server or tool), and each caller independently per target. That third axis is the one that matters: cap how hard any single agent can hit any single backend, per time window, with config-time floors so a misconfigured policy can't lock out the platform's own automation. Enterprise gateway implementations pair this with a fail-open availability guardrail — when the limiter itself is unhealthy, traffic flows and the incident gets paged, rather than the gateway becoming the outage.

This is also where CNCF's Embedded FinOps pillar stops being abstract. Every agent tool call has a token cost and a compute cost, and a gateway sitting on the single front door is the natural place to attribute both per caller, per team, per tenant. Cost intelligence "moves from bolt-on reporting to provisioning-time decisioning" — and for agent traffic, the gateway is where provisioning-time happens. Pre-deployment cost gates have an obvious sibling here: pre-call budget checks against a tenant's agent spend.

Tool scoping follows the same deny-by-default discipline. MCP servers natively expose every tool to every connected client; as one gateway's README puts it, "any connected client can call any tool with any arguments." The gateway inverts that: per-role tool policies where nothing is callable unless explicitly granted, so the read-only analytics agent never even sees the deploy_production tool in its federated list — not "sees it but gets rejected," but never receives it in tools/list at all. Hiding unauthorized tools from discovery, rather than merely rejecting their invocation, shrinks both the attack surface and the agent's confusion surface.

The threat model that forces this into the platform​

All of the above could, in principle, be per-team discipline. The reason CNCF put gateways in the platform layer is that per-team discipline has observably failed. The 2026 numbers:

  • A CyberPress analysis found 4,982 security issues across 2,259 public MCP servers, including template injection enabling server-side code execution, eval() calls inside tools, and path traversal in widely-starred servers. (CyberPress)
  • The MCPInspect academic study (DSN 2026) found 833 vulnerable MCP servers in the wild, 18 with suspicious or deliberately misleading tool descriptions — signs of intentional poisoning. (via Forkast)
  • The MCPTox benchmark measured tool-poisoning attack success up to 72.8% against capable agents — and more capable models were often more vulnerable, because the attack exploits superior instruction-following. (MCPTox writeup)
  • Real CVEs are already in the wild: CVE-2025-54136 (MCPoison persistent code execution in developer IDEs) and CVE-2026-26118 (AI tool hijacking) (tool-poisoning defense survey), plus CVE-2026-75130 (Context7 MCP prompt injection, August 2026) (incident archive).
  • The Cloud Security Alliance now calls MCP "one of the most rapidly weaponized attack surfaces in agentic AI deployments," and Akamai's September 2026 agentic-enterprise report dedicates its analysis to MCP as the expanding attack surface demanding zero-trust controls. (Akamai via AIGovernance)

Tool poisoning deserves a sentence of its own, because it is the attack the gateway exists to catch. A poisoned tool looks benign — Invariant Labs' canonical demo is a tool that appears to add two numbers — while its hidden description instructs the agent to exfiltrate credentials first. The description is the attack, which means no amount of server-side input validation catches it. Only something sitting between the server catalog and the agent can: the gateway, comparing served tool definitions against pinned, approved copies and alerting (or refusing to serve) on drift. Tools like mcpguard implement exactly this as a lockfile for tool definitions.

The second gateway-side defense is provenance. The official MCP Registry — a metadata-only metaregistry backed by Anthropic, GitHub, PulseMCP, and Microsoft — gives every server a namespace-verified identity (io.github.<user>/* via GitHub auth, reverse-DNS via domain verification). A gateway that only federates registry-verified servers, and pins their definitions, turns "don't install random servers" from a policy document into a routing rule.

The third is the audit trail: every call, with caller identity, agent identity, tool, and arguments. When — not if — a poisoning attempt lands, the blast-radius question ("which agents saw this tool definition, and what did they do with it?") is answerable in minutes instead of never.

What a self-hosted PaaS control plane already absorbs​

Here is the honest part, and the reason this post exists on a PaaS blog rather than a security-vendor blog. A self-hosted platform whose control-plane API is already the single chokepoint for every deploy, rollback, scale, and secret-rotation operation has already built half a gateway without calling it one:

Gateway responsibilityAbsorbed by the control-plane API?
Single front door for platform operationsYes — every deploy/rollback tool call goes through one API
Caller auth (human)Yes — the API already authenticates operators
Per-call auditMostly — API request logs with identity; needs agent-identity fields added
Rate limiting per callerPartially — usually per-API-key, needs the per-target axis
Tool scoping / discovery filteringNo — needs per-role tool policies on the federated list
Third-party server federationNo — tenants bring their own servers; the platform must front them
Tool-definition drift detectionNo — descriptions the platform didn't write need pinning and alerting

That table is the actual migration story. The control plane absorbs identity, the chokepoint, and most of the audit trail — which is precisely why "bake it into the platform" is cheaper than it looks for a PaaS, and precisely why per-team bolt-ons are the wrong answer. What remains is genuinely new work: federating servers the platform doesn't own, scoping tools per tenant role, and treating tool descriptions as untrusted input. None of that is a weekend project, but all of it compounds on infrastructure that already exists.

Note what this framing deliberately excludes: GPU/TPU allocation and model serving are separate items in the same CNCF pillar. A gateway doesn't schedule your GPUs. Don't let a gateway project sprawl into a serving project; they share the pillar, not the implementation.

The bake-in checklist​

If you're the platform team tasked with making "MCP gateway exposure" real in the next two quarters, here is the concrete list, in dependency order:

  1. Pick the gateway pattern for your scale. Laptop-fleet dev tooling: 1MCP or Docker MCP Gateway. Kubernetes-native with session-aware routing: Microsoft's gateway or agentgateway. Enterprise federation across MCP/A2A/REST: ContextForge-class. The pattern matters more than the product — one front door, federated list, routed calls.
  2. Stand up the OAuth issuer with a per-tool scope taxonomy. Remote servers must speak OAuth 2.1 + PKCE per the current spec. Define scopes like mcp:tool:<server>:<verb> before the first server onboards, and enforce them at the gateway, not in each server.
  3. Define rate-limit axes up front. Caller, target, caller-per-target, per window — with lockout-safeguard floors and a fail-open rule for limiter outages.
  4. Sink audit events somewhere queryable. Caller, agent identity, tool, arguments, decision. If you can't answer the blast-radius question in minutes, the trail isn't done.
  5. Pin tool definitions and alert on drift. Record approved tool descriptions in a lockfile at onboard time; refuse to serve (or loudly flag) definitions that change. This is your tool-poisoning control.
  6. Source servers from the verified registry. Only federate namespace-verified servers, and re-verify on a schedule. Random-server sprawl ends at the gateway's routing table.

Six items, each independently shippable, each a platform capability every team inherits the day it lands. That is what "baking guardrails into the platform layer instead of bolting them on" means when you stop waving your hands: one team builds the front door once, and every agent integration behind it gets auth, limits, scoping, provenance, and audit for free.

CNCF's Platform Engineering 2.0 framing treats agents as first-class platform consumers — which is exactly how Bex.co already treats them: every deploy, rollback, and scale operation goes through one API your agents can drive today. Star the repo on GitHub or deploy your first app and point an agent at it.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide