In July 2026, Trend Micro's corrected internet scan found 1,467 MCP servers publicly reachable with zero authentication or encryption — and only 8.5% of surveyed servers used OAuth at all. The protocol that lets AI agents touch production is, in most deployments, the least hardened thing in production.
The enterprise guidance has converged on one production shape: run the server on Kubernetes over HTTP transport, front it with OAuth, and keep an audit log of every tool call. Here is that shape as a ten-item checklist — the core deliverable of this post — with each item expanded below into something you can configure this week.
| # | Checklist item | Section |
|---|---|---|
| 1 | Serve Streamable HTTP on a single /mcp endpoint | Transport |
| 2 | Run stateless so every replica is interchangeable | Transport |
| 3 | Front the server with OAuth (authorization-code + PKCE) | Auth |
| 4 | Scope tokens per agent action, short-lived and least-privilege | Auth |
| 5 | Validate issuer, audience, and resource on every request | Auth |
| 6 | Log every tool call: who, what, when, result | Audit |
| 7 | Retain logs on a replayable schedule; never log credentials | Audit |
| 8 | Lock down the network: TLS everywhere, NetworkPolicy around the pods | Kubernetes |
| 9 | Rate-limit and minimize the tool surface | Kubernetes |
| 10 | Operate it like production: probes, autoscaling, hardened containers | Kubernetes |
Transport: Streamable HTTP, stateless (items 1–2)
MCP has two standard transports: stdio for local child-process servers and Streamable HTTP for anything remote. The older HTTP+SSE dual-endpoint transport was deprecated in the March 2025 spec and formally retired onto a 12-month deprecation clock by the 2026-07-28 revision. For a new server there is no decision left to make: serve Streamable HTTP on one endpoint.
The 2026-07-28 revision also made the protocol core stateless — no initialize/initialized handshake, no Mcp-Session-Id, no server-side session. That is a Kubernetes-shaped change: when no session state is pinned to a pod, a plain round-robin Service works and replicas scale horizontally with nothing fancier than a Deployment. Operators running the legacy stateful mode needed sticky sessions or cookie affinity to keep a stream on one pod; stateless servers delete that entire failure class. The Flux Operator MCP Helm chart, for example, now runs in stateless mode by default and removed the legacy SSE transport outright.
One caution travels with this: treat any session identifier a client sends as untrusted input. Do not bind authorization decisions to it. Auth rides on the token, which is the next section's entire subject.
Auth: your server is an OAuth resource server (items 3–5)
The MCP authorization spec places your server in a specific OAuth role: it is an OAuth 2.1 resource server, and clients discover how to get tokens from metadata your server publishes. The "bring your own token" era is over in the spec; the scans above show how far deployments lag behind it.
Concretely, item 3 means implementing the authorization-code flow with PKCE (S256) for all clients, served exclusively over HTTPS, with dynamic client registration (RFC 7591) so agents can onboard without a human pasting secrets, plus the two metadata documents that make discovery work: authorization-server metadata (RFC 8414) and protected-resource metadata (RFC 9728). If you run an enterprise IdP, the 2026 drafts add an identity-assertion grant path that puts Okta or Entra back in control of who the agent acts as — worth tracking if your agents already live behind corporate SSO.
Item 4 is where most teams under-invest: token scoping per agent action. A token that can list files should not also be able to restart deployments. Bind tokens to your server as the resource (RFC 8707 resource indicators) so a token stolen from one MCP server is not valid against another service, keep lifetimes short, and grant the minimum scope each tool needs. Least privilege is not a slogan here; it is the difference between a leaked token that reads a wiki and a leaked token that pushes to production.
Item 5 closes the loop on the server side: validate issuer, audience, and resource on every request, return proper WWW-Authenticate challenges on failure, and encrypt token storage. Validate on every request, not once per session — there are no sessions anymore, per the section above.
Audit: log every tool call like it will be replayed (items 6–7)
An MCP server that can restart a deployment is a production control plane wearing an agent costume. Item 6 says every tool call gets a structured log record with the same fields an on-call engineer needs to reconstruct an incident: who (agent identity and token subject), what (tool name and arguments), when (timestamp), and result (success, error, and what changed). A minimal record looks like this:
{
"ts": "2026-09-23T02:55:11Z",
"agent": "deploy-bot",
"token_sub": "agent:deploy-bot:ci-4821",
"tool": "kubernetes.restartDeployment",
"args": { "namespace": "shop", "deployment": "checkout-api" },
"result": "ok",
"duration_ms": 812
}Structured JSON to stderr keeps the log pipeline boring and keeps stdout clean for protocol traffic. When the record shape is fixed, the on-call query at 2 a.m. is one filter on tool and token_sub, not archaeology.
Item 7 is about retention and restraint. Keep those records long enough that a 2 a.m. page can replay exactly what an agent did — weeks, not hours — and never log credentials, tokens, or tool arguments that carry secrets. The authorization guidance is explicit on this point because the failure is so common: a beautifully complete audit log that contains bearer tokens is a credential leak with extra steps.
This is also the item that most sharply separates self-hosting from the proxied alternative, which the comparison section below takes up directly.
Kubernetes hardening around the pods (items 8–10)
Items 1–7 are MCP-specific; items 8–10 are the ordinary production discipline the server inherits by living on your cluster — and skipping them is how a well-authed server still gets owned.
Item 8: TLS everywhere, including inside the cluster if your threat model includes a compromised pod, and a NetworkPolicy that scopes exactly who can reach the server pods. The MCP port should be reachable from the ingress controller and the agents that call it, not from every namespace with a curious sidecar:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: mcp-server-ingress
namespace: agents
spec:
podSelector:
matchLabels:
app: mcp-server
policyTypes: [Ingress]
ingress:
- from:
- namespaceSelector:
matchLabels:
name: ingress-nginx
ports:
- port: 8080Item 9 has two halves. Rate-limit the endpoint — agents retry aggressively and a runaway loop can look exactly like a DDoS from the inside. And minimize the tool surface: expose the minimum tool set the agents need, and put every mutating tool through a human audit before it ships. Read-only tools get logging; mutating tools get logging plus a second pair of eyes.
Item 10 is the unglamorous rest: readiness and liveness probes so broken replicas leave the rotation, a HorizontalPodAutoscaler so a fleet of agents at 9 a.m. does not fall over the one replica you deployed at midnight, container hardening (non-root, read-only filesystem, pinned image digests), and secrets via your cluster's secret management rather than environment variables pasted into a manifest.
The alternative: a proxied remote endpoint
The ecosystem keeps offering to make all of the above someone else's problem. Railway's hosted MCP server at mcp.railway.com is the clearest example: one command (railway mcp install --remote) points your agent at a remote endpoint over OAuth with no local install, and recent CLI versions made the remote server the default. Zero YAML, zero on-call.
The tradeoff is exactly items 4, 6, and 7. Your token policy is Railway's token policy. Your audit trail of every production-touching tool call lives on someone else's side, queryable on their terms and retained on their schedule. For agents that manage Railway infrastructure itself, that consolidation is arguably correct — the audit trail lives next to the thing being changed. For agents that touch your production through your MCP server, the trail belongs where your other production evidence lives: your cluster, your log pipeline, your retention policy.
Choose the hosted endpoint when the infrastructure being operated is the host's own and the audit audience is one developer. Self-host the checklist above when agents touch production you own and an on-call engineer you employ has to answer for what they did.
Harden the hands that touch production
The through-line of every 2026 MCP security scan is that adoption outran hardening: tens of thousands of reachable servers, a single-digit share behind OAuth, registries that list servers with no security requirements at all. The spec has done its part — stateless core, OAuth 2.1 resource-server profile, discovery metadata. The remaining gap is operational, and it closes one checklist at a time: Streamable HTTP, OAuth with per-action scopes, an audit log you can replay, and the ordinary Kubernetes discipline around the pods.
The server that lets agents touch production deserves production's own hardening first. Everything else — more tools, more agents, more autonomy — stacks on top of that foundation or it stacks on sand.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



