Skip to main content

Your Deploy MCP Server Is Production Infrastructure: Scoring Golf's Built-In Auth, Tracing, and Telemetry

12 min readDora NodaDora Noda
Share
On this page

Picture the moment your deploy-from-chat roadmap stops being a demo. An agent holds a deploy tool, a scale tool, and a rollback tool, pointed at real tenant workloads. Now ask the three questions that decide whether that roadmap ships: which tenant is this agent allowed to touch, where is the record of what it did, and how do you debug the tool call that took down staging at 2 AM? A hand-rolled MCP shim over your API answers all three with "you build that yourself." Golf, an Apache-2.0 Python framework that describes itself as a "production-ready MCP server framework" with auth, observability, debugger, telemetry, and runtime included, claims a different answer. This post checks that claim against the checklist, row by row, and ends with what a self-hosted PaaS should require of its MCP layer before agents become operators.

The short version for readers in a hurry: Golf genuinely covers per-tenant auth scoping, per-tool tracing, Kubernetes probes, and a dev/prod build split out of the box. Audit export, RBAC, rate limiting, and PII redaction live in its separate Gateway product, not the framework. And the "debugger" in the tagline has no documented standalone component — the debugging story is dev-mode tracing plus the standard MCP Inspector. None of that disqualifies it, but it changes what "included rather than bolted on" means in practice.

The verdict up front​

Here is the production checklist for an MCP server that exposes dangerous tools, scored three ways: a hand-rolled shim over your API, the Golf framework, and the Golf Gateway product that sits in front of it.

RequirementHand-rolled shimGolf frameworkGolf Gateway
Per-tenant auth scopingYou implement token verification and scope checksJWT/JWKS with required_scopes, full OAuth Server mode, multi-server RemoteAuthAdds IdP SSO (Auth0, Entra ID, Descope) plus server and per-capability RBAC
Audit trail per agent-issued deployYou build the log pipelineBasic OpenTelemetry traces: timing, success/failure, errors per toolAudit log schema with export to Elasticsearch, OTLP, or Sentinel
Debug a failed tool callLogs and luckTraces plus dev-only input/output capture, structured context logging, Inspector-compatible endpointCentralized audit view across servers
Step-through debuggerNoPartial: no documented standalone debugger; dev-mode tracing is the debugging storyNo — the gateway is policy, not a debugger
Health and readiness probesYou add /healthhealth.py/readiness.py generate /health and /ready with Kubernetes 200/503 semanticsGateway health and metrics endpoints
Dev/prod build splitYour Dockerfilegolf build dev copies .env; golf build prod expects runtime envNot applicable
Horizontal scale storyYour problemStateless MCP 2026-07-28 needs no session affinity, so replicas scale; probes includedPer-user, per-server, and global rate limiting
PII redactionYou build itdetailed_tracing warns about sensitive data and stays dev-onlyPII scrubbing on MCP responses

Two honest cells in that table do most of the work: the debugger is partial, and four rows say "gateway." The rest of this post substantiates every cell, starting with the foundation everything else depends on: auth.

Auth is five decisions, not one​

Golf's most substantial "included" claim is authentication, and it is also where hand-rolled servers most often ship a placeholder. The MCP authorization specification for the current 2026-07-28 revision grounds HTTP-transport auth in OAuth 2.1 (draft-13) with Bearer usage (RFC 6750), authorization server metadata (RFC 8414), dynamic client registration (RFC 7591), resource indicators (RFC 8707), protected resource metadata (RFC 9728), and issuer identification (RFC 9207). Authorization is optional in the spec and HTTP servers should conform — which in practice means every production MCP server replays the same decision: validate tokens yourself, run an authorization server, or delegate to an IdP. Golf's answer is a dedicated auth.py file, separate from tool code, with five modes behind one configure_auth call.

JWT validation is the production default: point it at a JWKS URI, issuer, and audience via environment variables, declare required_scopes, and Golf validates tokens per RFC 7519 before any tool runs. For a deploy MCP server, this is exactly the per-tenant scoping primitive — a token carrying deploy:staging but not deploy:production never reaches the rollback tool, and the check lives in the framework rather than in every tool body.

python
from golf.auth import configure_auth, JWTAuthConfig
 
configure_auth(JWTAuthConfig(
    jwks_uri_env_var="JWKS_URI",
    issuer_env_var="JWT_ISSUER",
    audience_env_var="JWT_AUDIENCE",
    required_scopes=["read", "write"],
))

The more surprising mode is OAuth Server: Golf v0.2.0 can act as a complete OAuth 2.0 authorization server, issuing JWTs and serving authorization, token, and revocation endpoints for its own clients. That inverts the usual shape — instead of your MCP server validating tokens from someone else's IdP, small deployments get tokens minted at the edge of the MCP server itself, with valid, default, and required scopes configured in the same file.

RemoteAuth covers the opposite end: distributed validation across multiple resource servers sharing authorization servers, the microservices shape where several MCP servers trust the same issuers. API-key passthrough extracts a key from a header and hands it to tools via get_api_key() for upstream calls, with actual authentication happening at the destination API. Static dev tokens round out the set for local development.

Two details signal production seriousness. First, tools read identity through helpers (get_auth_token(), get_api_key()) rather than parsing headers, so the trust boundary stays in one place. Second, the README carries an explicit token-hygiene warning: an inbound MCP JWT or OAuth bearer must never be forwarded to an upstream API — use a separate upstream credential or a token-exchange flow. A framework that tells you where its own auth primitive must stop is one whose authors have watched a confused-deputy bug in the wild.

What Golf's auth does not give you is identity lifecycle: no user directory, no group membership, no "who approved this connection." Those live in the Gateway's IdP integrations and RBAC. For a single-team deploy-from-chat pilot, framework auth plus scopes is enough. For multi-tenant agents touching production, the gateway column of the table is load-bearing, and the honest version of the pitch prices it in from day one.

Observability, and the debugger asterisk​

If auth is Golf's strongest chapter, observability is where the tagline needs a careful read. The framework ships OpenTelemetry integration with two deliberately separated levels. Basic telemetry — execution timing, success and failure rates, error information per tool, resource, and prompt — is the production setting. Detailed tracing adds input and output capture: tool parameters and return values, resource content, even full elicitation and sampling conversations. The docs recommend detailed mode for development and testing only, with safe serialization and size limits, and show a production config that keeps opentelemetry_enabled on while detailed_tracing stays off. That split is the correct default for a deploy MCP server: you always want to know that rollback failed and how long it took, but capturing the full arguments of every agent-issued deploy into your trace backend is a retention and PII decision, not a free default.

json
{
  "opentelemetry_enabled": true,
  "detailed_tracing": false
}

Health checking follows the same "boring in the right way" pattern. Drop a health.py and a readiness.py in the project root, each exporting a check() function, and Golf serves /health and /ready with Kubernetes probe semantics: HTTP 200 keeps the pod alive and in rotation, 503 restarts it or pulls it from the load balancer. Without custom files, /ready returns a passing status and /health a plain OK. For a PaaS team, this is the difference between an MCP server that slots into existing fleet monitoring and one that needs a bespoke sidecar before it can run on the cluster at all.

Now the asterisk. The repo tagline promises a Debugger alongside auth, observability, telemetry, and runtime — but no standalone debugger component appears in the README, the CLI reference (init, build, run, telemetry), or any framework docs page. What debugging actually looks like on Golf is a composition: dev-mode detailed tracing that captures the exact inputs and outputs of the failing tool call, structured logging through get_current_context().logger inside tools, a local dev server on localhost:3000, and a standard streamable-HTTP endpoint that the stock MCP Inspector can attach to. That is a legitimate debugging story — arguably the same story most frameworks offer — but it is not a built-in debugger, and the verdict table scores it partial for exactly that reason. When your roadmap's week-one incident is "the agent's scale call returned an error and nobody knows why," what saves you is the trace with the captured arguments, not a step-through UI. Budget for traces, not for a debugger that ships in the box.

Build, runtime, and the scale story​

Golf projects follow a convention-over-config layout that will feel familiar to anyone who has scaffolded a web framework. After pip install golf-mcp, golf init creates tools/, resources/, and prompts/ directories plus golf.json and auth.py; each Python file defines one component with its module docstring as the description, and component IDs derive from the file path (tools/payments/submit.py becomes submit_payments). The build step compiles those files into a runnable server, and the dev/prod split is where the production thinking shows: golf build dev copies your .env into dist/ for local runs, while golf build prod copies nothing and expects environment variables at runtime. Secrets ride the platform's env injection in production instead of a file baked into the artifact — the unglamorous detail that separates a demo server from one you can hand to a fleet.

The transport story tracks the current spec. Golf targets FastMCP 4.0.0 and MCP 2026-07-28, serves streamable HTTP with stdio available, keeps SSE only as a deprecated legacy transport, and negotiates legacy handshake-era clients through a compatibility mode. The 2026-07-28 revision matters for scaling because it made the protocol stateless: no initialize handshake, no session ID, every request self-describing. A sessionless MCP server needs no session affinity, so horizontal scaling is the boring kind — run N replicas behind the load balancer, let /ready admit them, and any replica can serve any request. Combined with RemoteAuth for fleets of resource servers sharing issuers, the framework gives you the primitives; the Gateway's rate limiting (per-user, per-server, global) is what keeps one enthusiastic agent from spending the whole fleet's budget.

One caveat travels with the spec target. The 2026-07-28 ecosystem is young — FastMCP 4.0 is stable but moving fast, with v4.0.10 shipping September 25, 2026, days after v4.0.7 — so pinning versions and reading changelogs is not optional. Golf's compatibility mode for legacy clients softens the edge, but any team adopting it this year should treat the protocol layer as actively settling rather than finished.

Where the framework ends, and what to require​

Draw the product boundary crisply, because the marketing blurs it. The open-source framework gives you the server: auth modes, OTel traces, probes, dev/prod builds, elicitation and sampling utilities, and a runtime that speaks the current spec. The Gateway and Control Plane product gives you the governance layer the framework's own docs describe as the missing control plane: which AI tools connect to which systems, who approved them, what flowed through, server and per-capability RBAC, PII scrubbing, audit export, and rate limits. A self-hosted PaaS evaluating Golf should evaluate both halves and price the gateway into any multi-tenant plan — framework auth scopes stop the wrong tool call, but only the audit export answers "show me every deploy this agent issued last month" at 2 AM.

Two smaller caveats belong in the evaluation notes. Golf is v0.2.x-era software: auth.py itself was a breaking change from the v0.1.x auth API, which is healthy evolution but also a reminder to pin and read migration notes. And the CLI collects anonymous usage telemetry (commands run, success/failure, versions, OS) with opt-out via golf telemetry disable or --no-telemetry — reasonable, documented, and worth knowing before it phones home from your CI.

With the boundary drawn, here is the liftable checklist for any deploy-from-chat roadmap, Golf or otherwise:

  • Per-tenant scopes enforced before tool dispatch, not inside each tool body.
  • An audit trail that records every agent-issued deploy, scale, and rollback with identity attached — traces for debugging, exportable logs for accountability.
  • A trace per tool call with timing and errors always on; input/output capture confined to development.
  • Kubernetes liveness and readiness probes with documented 200/503 semantics.
  • A production build that never bakes .env secrets into the artifact.
  • Token hygiene: inbound agent tokens validated at the edge, never forwarded upstream.
  • Gateway-grade RBAC, PII handling, and rate limiting before agents from more than one tenant share the server.

Adopting a production-grade MCP framework does not remove a single row from that checklist — it changes who implements each row. The hand-rolled shim leaves all seven to your team, due the week an agent misbehaves. Golf's framework takes the auth, tracing, probe, build, and hygiene rows off your plate structurally, splits the audit row with the Gateway (traces here, exportable logs there), and names multi-tenant governance as gateway territory. That is a meaningful shift of risk, priced honestly: less custom auth and plumbing to write, a clear-eyed view of what still needs a governance layer, and no imaginary debugger. For a platform whose roadmap ends with agents as first-class operators, that trade is worth taking deliberately rather than discovering accidentally.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agents are first-class operators on that roadmap, and a production-grade MCP layer is how deploy-from-chat stops being a demo. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide