Pinterest's engineers run 66,000 agent tool calls a month through a central registry (Pinterest Engineering) — because the week two teams shared one tool surface, "just point at the endpoint" stopped being an answer. On July 31, 2026, MLflow 3.15.0 shipped that lesson as an off-the-shelf feature: the MCP Registry, an experimental but fully worked catalog that registers MCP servers from server.json manifests, versions them with semver, promotes them through staging → production aliases, snapshots their tool lists, and links traces back to the exact server version that served them.
The whole pitch fits in one loop. Register version 1.0.0 of your infrastructure-tools server as a draft. Flip it active, point the staging alias at it, verify, then move the production alias — no client config changes, because endpoints resolve the alias, not the version string. Ship 1.1.0 the same way, and deprecate 1.0.0 without deleting its history. That is CI/CD's version-and-promote discipline applied to agent tool dependencies, and it is the core deliverable of this post: a worked walkthrough of that loop, the honest boundary of what the registry governs versus what it leaves to you, and why the registry question arrives the moment a second tenant's agent shares your tool surface.
The three entities, precisely
The registry is organized around three nouns, and every later behavior hangs off them:
| Entity | What it is | Example |
|---|---|---|
MCPServer | The logical server entry, named in reverse-DNS namespace format | io.github.anthropic/brave-search, com.acme/internal-tools |
MCPServerVersion | One immutable semver configuration: the frozen server_json payload, a tool snapshot, and a status | 1.0.0, status draft → active → deprecated → deleted |
MCPAccessEndpoint | A connection record — URL plus transport (streamable-http or sse) — pinned to a version or an alias | https://mcp.internal/brave-search → alias production |
Two design choices here do most of the governance work. First, MLflow stores the publisher-declared semantic version from server_json["version"] as the canonical version — it does not invent a second MLflow-specific version number, so the version your agents resolve is the version the server author shipped. Second, each version's server_json is frozen at registration: a configuration change means registering a new version, never silently mutating the old one. That immutability contract is what makes "which tool version served this request" an answerable question six months later.
The server.json format itself follows the open MCP registry standard, so you can import servers straight from the official public registry or register your own private ones. Per the design RFC, the spec references track registry repo v1.6.0 (April 15, 2026), with the upstream API frozen since October 2025 — a stable enough target to build a system of record on.
The promotion walkthrough, end to end
Here is the loop from the hook, as real SDK calls against a self-hosted tracking server:
import mlflow
# 1. Register v1.0.0. Parent MCPServer is auto-created; status starts as draft.
version = mlflow.genai.register_mcp_server(
server_json={
"name": "com.acme/deploy-tools",
"version": "1.0.0",
"description": "Read-only deploy pipeline tools for tenant agents",
"remotes": [
{"url": "https://mcp.internal/deploy-tools", "type": "streamable-http"},
],
},
)
# 2. Promote draft -> active, then attach the staging alias.
mlflow.genai.update_mcp_server_version(
name="com.acme/deploy-tools", version="1.0.0", status="active",
)
mlflow.genai.set_mcp_server_alias(
name="com.acme/deploy-tools", alias="staging", version="1.0.0",
)Status transitions are constrained, not freeform: draft can go to active or deleted, active to draft or deprecated, deprecated back to active or on to deleted. Deletion is a soft delete — the record and its history survive for audit, which matters the first time a postmortem asks what 1.0.0 contained.
# 3. Record the approved connection path, pinned to an alias — not a version.
mlflow.genai.create_mcp_access_endpoint(
server_name="com.acme/deploy-tools",
url="https://mcp.internal/deploy-tools",
transport_type="streamable-http",
server_alias="production",
)
# 4. Promote: move the alias. Clients and bindings follow without edits.
mlflow.genai.set_mcp_server_alias(
name="com.acme/deploy-tools", alias="production", version="1.0.0",
)
# 5. Ship 1.1.0 through the same staging gate, then retire 1.0.0 (history kept).
mlflow.genai.set_mcp_server_alias(
name="com.acme/deploy-tools", alias="production", version="1.1.0",
)
mlflow.genai.update_mcp_server_version(
name="com.acme/deploy-tools", version="1.0.0", status="deprecated",
)Step 4 is the payoff line. Because the access endpoint resolves the production alias rather than a pinned version string, promotion is a pointer move — the registry equivalent of retagging stable in an artifact repository. Agents that discovered the endpoint keep working; the rollback is the same call in reverse. There is also a reserved latest alias that always resolves to the highest semver among active versions (falling back to the highest non-deleted version when nothing is active), for consumers that explicitly opt into floating.
Tool snapshots and the refresh workflow
Versioning the manifest is half the inventory problem. The other half is knowing what a version actually exposes — MCP servers can and do change their tool lists — and the registry answers that with snapshots plus explicit refresh.
When you register a version without hand-supplying tools, MLflow performs live auto-discovery: it connects to the first usable remote in server_json.remotes[], calls the server's tool list, and stores the result on the version record. When the live server drifts after registration, you re-discover on demand:
# Preview what changed before committing to it.
preview = mlflow.genai.refresh_mcp_server_version_tools(
name="com.acme/deploy-tools", version="1.0.0", dry_run=True,
)
# Protected servers take auth headers for the discovery call.
updated = mlflow.genai.refresh_mcp_server_version_tools(
name="com.acme/deploy-tools",
version="1.0.0",
mcp_server_access_headers={"Authorization": "Bearer <token>"},
)Note the documented caveat, stated plainly in the docs: discovered tools are a point-in-time snapshot and are never automatically kept in sync with the live server. Discovery also needs the mcp extra installed (pip install 'mlflow[mcp]'). This is the right default for a system of record — silent re-sync would destroy exactly the "what did this version expose" evidence the snapshot exists to preserve — but it means your runbook owns a refresh step. Stale snapshots that nobody refreshes are worse than no snapshots, because they answer confidently and wrongly. The dry_run preview is the cheap version of that runbook step: diff first, save second.
Trace linking is the lineage story
Snapshots answer "what could this version do." Traces answer "what did it actually do, for whom" — the lineage leg of the promise. When a runtime participates in MLflow-aware tracing, tool calls can be associated with the governed {workspace, name, version} identity:
client.link_mcp_server_versions_to_trace(
trace_id="tr-abc123",
mcp_servers=[mcp_server_version],
)The GenAI UI surfaces this as an "MCP Servers" tab next to the existing "Prompts" tab on each trace, with name, version, and links back to the server detail page — and the reverse lookup works too, so a version page can show the traces it served. The subtle win is identity stability: because traces bind to the governed version identity rather than a raw endpoint URL, they keep rolling up correctly when endpoints move or aliases get re-pointed. Rollback analysis ("which agents touched the bad version before we deprecated it?") and lightweight auditing of which definitions were in use both fall out of the same association.
The precondition is real, though: looking a server up in the registry does not by itself create any trace association. The client or runtime has to emit MLflow-aware traces, with context propagated (for example traceparent over HTTP) if you want caller-side and runtime-side spans correlated. Lineage here is a protocol the runtime opts into, not something the registry extracts by observing traffic — because, as the next section stresses, the registry never sees traffic at all.
Registry vs gateway vs nothing: the honest boundary
This is the section that decides whether the registry fits your threat model. The design doc (RFC 0004, authored by Jon Burdo, Dan Kuc, and Matthew Prahl against MLflow issue #22625) draws the line explicitly: the registry is the control-plane system of record, and a future MLflow MCP Gateway would be the data-plane that mediates live traffic. The registry layer ships first; the gateway is designed-for, not delivered.
| Capability | Hardcoded endpoint | DNS TXT record | MLflow MCP Registry | Full MCP gateway |
|---|---|---|---|---|
| Version pinning | No — URL is the version | No | Yes — semver versions, immutable manifests | Yes |
| Promotion without client change | No — redeploy clients | Partial — retarget the record | Yes — move the alias | Yes |
| Tool inventory ("what does it expose") | No | No | Yes — snapshots + refresh | Usually |
| Lineage ("who called which version") | Logs, if any | No | Yes — trace linking, runtime permitting | Yes — sees all traffic |
| Runtime auth enforcement | Whatever the server does | None added | None added | Yes — policy at request time |
| Proxying / usage control | No | No | Explicitly out of scope | Yes |
Read the registry row carefully: it governs identity, lifecycle, and approved connection paths, but it does not provision, host, proxy, or police anything. The RFC's out-of-scope list says it outright — consumers can still connect directly to an MCP endpoint unless a gateway or proxy mediates access. An access endpoint record is an approved path, not an enforced one. For a platform letting tenant agents call infrastructure tools, that means the registry answers "which versions exist, which are blessed, and who called what" while the server (or a future gateway) still answers "who is allowed."
That boundary is also why the gateway integration is designed as derived state: when a gateway eventually ships, its deployment records will target the same governed versions and aliases rather than introducing a second catalog. One system of record now, one enforcement point later — instead of two catalogs that drift.
Two more honest footnotes. The feature is experimental as of 3.15.0, so expect API and behavior changes. And the upstream-spec compatibility router (serving the public GET /v0.1/servers API shape so IDE plugins and agent frameworks can discover MLflow-registered servers without MLflow-specific code) is deferred to Phase 2 — native MLflow APIs, with MLflow auth, workspaces, and per-resource RBAC, are the interface today.
When the second tenant arrives
So when does a platform actually need this? The spec's answer is precise: the moment a second tenant's agent shares the same tool surface. One tenant and one agent can hardcode an endpoint URL and survive, because every tool change has exactly one consumer to coordinate with. The second tenant breaks that in a specific, predictable way: an unversioned tool change that fixes tenant A's workflow silently breaks tenant B's, and with no versioned identity on either side, the postmortem cannot even name what changed under whom.
The registry's answer to that failure is per-tenant pins through the same alias machinery: tenant A's endpoint resolves production at 1.1.0 while tenant B's stays on 1.0.0 until their agent is validated against the new tools — then the alias move promotes them with a rollback that is one call, not a redeploy. Deprecation without deletion keeps the old version servable for the laggard while signaling the migration. And trace linking tells you exactly when the last 1.0.0 call happened, so "can we delete it yet" is a query, not a guess.
That is the version-and-promote discipline CI/CD pipelines already apply to artifacts, finally applied to the tool dependencies agents resolve at runtime. Pinterest built its own registry because at 66,000 calls a month there was no alternative; MLflow's bet is that the governance shape — namespaced identity, immutable versions, aliases, snapshots, trace lineage — is standard enough to ship once and run everywhere, including on a self-hosted tracking server next to your fleet. Adopt it for the catalog and the promotion loop today; keep the enforcement story honest about what still belongs to the server or the gateway to come.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agent-operated deploys need governed tool surfaces like the ones above, not hardcoded endpoints. Star the repo on GitHub or deploy your first app today.



