On September 9, 2026, every MCP connector OptimNow ran went dark at the same instant — not with a timeout or a 500, but with an HTTP 402 on initialize. One shared free quota on their hosted MCP platform ran dry, and because the quota was pooled across connectors, a single billing boundary became a single point of failure for all of their agents' tools at once. The fix landed the same day: a 14-file commit moving the Python connector off the hosted platform onto Fly.io with scale-to-zero, a Dockerfile that syncs bundled data at build time and fails the build if the sync comes out short, and a clean-environment verification pass.
Here is the verdict before the why: if your agents depend on an MCP server in production, own its quota boundary. A hosted connector platform is a fine place to launch, but a shared pool means your tools inherit someone else's accounting — one exhausted quota, and initialize fails everywhere simultaneously with no graceful degradation. This post walks through the incident, why 402-on-initialize is a total-loss failure mode, the anatomy of the actual 14-file move, and the self-hosting checklist it implies.
September 9: every connector returns 402
OptimNow's cloud-finops-skills is an open-source FinOps knowledge skill and MCP server for AI agents — Claude, ChatGPT, Gemini, Cursor, anything MCP-compatible — covering cloud cost optimization across AWS, Azure, GCP, and OCI plus AI cost management and Kubernetes spend. The team hosts its connectors on Alpic, the cloud platform for MCP Apps from the company behind Skybridge: one-command deploys, Streamable HTTP endpoints on *.alpic.live, analytics, logs, and a public playground. For a small team shipping agent tools, that bargain is exactly right — until the billing boundary bites.
The bite came on September 9. Alpic's shared free quota, pooled across every OptimNow connector, exhausted — and every connector started answering initialize with HTTP 402.
Note what that means mechanically: initialize is the first handshake in the MCP session lifecycle, the call where client and server agree on protocol version and capabilities. A 402 there doesn't degrade one tool or slow one query. It means the client never gets a capability list at all — from the agent's perspective, the tools don't exist.
The response was swift and public. The same day, OptimNow merged PR #195, "Move the hosted connector from Alpic to Fly.io": 14 files changed, 123 additions, 108 deletions, moving the Python connector to https://cloud-finops-mcp.fly.dev/mcp. The commit message states the cause plainly — "Alpic's shared free quota took every OptimNow connector down" — and the verification line reads like a release checklist: six tools, 35 references, 1.36.0 content, sandbox domain matching the new URL.
This is a small incident at a small project, and that is precisely why it matters. There was no cascading microservices topology, no multi-region failover to second-guess. One pool, one exhaustion, total loss — the failure mode in its purest form.
A shared quota is a correlated-failure mode
The hosted-connector bargain looks like this: the platform handles TLS, process supervision, logging, and a public URL, and you pay per request, per connector, or per seat — or ride a free quota while you grow. What the OptimNow incident exposes is the correlation structure hiding inside the word "shared." When N connectors draw from one pool, their availabilities are not independent. The pool is a single fate-sharing domain, and exhaustion is synchronized by construction.
For MCP specifically, the blast radius is worse than it looks, because where the 402 lands in the lifecycle determines everything:
| Failure point | What the agent sees | Severity |
|---|---|---|
402 on tools/call (one tool) | That tool errors; the agent can retry, replan, or use other tools | Degraded |
402 on tools/list | No tools advertised; session alive but useless | Severe |
402 on initialize | No session at all; the connector might as well not exist | Total loss |
402 on initialize, N connectors, one pool | Every connector gone in the same instant | Total, correlated loss |
A per-connector quota or a per-key budget fails one connector at a time — still an outage, but a legible one with a blast radius you can reason about. A shared pool fails everything at once, including the connectors you weren't watching, at whatever hour the pool happens to run dry. Your agent's tools inherit someone else's quota pool, and quota pools don't page you before they empty.
This is not an argument that hosted MCP is broken. It is an argument that the quota boundary is part of your availability design, whether you chose it deliberately or accepted the default. An April 2026 analysis of 2,181 public remote MCP endpoints found 52% completely dead and only 9% fully healthy — hosting an MCP server is easy, keeping one alive is the actual problem. And as The New Stack put it, when remote MCP servers fail it is a systemic failure that cascades across the whole agentic workflow: one upstream going offline can stall an entire plan execution. A shared quota turns that single-upstream risk into an all-upstreams-at-once event.
Anatomy of the 14-file move
The fix is worth studying file by file, because it is a complete template for "hosted connector to self-hosted container" in miniature. The commit does five things:
1. A Dockerfile that builds from the checkout and pins data at build time. The image bundles the connector's reference data by syncing it during the build — and the build fails if the sync comes out short. That guard is the most underrated line in the whole diff: it converts "silently shipped with stale or partial data" from a runtime mystery into a build-time error. Data drift becomes a red CI job instead of a wrong answer from your agent weeks later.
2. A fly.toml that scales to zero. The Fly.io config stops the machine when idle and starts it on the first request (auto_stop_machines, auto_start_machines, min_machines_running = 0). For an MCP connector — bursty, session-oriented traffic with long idle stretches — this is the correct cost shape: you pay for compute only while an agent is actually mid-session, instead of renting a 24/7 process or drawing from someone else's free pool. The tradeoff is cold start on the first initialize after idle, typically seconds on Fly's Machines platform — acceptable for an interactive agent tool, and a known quantity you control rather than a quota you share.
3. A .dockerignore that trims the build context. Small, but load-bearing: the image contains the connector and its synced data, not the repo's docs, tests, and history. Reproducible builds start with knowing exactly what went into the context.
4. The URL-derived constant and everything downstream of it. The connector URL moves from *.alpic.live to cloud-finops-mcp.fly.dev/mcp, and the commit chases that change through the constant deriving the MCP Apps sandbox domain (ui.domain), its test, the README, installation docs, the PyPI README, server.json, the dependency map, and the release checklist — while deleting alpic.json and the old Horizon entrypoint. This is the unglamorous 80% of any hosting move: the endpoint URL is load-bearing in docs, client configs, and security boundaries (the sandbox domain gates widget rendering), so a move is really a coordinated rename with a deploy attached. Pinning the derived domain with a test is what makes the rename safe to repeat.
5. Verification in a clean environment. Before calling it done, the team re-ran the same steps as the Dockerfile in a fresh environment and asserted the outcome: six tools, 35 references, 1.36.0 content, sandbox domain matching the new URL. That is a deploy smoke test expressed as content invariants — not "did the container start" but "does this server serve exactly the tools and data we shipped." Any MCP deploy pipeline can steal this pattern verbatim.
Fourteen files, net +15 lines, and the connector is back with its quota boundary owned outright. The whole move is smaller than most teams' Terraform for a single load balancer.
The self-hosting checklist for production MCP servers
Generalize the diff and you get a checklist for any team running MCP servers its agents depend on in production. Each item maps to something the September 9 incident proved necessary:
- Own the quota and billing boundary. Know exactly which pool your connector draws from, who else draws from it, and what happens at exhaustion — which lifecycle call fails, and whether it fails one connector or all of them. If the answer is "a shared free pool," you have a correlated-failure mode with a calendar date on it. A paid tier with per-connector budgets, or your own machines where the only quota is the one you set, both fix the correlation.
- Pin data at build time; fail the build if it comes out short. If your server bundles reference data, sync it during the image build with an assertion on completeness. Never let a deploy succeed with partial data — a connector serving 30 of 35 references doesn't error, it just answers wrong.
- Scale to zero when idle. Session-oriented agent traffic is bursty.
min_machines_running = 0with autostart gives you near-zero idle cost on your own infrastructure, which removes the economic argument for parking production tools on someone else's free tier. Measure the cold-start latency on firstinitializeand decide explicitly whether your agents tolerate it. - Smoke-test content invariants, not just liveness. After every deploy, assert the things your agents actually consume: tool count and names, reference/data version, the sandbox or auth domain matching the live URL. A 200 on
/mcpproves the process runs; it says nothing about whether the tools are the ones you shipped. - Treat the endpoint URL as a coordinated rename. Document every place the URL is load-bearing — client configs, connector registries, sandbox-domain derivations, docs — and cover the derived values with tests. The next move (and there is always a next move) should be a find-and-test pass, not an archaeology expedition.
- Keep a local escape hatch. The OptimNow connector also ships as a PyPI package, so any user can run it over stdio with no hosting dependency at all. For critical agent tools, a stdio or local-docker path means a hosting outage degrades to "reconfigure the client," not "the tools are gone."
None of this requires operating Kubernetes or hiring a platform team. It requires a Dockerfile, a fly.toml, and the discipline to assert on content — roughly one focused afternoon, judging by the size of the diff that proved it.
When hosted still wins
Self-hosting everything on day one would be its own failure mode — premature infrastructure is how side projects die. The honest decision table looks like this:
| Stage | Hosted connector platform | Self-hosted (Fly.io, VPS, your cluster) |
|---|---|---|
| Prototype / demo | Yes — one-command deploy, playground, zero ops | Overkill |
| Growing usage, tools still optional | Fine, but move off shared free pools to paid per-connector budgets | Consider it |
| Agents depend on these tools in production | Only with isolated quotas, usage alerts, and an exit plan | Yes — own the boundary |
| Regulated data / custom auth / VPC | Usually no — you need the network and the audit trail | Yes |
The through-line: match the hosting to the blast radius. When the tools are a demo, borrow someone else's quota pool with open eyes. The moment an agent's production workflow stalls because initialize returned a billing status code, you have outgrown the pool — and the fix, as September 9 showed, is a 14-file afternoon.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Self-hosting your agent's MCP servers is exactly the kind of deploy-from-git workflow it exists for. Star the repo on GitHub or deploy your first app today.



