Skip to main content

The Agent That Burned $4,200 in 63 Hours: What a Deploy/Rollback MCP Server Needs for Spend Circuit Breakers

13 min readDora NodaDora Noda
Share
On this page

Friday, 5pm: a developer kicks off an agentic coding session and heads home. Monday, 8am — 63 hours later — the bill is $4,200. Nobody watched it happen: every API call was small, every step plausibly useful, the loop never errored. It just never stopped.

That $4,200 weekend is real — the LeanOps spring-2026 audit of 30 engineering teams found median spend $480/developer/month, 90th percentile $1,650, and outliers past $4,200, one over a single long weekend. The postmortem traces it to one load-bearing assumption: the user will notice and stop it in time. No hard spend ceiling sat anywhere between the agent's decisions and the infrastructure it consumed — so at 3am Sunday, with no human watching, the loop kept going.

This is a class, not an anecdote: an estimated $400M in unbudgeted Fortune-500 cloud spend from agentic overruns, Uber's entire 2026 AI budget gone in four months of Claude Code rollout, a $47K multi-agent loop that ran 11 days with every per-request limit green. OWASP's agentic guidance now mandates circuit breakers in every agentic workflow — because "a human will stop it" keeps proving to be the same unsafe assumption underneath incident after incident.

Here is this post's core deliverable, up front: the five-layer circuit-breaker design a deploy/rollback MCP server sitting in front of a Cluster API fleet needs. Everything after it is the evidence and the mechanism behind each row.

LayerControlExample knobWhat it stops
1. Identity + ledgerPer-agent identity with a server-side cumulative spend ledgerAtomic dollar/token counters in shared state; fail closed on ledger failureThe $4,200 weekend: no agent spends without an account it is charged against
2. Chain budgetsPer-session/per-chain cumulative caps, not per-request capsmax_budget_usd per session; token cap; input/output-ratio tripwire above ~30:1Multi-step loops and fan-out that stay green on every single call
3. Tool quotas + approvalsPer-tool call-count caps, wall-clock timeouts, human approval on irreversible actionsMax N calls per tool per run; deadline kills the runaway; MCP approval gate on destroy/rollback-to-prodRetry storms; the 9-second production-volume deletion
4. Infrastructure backstopPer-agent-namespace CPU/memory quotas and default container limitsResourceQuota + LimitRange; default-deny egress; job deadlinesThe runaway that escapes the software ledger gets throttled or OOM-killed
5. Audit + kill switchImmutable action log and a tested global stopWho did what, when, at what cost; one command halts the fleetPost-incident reconstruction; the 3am "make it stop" moment

Why "a Human Will Notice" Fails​

The assumption has two halves — detection and reaction — and production data breaks both.

Detection is slow: DigitalApplied's H1 2026 retrospective across 50+ public AI incidents found median time-to-detect ~4.5 hours and median time-to-contain ~12 hours, with cost spikes as the dominant signal. Note the shape: the bill is the alarm. Tool-misuse incidents were minor when caught at minute 4, catastrophic at hour 4 — and the $4,200 weekend was caught at hour 63.

Reaction is slower than the damage. The PocketOS incident of April 2026 is the extreme case: a coding agent hit a credential mismatch in staging and, with no confirmation prompt and a CLI token scoped to everything, deleted the production volume and all volume-stored backups in 9 seconds. Three months of booking data, gone in 9 seconds — no on-call rotation reacts that fast (Railway's CEO personally restored it after the story went public). When the loop between decision and irreversible side effect is measured in seconds, "the user will notice" is not a control. It is a hope.

And the model will not save itself. The BAGEN benchmark (May 2026) tested frontier agents on budget awareness across four environments and found capability does not correlate with budget discipline (r=0.35): models are systematically over-optimistic about remaining budget, and calibration tops out at 47% even after targeted training. OWASP's AISVS guidance draws the only possible conclusion: budget enforcement must live in the runtime, not in the model's judgment. A deploy-authority MCP server is exactly that runtime — the choke point every agent action flows through — which is why the ceiling belongs there, server-side, where the agent cannot negotiate with it.

The Anatomy of a Runaway​

Per-request limits feel like protection. Every incident in this class shows why they are not: the spend that kills you is cumulative across a chain, and each link looks innocent.

First, context compounding. LeanOps' decomposition is the number to remember: a 5-step agent loop costs ~3.2× a single call, a 50-step loop exceeds 30×, and a 200-step loop exceeds 100× — because each step re-sends the entire conversation history. Fully 62% of agent bills in their sample was re-sent context. Step caps alone leave most of that curve unenforced; the April 2026 Claude Code CLI bug made it visceral, ingesting ~30M input tokens in one session at a 74:1 input/output ratio (normal sessions sit at 5:1 to 15:1) and flipping an API balance from +$8.46 to −$13.23 while the assistant's visible outputs stayed small.

Second, delegation fan-out. An arXiv catalog of 63 confirmed production budget-overrun incidents across 21 frameworks (June 2026) attributes 11 of them to the "delegation-fanout race": sub-agents spending concurrently against a budget none of them can see in full. The $47K LangChain loop is the canonical shape — four agents, each under its own per-request limits, collectively burning $47K over 11 days with no cross-chain counter anywhere. Concurrent budget delegation without a shared ledger is just four $4,200 weekends happening at once.

Third, retry storms: the February 2026 $47K incident needed no intelligence, just a loop that retried failures with the full context attached — each retry re-tokenizing history at ~10× the cost, surfacing as a budget anomaly that error-rate alerts never catch.

A human-paced dashboard shows all of this beautifully, after the fact. The five layers below are ordered so each catches what the previous one misses.

Layers 1–2: Identity, Ledger, and Chain Budgets​

Layer 1 is the foundation everything else charges against: every agent gets an identity, and every identity gets a server-side ledger that accumulates dollars and tokens across all of its sessions, chains, and sub-agents. Three properties matter. The ledger must be server-side — enforced by the MCP server or a gateway in front of it, never self-reported by the agent. Its counters must be atomic and shared across replicas (Redis-backed, not per-process memory), or one replica tripping the breaker won't stop the others. And it must fail closed: if the ledger is unreachable, new spend stops rather than flowing unmeasured. A budget control that fails open is a suggestion.

Layer 2 puts the actual ceilings on that ledger, and the key word is cumulative. A per-request cap of $1 means nothing to a loop that makes 4,200 requests; a per-session max_budget_usd ends that loop at a number finance chose in advance. The same goes for tokens: per-session and per-chain token caps bound the context-compounding curve from the previous section. And the cheapest anomaly signal in the whole design is the input/output ratio — sustained ratios above ~30:1 mean the agent is ingesting a runaway transcript while producing little, the exact signature of the Claude Code 74:1 incident. That tripwire costs one division per session and would have paged someone on Saturday morning instead of Monday.

One caveat: provider token counts are not a trustworthy sole input — a May 2026 client regression silently inflated them ~40%. Cross-check against gateway-side estimation, pin client versions, and alert on divergence. The ledger is only as honest as its cheapest input.

Layer 3: Tool Quotas and Approval Gates​

Some damage is not measured in dollars. PocketOS's 9-second deletion cost nothing in tokens and destroyed three months of data — which is why Layer 3 bounds actions, not just spend.

Per-tool quotas are the mechanical part: each tool gets its own call-count cap per run, its own wall-clock timeout, and its own compute boundary. A deploy tool that may be called 3 times per run cannot retry-storm; a tool call with a 30-second deadline that actually kills the process (not just the client connection — Cloud Run's docs warn explicitly that an HTTP 504 can close the connection while the container keeps processing) cannot run past its supervision. OWASP's AISVS 9.1.1 says to verify these per tool with negative probes — a CPU spinner, a memory allocator, a disk filler, a denied egress attempt, a sleeper — and to record that the breach terminated only that tool while a follow-up call to another tool still succeeded. That last check is what proves the quota is per-tool rather than node-wide.

Approval gates are the human part, and they belong on exactly the actions whose blast radius is irreversible: destroy, delete without backup, rollback-to-production, privilege grants, outbound payments. The MCP protocol already has the primitive — elicitation/approval flows that pause a tool call for human confirmation — The May 2026 Bankr/Grok wallet incident ($155K+ moved with no per-transaction cap or allowlist) shows what "valid-looking action, no gate" costs when tools move value instead of tokens. Scope the credentials the same way: the PocketOS agent found a CLI token whose blast radius matched its autonomy. Least-privilege tool credentials plus a confirmation prompt on the destructive five percent of actions would have reduced that incident to a denied call and a log line.

Layer 4: The Infrastructure Backstop​

Every software control above can fail — the ledger can be bypassed, the gateway can be misconfigured, the agent can find a tool nobody capped. Layer 4 is the floor that catches the fall: Kubernetes-native resource bounds on the namespace each agent runs in.

The pairing is ResourceQuota (the namespace ceiling: total CPU, memory, storage, object counts the agent's workloads may consume) plus LimitRange (per-pod guardrails: default requests/limits for containers that don't specify them, min/max bounds that reject absurd values). Together they mean a runaway that escapes every dollar-denominated control still gets CPU-throttled by the kernel cgroup or OOM-killed on memory — noisy and visible, which is precisely what you want at 3am. Add default-deny egress NetworkPolicy with narrow allow rules so a compromised loop cannot exfiltrate or phone home freely, and activeDeadlineSeconds on Jobs so runaway batch work terminates instead of accumulating.

The point is different failure semantics: the kernel and the kubelet enforce this layer, not any policy process the agent shares fate with. When your Cluster API fleet provisions per-tenant or per-agent namespaces, quota and limit-range objects should ship in the same manifest as the namespace itself — no namespace without a ceiling, the way no agent gets Layer 1 without an identity. The platform team that owns the machines can guarantee this in a way no SaaS dashboard ever could: the machine is the last budget enforcer.

Layer 5: Audit Trail and Kill Switch​

After every incident above, the first question was some version of "what did it do?" Layer 5 answers that question in minutes instead of days: an append-only log of every tool call with agent identity, timestamp, arguments, and cost, feeding both the Layer 2 counters and a human-readable trail. If a human cannot answer what happened, why, and what changed, the system is not production-ready — and with 65% of organizations reporting an agent-caused incident last year but only 21% able to say what their agents are doing at runtime, most fleets fail that test today.

The kill switch is the other half: one tested command halting all agent spend fleet-wide, exercised in game days — and a breaker firing mid-deploy must leave state consistent, which is why OWASP pairs circuit breakers with transactional rollback. Mind the fox guarding the henhouse: two June 2026 LiteLLM vulnerabilities (CVE-2026-42208 dumping the keys that bypass spend limits, exploited within 36 hours; CVE-2026-42271, RCE on the proxy host) proved the budget gateway is itself high-value attack surface. Its patch level is budget evidence, not just hygiene.

What Each Layer Would Have Caught​

IncidentWithout controlsLayer that stops itResidual damage
$4,200 long weekend (63h loop)Discovered Monday; full $4,200 goneLayer 2 session dollar cap + ratio tripwireSpend stops at the cap (e.g. $50); page fires Saturday morning
PocketOS 9-second deletionProduction volume + backups destroyedLayer 3 approval gate on destroy + scoped tokenOne denied call; a log line; an annoyed agent
$47K LangChain loop (11 days)Four agents, all per-request limits greenLayer 1 shared cross-chain ledgerLoop halts at the fleet-wide cumulative ceiling
$47K retry storm (2.3M calls)Error-rate alerts never firedLayer 3 per-tool call-count cap + wall-clock deadlineStorm ends after N calls; Layer 4 throttles the rest
Uber-scale fleet burn ($2.5–10M/mo)Annual budget gone in 4 monthsLayers 1+2 per-engineer/per-agent gateway budgetsToken volume compounds against ceilings, not against finance's surprise

No single layer covers the table. That is the argument for building all five: each row's "layer that stops it" is a different layer, and the incidents keep arriving in whichever column you left empty.

Build vs. Borrow​

Good news: roughly half of this design exists as off-the-shelf gateway features. LiteLLM supports tag- and agent-scoped budgets with per-session dollar caps; MLflow AI Gateway added budget policies with daily/weekly/monthly windows on Redis-backed shared state; Traefik Hub's Triple Gate does proactive token estimation that blocks abusive requests before they reach the model. Adopt one of these for Layers 1–2 rather than hand-rolling a ledger — but treat the gateway binary, its config, and its patch level as part of the control, per the LiteLLM CVEs above.

What you must build is the deploy-authority part: approval gates on destructive tools, per-agent-namespace quota wiring in your Cluster API manifests, the audit schema joining "agent X called deploy Y" to "it cost Z," and the kill switch. Nobody else's gateway knows which of your tools deletes a production volume. That mapping — tool inventory to blast radius to approval policy — is the platform-specific work, and it is also the highest-leverage: the PocketOS row in the table above is one gate on one tool.

Start this week with three moves needing no new infrastructure: a per-session dollar cap on every agent session, the 30:1 ratio alert, and a human gate on your single most destructive tool. Those three cover four of the five rows above. The full five layers can follow — but the next $4,200 weekend is already ticking.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agent operators that deploy to it deserve infrastructure that bounds what an agent can spend and destroy; star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide