Skip to main content

Two Incidents in Four Days: What Hetzner's Provisioning Delay Actually Costs When You Don't Multi-Cloud

12 min readDora NodaDora Noda
Share
On this page

Two unrelated Hetzner incidents landed four days apart in late July 2026. One took machines offline. The other made it slow to create new ones. If you run a single-provider fleet on Cluster API Provider Hetzner, those two sentences sound similar on a status page and cost you completely different things.

On July 24, a network-switch fault in nbg1-cloud1-leaf88 interrupted connectivity for a slice of Nuremberg. On July 28, Hetzner's status page posted a separate "Cloud Resource Creation Delay" — the provisioning path, not the data plane, was slow. New servers queued. Existing ones kept running. For a fleet that deliberately doesn't multi-cloud, the interesting question isn't "was Hetzner down twice in a week?" It's "which of those two failure modes would have actually hurt my tenants, and which would have just made my autoscaler wait?"

This post gives you the split.


What actually happened: two incidents, two fault domains​

SignalJuly 24 — nbg1-cloud1-leaf88July 28 — Cloud Resource Creation Delay
Status wordingNetwork-switch fault, rack-level connectivity loss"Delays creating cloud resources" via HCloud API
Fault domainData plane — packets to/from existing hostsControl plane — POST /v1/servers and related creation calls
Duration on status pageSingle-digit hours (resolved same day)Hours-long degraded creation path (no compute outage)
Existing VMsIntermittent unreachable for affected hostsUnaffected — running servers stayed reachable
New VMsCreation also affected inside the rackCreation queued/slow globally for the provisioning path
Fleet symptomNode NotReady, pod rescheduling pressureHetznerMachine stuck Provisioning, autoscaler queue growth

Neither incident was declared a full-region outage. Both were scoped. The reason they feel like a pattern is the calendar, not the failure mode — four days is short enough for human pattern-matching to fire and long enough for two independent subsystems to fail for independent reasons.

Hetzner's public status history supports that base rate: isolated leaf/rack events and short provisioning degradations have appeared intermittently across 2024–2026 without sharing a root cause. Two tickets in one week is unusual, but the provider's incident taxonomy treats a leaf switch and a creation queue as different services. Your runbook should too.


The only table that matters: provisioning delay vs. outage​

For a CAPH fleet, "Hetzner had an incident" is underspecified. The fleet has two distinct interfaces to Hetzner: the running-machine interface (packets, disks, load balancers) and the provisioning interface (the HCloud API that turns a HetznerMachine into a real server). An incident on one does not imply an incident on the other.

DimensionProvisioning delay (July 28)Network/compute outage (July 24)
CAPH object stateHetznerMachine stays Provisioning / Machine stays ProvisioningHetznerMachine is Running but Machine / Node goes NotReady
Kubernetes-visiblePending pods grow; no evictionsRunning pods may be evicted / rescheduled; PV attaches can stall
Tenant-visibleDeploys and scale-outs are slowRequests may fail or latency-spike for tenants on affected nodes
Duration costTime-to-capacity stretches (see next section)Downtime/downstream error budget burned
Operator actionThrottle scale-out, queue, notify "deploy latency, not downtime"Cordon/drain if needed, reschedule, incident comms for downtime
When multi-cloud would helpBarely — other provider also needs time to provisionModerately — traffic could shift if you had a second substrate

If you alert on "Hetzner incident" as a single boolean, you will page for a provisioning delay the same way you page for an outage and then spend the incident explaining to tenants that nothing was actually down. Separate the two.


What a provisioning delay costs mid-scale-out​

Take a concrete fleet: 20 Hetzner Cloud servers under CAPH, spread across nbg1, fsn1, and hel1, running a multi-tenant Bex-style PaaS. Typical day: 18 machines steady-state, 2 spare. On July 28, coincident with the creation delay, tenant traffic pushes the cluster-autoscaler to request 10 more machines to go from 20 to 30.

Normal path without a delay:

  • Machine -> HetznerMachine -> HCloud POST /v1/servers -> server running -> kubeadm join -> Node Ready
  • End-to-end p50 about 60–90 seconds per machine on Hetzner Cloud (image already cached, Server creation is the long pole).
  • 10 machines in parallel land in roughly 2–3 minutes to Ready.

Delayed path during the July 28 creation slowdown:

  • POST /v1/servers returns slower, or queues with 429 / extended provisioning state on the Hetzner side.
  • CAPH's reconciler requeues with backoff. HetznerMachine conditions stay Provisioning: True longer. Machine never reaches Running.
  • The autoscaler sees pending pods and keeps the scale-out desired, but new capacity doesn't arrive. Pending-pod queue depth grows linearly with incoming deploys.

What tenants see:

  • New git push deploys that need a fresh placement stay Pending longer. If your PaaS pins builds to new nodes or needs headroom, bex deploy hangs at "waiting for capacity" rather than failing.
  • Already-running services are untouched. No restarts, no connection drops, no PV reattaches. Your error budget for availability doesn't move. Your error budget for deploy latency does.

Stretch factor observed in analogous CAPH creation-delay issues (rate-limit and provisioning backoff in cluster-api-provider-hetzner issues #2097 and related discussions): time-to-capacity can stretch from ~90 seconds to 10–30 minutes per machine while the provider's creation queue drains. Ten machines that would normally be ready in 3 minutes now arrive over 20–45 minutes, serializing behind the provider's throttle.

What it does not cost:

  • No extra Hetzner bill. You are not charged for a server that hasn't entered running. You pay for time-to-capacity, not time-in-queue.
  • No pod evictions or data-plane retries. kubelet on existing nodes never notices.
  • No multi-cloud failover logic to exercise. Shifting a pending HetznerMachine to another provider still requires a full create path on that provider — you save nothing unless you had warm standby capacity already running (which is a different cost).

Sensitivity check — because "a provisioning delay costs X" depends on whether you were scaling:

Fleet state during delayCost
Not scaling, no deploys needing new nodes~0 — queue exists but nobody is waiting
Scaling +1–2 machines for routine headroom10–20 min extra deploy latency, no downtime
Scaling +10 machines for a traffic burst20–45 min to full capacity, pending-pod backlog, possible tenant deploy timeouts
Autoscaler + CI burst (many tenants pushing at once)Queue amplifies — each new Machine waits behind the same creation throttle

The honest summary: a provisioning delay taxes elasticity, not availability. If your fleet wasn't trying to grow that hour, you may not have noticed at all.


What an outage costs running tenants​

Same fleet, July 24 fault. A leaf switch in nbg1-cloud1 blips. Machines already Running on that leaf lose connectivity intermittently.

  • Node goes NotReady after the kubelet grace period. Pods with PodDisruptionBudgets reschedule if capacity exists in fsn1/hel1. Pods without spare capacity stay Pending until the leaf recovers or you cordon.
  • Hetzner Load Balancers fronting the API server may briefly lose a backend. Cloud Controller Manager node status flaps.
  • Any Volume whose attachment is tied to the affected hypervisor host stalls until the host returns.

For a fleet with no spare headroom, a 10-minute network fault can produce 15–25 minutes of tenant-visible error rate as pods reschedule and warm caches. For a fleet with 20% headroom spread across three facilities, the same fault can be a 1–2 minute latency bump with zero hard downtime outside the affected rack.

The dollar math differs from a provisioning delay in kind, not just degree. Downtime burns error budget and, for some tenants, revenue. Deploy latency burns developer time and defers capacity but doesn't break in-flight requests. Conflating them in a postmortem makes two small provisioning delays look like two outages.


The 4-day cluster trap​

Two incidents in four days feels like a pattern. Statistically, for a provider that posts a handful of scoped incidents per quarter, two in one week is well within variance — especially when the two are in different subsystems. The availability literature calls this the clustering illusion: humans overweight recent co-occurrence and underweight base rate.

A practical base-rate check for your own fleet:

  • Pull Hetzner's status history for the last 12 months. Count creation-path degradations vs. network/compute degradations separately. You will typically find they don't correlate — a creation delay doesn't predict a leaf failure and vice versa.
  • Count your fleet's incident-induced symptoms the same way: Provisioning stalls vs. NotReady nodes. If both series spike together, you have a systemic provider issue. If only one spikes, you have a subsystem issue.
  • Normalize by exposure window. A fleet that scales aggressively will hit provisioning delays more often than a static fleet, even on the same provider. That's not the provider getting worse — it's you exercising the creation path more.

Two in four days is worth a ticket. It is not, by itself, worth a multi-cloud migration.


An operator's framework: noise or pattern worth escalating?​

Use three signals before you decide "we need to change providers or add a second one."

1. Failure-domain overlap. Did the two incidents share a facility, API, or component? nbg1-cloud1-leaf88 and a global creation queue share almost nothing — different racks, different control planes. Overlap near zero argues for independent causes.

2. Blast-radius trend. Is the consequence growing? A 10-minute provisioning delay that becomes a 60-minute delay next quarter is a degradation worth escalating. Two short delays with stable MTTR is a stable provider.

3. Recovery SLO. Is time-to-healthy getting longer? For provisioning delays, track Machine Provisioning duration p50/p95. For outages, track Node NotReady duration. A rising tail, not a single point, is the pattern.

Metrics to alert on separately:

  • capi_hetzner_machine_provisioning_duration_seconds (or your equivalent derived from HetznerMachine condition transitions)
  • Pending-pod queue depth and age
  • HCloud API 429 / 5xx rate on POST /v1/servers
  • Node NotReady count and duration, by facility

When to escalate with Hetzner vs. when to mitigate locally:

  • Escalate when MTTR is growing, overlap is high, or either series breaches your SLO for two consecutive quarters. Bring the two incident timelines, the distinct fault domains, and your provisioning-duration histogram. Ask for creation-path capacity planning, not generic "reliability."
  • Mitigate locally first when overlap is low and MTTR is flat. Cheaper mitigations: spread MachineDeployments across nbg1/fsn1/hel1 so a leaf fault can't take a majority, keep 15–20% pre-warmed headroom so a provisioning delay has somewhere to queue, and cap autoscaler burst size so one traffic spike doesn't enqueue 20 creates at once.

Why this fleet deliberately doesn't multi-cloud​

A second provider would have turned neither July incident into zero impact without paying for it continuously.

ChoiceWhat you payWhat you get for these two incidents
Single provider (Hetzner) + regional spreadOne CAPH, one set of HetznerMachineTemplates, one bill, one API to operateJuly 24: leaf fault isolated to one facility, other regions carry load. July 28: you wait — but so would any provider's creation path under load
Add a second provider (e.g., Hetzner + second EU bare-metal)Second Cluster API provider, second machine template family, second set of images/networks, doubled upgrade/test surface, cross-provider scheduling/NAT complexityJuly 24: could shift new workloads, but existing nbg1 pods still need rescheduling. July 28: still waiting — second provider also has to create machines, not instantly. Warm standby on provider two would help, but warm standby on Hetzner itself would help for less money

Multi-cloud's strongest case is against a full-region loss, not a leaf fault or a creation queue. For the two July failure modes, the cheaper resilience is regional spread within Hetzner plus headroom, not a second vendor. That calculus flips when your threat model includes provider-wide control-plane loss or jurisdiction risk — then a second substrate is about blast-radius independence, not incident frequency.

Bex is explicitly built on that bet: own the machines on one EU provider via Cluster API, keep the provisioning path simple and declarative, and let regional spread plus spare capacity absorb rack-level faults. The July 28 delay is the incident that validates the bet — control-plane slowness is the cheapest possible bad day to have when you don't multi-cloud, because waiting beats failing.


Runbook: the next "Cloud Resource Creation Delay"​

Copy this into your fleet's on-call notes:

  1. Detect the right signal. Check HetznerMachine conditions, not node health. If machines are Provisioning and nodes are Ready, it's a creation delay — not an outage. Don't page tenants for downtime.
  2. Confirm provider-side. Check status.hetzner.com for a creation-path incident and sample POST /v1/servers latency/429 rate. Log the incident start in your own timeline.
  3. Queue, don't thrash. Cap MachineDeployment scale-out burst (e.g., max 3 concurrent creates). Let CAPH's requeue backoff work. Creating 20 machines at once during a provider throttle makes the tail longer.
  4. Communicate accurately. Tenant message: "Deploys needing new capacity are delayed ~15–30 min; running services are unaffected." Not: "Hetzner is down."
  5. Protect running capacity. Don't cordon healthy nodes or drain during a pure provisioning delay — you will create the outage you feared.
  6. Log for the vendor conversation. Record p50/p95 provisioning duration, queue depth, and API error rate for the window. If the tail grows quarter over quarter, that's your escalation artifact — not "two incidents in one week."

Four days, two incidents, one lesson: not every Hetzner status post is the same incident type, and treating them as one series is how a single-provider fleet talks itself into a second provider it doesn't need. Measure provisioning duration and node health as separate series, keep enough headroom that waiting is an option, and save the multi-cloud conversation for the failure mode where waiting isn't.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex