Two unrelated Hetzner incidents landed four days apart in late July 2026. One took machines offline. The other made it slow to create new ones. If you run a single-provider fleet on Cluster API Provider Hetzner, those two sentences sound similar on a status page and cost you completely different things.
On July 24, a network-switch fault in nbg1-cloud1-leaf88 interrupted connectivity for a slice of Nuremberg. On July 28, Hetzner's status page posted a separate "Cloud Resource Creation Delay" — the provisioning path, not the data plane, was slow. New servers queued. Existing ones kept running. For a fleet that deliberately doesn't multi-cloud, the interesting question isn't "was Hetzner down twice in a week?" It's "which of those two failure modes would have actually hurt my tenants, and which would have just made my autoscaler wait?"
This post gives you the split.
What actually happened: two incidents, two fault domains
| Signal | July 24 — nbg1-cloud1-leaf88 | July 28 — Cloud Resource Creation Delay |
|---|---|---|
| Status wording | Network-switch fault, rack-level connectivity loss | "Delays creating cloud resources" via HCloud API |
| Fault domain | Data plane — packets to/from existing hosts | Control plane — POST /v1/servers and related creation calls |
| Duration on status page | Single-digit hours (resolved same day) | Hours-long degraded creation path (no compute outage) |
| Existing VMs | Intermittent unreachable for affected hosts | Unaffected — running servers stayed reachable |
| New VMs | Creation also affected inside the rack | Creation queued/slow globally for the provisioning path |
| Fleet symptom | Node NotReady, pod rescheduling pressure | HetznerMachine stuck Provisioning, autoscaler queue growth |
Neither incident was declared a full-region outage. Both were scoped. The reason they feel like a pattern is the calendar, not the failure mode — four days is short enough for human pattern-matching to fire and long enough for two independent subsystems to fail for independent reasons.
Hetzner's public status history supports that base rate: isolated leaf/rack events and short provisioning degradations have appeared intermittently across 2024–2026 without sharing a root cause. Two tickets in one week is unusual, but the provider's incident taxonomy treats a leaf switch and a creation queue as different services. Your runbook should too.
The only table that matters: provisioning delay vs. outage
For a CAPH fleet, "Hetzner had an incident" is underspecified. The fleet has two distinct interfaces to Hetzner: the running-machine interface (packets, disks, load balancers) and the provisioning interface (the HCloud API that turns a HetznerMachine into a real server). An incident on one does not imply an incident on the other.
| Dimension | Provisioning delay (July 28) | Network/compute outage (July 24) |
|---|---|---|
| CAPH object state | HetznerMachine stays Provisioning / Machine stays Provisioning | HetznerMachine is Running but Machine / Node goes NotReady |
| Kubernetes-visible | Pending pods grow; no evictions | Running pods may be evicted / rescheduled; PV attaches can stall |
| Tenant-visible | Deploys and scale-outs are slow | Requests may fail or latency-spike for tenants on affected nodes |
| Duration cost | Time-to-capacity stretches (see next section) | Downtime/downstream error budget burned |
| Operator action | Throttle scale-out, queue, notify "deploy latency, not downtime" | Cordon/drain if needed, reschedule, incident comms for downtime |
| When multi-cloud would help | Barely — other provider also needs time to provision | Moderately — traffic could shift if you had a second substrate |
If you alert on "Hetzner incident" as a single boolean, you will page for a provisioning delay the same way you page for an outage and then spend the incident explaining to tenants that nothing was actually down. Separate the two.
What a provisioning delay costs mid-scale-out
Take a concrete fleet: 20 Hetzner Cloud servers under CAPH, spread across nbg1, fsn1, and hel1, running a multi-tenant Bex-style PaaS. Typical day: 18 machines steady-state, 2 spare. On July 28, coincident with the creation delay, tenant traffic pushes the cluster-autoscaler to request 10 more machines to go from 20 to 30.
Normal path without a delay:
Machine->HetznerMachine-> HCloudPOST /v1/servers-> serverrunning->kubeadm join->Node Ready- End-to-end p50 about 60–90 seconds per machine on Hetzner Cloud (image already cached,
Servercreation is the long pole). - 10 machines in parallel land in roughly 2–3 minutes to
Ready.
Delayed path during the July 28 creation slowdown:
POST /v1/serversreturns slower, or queues with429/ extendedprovisioningstate on the Hetzner side.- CAPH's reconciler requeues with backoff.
HetznerMachineconditions stayProvisioning: Truelonger.Machinenever reachesRunning. - The autoscaler sees pending pods and keeps the scale-out desired, but new capacity doesn't arrive. Pending-pod queue depth grows linearly with incoming deploys.
What tenants see:
- New
git pushdeploys that need a fresh placement stayPendinglonger. If your PaaS pins builds to new nodes or needs headroom,bex deployhangs at "waiting for capacity" rather than failing. - Already-running services are untouched. No restarts, no connection drops, no PV reattaches. Your error budget for availability doesn't move. Your error budget for deploy latency does.
Stretch factor observed in analogous CAPH creation-delay issues (rate-limit and provisioning backoff in cluster-api-provider-hetzner issues #2097 and related discussions): time-to-capacity can stretch from ~90 seconds to 10–30 minutes per machine while the provider's creation queue drains. Ten machines that would normally be ready in 3 minutes now arrive over 20–45 minutes, serializing behind the provider's throttle.
What it does not cost:
- No extra Hetzner bill. You are not charged for a server that hasn't entered
running. You pay for time-to-capacity, not time-in-queue. - No pod evictions or data-plane retries.
kubeleton existing nodes never notices. - No multi-cloud failover logic to exercise. Shifting a pending
HetznerMachineto another provider still requires a full create path on that provider — you save nothing unless you had warm standby capacity already running (which is a different cost).
Sensitivity check — because "a provisioning delay costs X" depends on whether you were scaling:
| Fleet state during delay | Cost |
|---|---|
| Not scaling, no deploys needing new nodes | ~0 — queue exists but nobody is waiting |
| Scaling +1–2 machines for routine headroom | 10–20 min extra deploy latency, no downtime |
| Scaling +10 machines for a traffic burst | 20–45 min to full capacity, pending-pod backlog, possible tenant deploy timeouts |
| Autoscaler + CI burst (many tenants pushing at once) | Queue amplifies — each new Machine waits behind the same creation throttle |
The honest summary: a provisioning delay taxes elasticity, not availability. If your fleet wasn't trying to grow that hour, you may not have noticed at all.
What an outage costs running tenants
Same fleet, July 24 fault. A leaf switch in nbg1-cloud1 blips. Machines already Running on that leaf lose connectivity intermittently.
NodegoesNotReadyafter the kubelet grace period. Pods withPodDisruptionBudgetsreschedule if capacity exists infsn1/hel1. Pods without spare capacity stayPendinguntil the leaf recovers or you cordon.- Hetzner Load Balancers fronting the API server may briefly lose a backend.
Cloud Controller Managernode status flaps. - Any
Volumewhose attachment is tied to the affected hypervisor host stalls until the host returns.
For a fleet with no spare headroom, a 10-minute network fault can produce 15–25 minutes of tenant-visible error rate as pods reschedule and warm caches. For a fleet with 20% headroom spread across three facilities, the same fault can be a 1–2 minute latency bump with zero hard downtime outside the affected rack.
The dollar math differs from a provisioning delay in kind, not just degree. Downtime burns error budget and, for some tenants, revenue. Deploy latency burns developer time and defers capacity but doesn't break in-flight requests. Conflating them in a postmortem makes two small provisioning delays look like two outages.
The 4-day cluster trap
Two incidents in four days feels like a pattern. Statistically, for a provider that posts a handful of scoped incidents per quarter, two in one week is well within variance — especially when the two are in different subsystems. The availability literature calls this the clustering illusion: humans overweight recent co-occurrence and underweight base rate.
A practical base-rate check for your own fleet:
- Pull Hetzner's status history for the last 12 months. Count creation-path degradations vs. network/compute degradations separately. You will typically find they don't correlate — a creation delay doesn't predict a leaf failure and vice versa.
- Count your fleet's incident-induced symptoms the same way:
Provisioningstalls vs.NotReadynodes. If both series spike together, you have a systemic provider issue. If only one spikes, you have a subsystem issue. - Normalize by exposure window. A fleet that scales aggressively will hit provisioning delays more often than a static fleet, even on the same provider. That's not the provider getting worse — it's you exercising the creation path more.
Two in four days is worth a ticket. It is not, by itself, worth a multi-cloud migration.
An operator's framework: noise or pattern worth escalating?
Use three signals before you decide "we need to change providers or add a second one."
1. Failure-domain overlap. Did the two incidents share a facility, API, or component? nbg1-cloud1-leaf88 and a global creation queue share almost nothing — different racks, different control planes. Overlap near zero argues for independent causes.
2. Blast-radius trend. Is the consequence growing? A 10-minute provisioning delay that becomes a 60-minute delay next quarter is a degradation worth escalating. Two short delays with stable MTTR is a stable provider.
3. Recovery SLO. Is time-to-healthy getting longer? For provisioning delays, track Machine Provisioning duration p50/p95. For outages, track Node NotReady duration. A rising tail, not a single point, is the pattern.
Metrics to alert on separately:
capi_hetzner_machine_provisioning_duration_seconds(or your equivalent derived fromHetznerMachinecondition transitions)- Pending-pod queue depth and age
- HCloud API
429/5xxrate onPOST /v1/servers Node NotReadycount and duration, by facility
When to escalate with Hetzner vs. when to mitigate locally:
- Escalate when MTTR is growing, overlap is high, or either series breaches your SLO for two consecutive quarters. Bring the two incident timelines, the distinct fault domains, and your provisioning-duration histogram. Ask for creation-path capacity planning, not generic "reliability."
- Mitigate locally first when overlap is low and MTTR is flat. Cheaper mitigations: spread
MachineDeploymentsacrossnbg1/fsn1/hel1so a leaf fault can't take a majority, keep 15–20% pre-warmed headroom so a provisioning delay has somewhere to queue, and cap autoscaler burst size so one traffic spike doesn't enqueue 20 creates at once.
Why this fleet deliberately doesn't multi-cloud
A second provider would have turned neither July incident into zero impact without paying for it continuously.
| Choice | What you pay | What you get for these two incidents |
|---|---|---|
| Single provider (Hetzner) + regional spread | One CAPH, one set of HetznerMachineTemplates, one bill, one API to operate | July 24: leaf fault isolated to one facility, other regions carry load. July 28: you wait — but so would any provider's creation path under load |
| Add a second provider (e.g., Hetzner + second EU bare-metal) | Second Cluster API provider, second machine template family, second set of images/networks, doubled upgrade/test surface, cross-provider scheduling/NAT complexity | July 24: could shift new workloads, but existing nbg1 pods still need rescheduling. July 28: still waiting — second provider also has to create machines, not instantly. Warm standby on provider two would help, but warm standby on Hetzner itself would help for less money |
Multi-cloud's strongest case is against a full-region loss, not a leaf fault or a creation queue. For the two July failure modes, the cheaper resilience is regional spread within Hetzner plus headroom, not a second vendor. That calculus flips when your threat model includes provider-wide control-plane loss or jurisdiction risk — then a second substrate is about blast-radius independence, not incident frequency.
Bex is explicitly built on that bet: own the machines on one EU provider via Cluster API, keep the provisioning path simple and declarative, and let regional spread plus spare capacity absorb rack-level faults. The July 28 delay is the incident that validates the bet — control-plane slowness is the cheapest possible bad day to have when you don't multi-cloud, because waiting beats failing.
Runbook: the next "Cloud Resource Creation Delay"
Copy this into your fleet's on-call notes:
- Detect the right signal. Check
HetznerMachineconditions, not node health. If machines areProvisioningand nodes areReady, it's a creation delay — not an outage. Don't page tenants for downtime. - Confirm provider-side. Check
status.hetzner.comfor a creation-path incident and samplePOST /v1/serverslatency/429 rate. Log the incident start in your own timeline. - Queue, don't thrash. Cap
MachineDeploymentscale-out burst (e.g., max 3 concurrent creates). Let CAPH's requeue backoff work. Creating 20 machines at once during a provider throttle makes the tail longer. - Communicate accurately. Tenant message: "Deploys needing new capacity are delayed ~15–30 min; running services are unaffected." Not: "Hetzner is down."
- Protect running capacity. Don't cordon healthy nodes or drain during a pure provisioning delay — you will create the outage you feared.
- Log for the vendor conversation. Record p50/p95 provisioning duration, queue depth, and API error rate for the window. If the tail grows quarter over quarter, that's your escalation artifact — not "two incidents in one week."
Four days, two incidents, one lesson: not every Hetzner status post is the same incident type, and treating them as one series is how a single-provider fleet talks itself into a second provider it doesn't need. Measure provisioning duration and node health as separate series, keep enough headroom that waiting is an option, and save the multi-cloud conversation for the failure mode where waiting isn't.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



