On March 30, 2023, Hetzner's backbone link between Nuremberg and Falkenstein faulted — twice in one day, once for nearly four hours (StatusGator's Hetzner Networks log still carries both entries). Nobody's servers caught fire. But for a few hours, two of Hetzner's three EU datacenters had a degraded path between them, and every operator with infrastructure split across that link got an unplanned answer to the question: what happens to my cluster when a datacenter stops cooperating?
Most single-DC clusters never have to answer that question, right up until they do. A reference project by @misterkuka — k0s-hetzner-boilerplate-multizone — answers it in advance: a fully-HA k0s cluster spread across three Hetzner EU datacenters (Falkenstein, Nuremberg, Helsinki), provisioned with Terraform, configured with Ansible, and run day-to-day by ArgoCD, in which no single node and no single datacenter can take the cluster down. This post prices that promise honestly: the infra bill, the latency tax, the idle headroom, and the complexity you now own — against the single-DC default and its backup-plus-reprovision escape hatch.
The answer up front
| Single-DC k0s (the default) | 3-DC multizone k0s (the reference) | |
|---|---|---|
| Infra cost | ~$100/mo and up for an HA-shaped build; the project's own single-zone sibling is the baseline | ≈ $103/mo for 14 servers: 3 controllers, 3 workers, 3 Postgres, 2 DRBD NFS, 2 LBs, 1 backoffice box (README cost section) |
| Price of the 3-DC spread itself | — | $0 in traffic: private-network and same-network-zone traffic is unbilled (Hetzner billing FAQ) |
| etcd member RTT | Sub-millisecond inside one DC | 25–27ms between Helsinki and the German DCs (hetzner-k3s measurements) vs etcd's 100ms heartbeat / 1000ms election-timeout defaults |
| Failure survived | Any one node | Any one node or one entire DC, etcd quorum included |
| Usable capacity | All of it, until the DC goes | Steady state must fit on 2 of 3 DCs — a DC loss takes 1/3 of workers with it |
| Ops model | Terraform + Ansible + ArgoCD, one failure domain | Same three tools, plus keepalived VIPs, DRBD replication, and Patroni failover stretched across DCs |
| Outage RTO for a DC loss | Hours (reprovision + restore) | Seconds to minutes (quorum holds, VIPs and replicas fail over) |
The short version: Hetzner charges you nothing extra for the spread, physics charges you ~25ms per etcd round trip, and the real bill arrives as operational complexity. The rest of this post shows its work.
What the multizone reference actually builds
The architecture is one private network (10.0.0.0/16) spanning Hetzner's eu-central network zone — fsn1, nbg1, hel1 — with every layer given a cross-DC failover story. Condensed from the project's own HA table:
| Layer | How the DC-level SPOF dies |
|---|---|
| Control plane | 3 k0s controllers + etcd quorum, one per DC, behind keepalived VIP 10.0.0.240 |
| Ingress / API LB | LB pair split across fsn1 + nbg1; keepalived moves the API VIP cross-DC; public ingress is round-robin A records across both LB public IPs |
| Workers | 3 workers, one per DC, plus the Hetzner cluster-autoscaler for burst |
| Storage (RWX) | DRBD NFS pair in fsn1 + nbg1 with a diskless quorum tiebreaker in hel1 |
| Database | 3-node Patroni Postgres, one per DC, automatic failover |
| GitOps | ArgoCD app-of-apps with self-heal (sealed-secrets, Traefik, monitoring, CI, and more) |
| Egress | NAT-HA on the LB pair |
| Admin access | WireGuard bastion on a backoffice box in hel1 (with monitoring and DB-backup tooling) |
Two design choices deserve notice before the cost discussion. First, public ingress failover is DNS round-robin — free and 2-DC reachable, but a browser retrying the live LB is not health-checked failover; the repo ships a Cloudflare Load Balancing template for services that need instant cutover. Second, the backoffice box (admin VPN, backup scheduler) lives in hel1 alone, so losing Helsinki costs you the VPN path and scheduled backups until it returns — the one layer that is node-redundant but not DC-redundant.
Cost #0: the traffic bill, verified rather than assumed
"Spreading these same nodes across three datacenters adds $0" is the claim that makes the whole topology plausible, so it is worth verifying against Hetzner's docs rather than the README alone. The billing FAQ states it plainly: "We only bill for outgoing traffic. Incoming and internal traffic is free." Internal traffic explicitly includes both private-network traffic (the Networks feature, which is what the reference uses) and traffic between servers in the same network zone over their public IPs. Since fsn1, nbg1, and hel1 all sit in eu-central, etcd replication, DRBD sync, and Patroni streaming between them never touch the meter.
Know the boundaries of that freebie, because they are where a careless variant of this topology starts paying. Traffic to a different network zone over public IPs counts as billable outgoing traffic — so a fourth node in Ashburn (us-east) talking to the fleet would meter. Public egress to the internet draws down each server's 20 TB EU allotment as usual. And Hetzner managed load balancers carry their own smaller traffic allotments, which is one reason the reference builds its LB pair from plain servers with keepalived and HAProxy instead. Confirm all of this on your own invoice or with the author's companion hetzner-cost-monitor before betting a production budget on it — but the docs are unambiguous about the core claim.
Cost #1: etcd at 25 milliseconds
Every Kubernetes write waits on etcd, and etcd quorum writes wait on the slowest member needed for majority. Inside one datacenter that round trip is sub-millisecond. Between Helsinki and the German DCs it is 25–27ms, measured and documented by the hetzner-k3s project, which supports masters in different locations and worked through exactly this math.
Against etcd's defaults, 25–27ms fits — but it is not a rounding error. The default heartbeat interval is 100ms and the default election timeout is 1000ms, and etcd's own tuning guidance says the election timeout should be at least 10x the RTT between members: 10x of ~27ms is ~270ms, comfortably under 1000ms. So the cluster is stable with defaults, and every write simply pays a ~25ms floor that a single-DC cluster never sees. That floor shows up in kubectl responsiveness, controller reconcile latency, and ArgoCD sync times — a constant background tax, not a spike.
The hetzner-k3s precedent also carries a warning worth heeding: its releases let operators pin all masters to just the German locations or just Helsinki because of the hel1↔DE latency. A 3-DC etcd quorum cannot take that option — one member must sit in Helsinki, and a network partition that isolates Helsinki leaves the two German members holding a bare 2-of-3 quorum with zero further margin. The topology survives a full DC loss, but during the survival it runs without a net.
Cost #2: idle headroom and location-pinned storage
Quorum math is only half the capacity story. The reference runs 3 workers, one per DC — so the DC outage you built this for instantly removes a third of your compute. The same applies to Patroni (a replica, or the primary itself, vanishes and must re-elect) and to any pod with a single replica that happened to schedule in the dead DC. Steady state must therefore fit on two DCs: run hotter than ~66% of total worker capacity and a DC loss becomes a DC loss plus an eviction storm.
That headroom is the permanent, invisible line item — roughly a third of the worker bill spent on capacity whose job is to sit idle until the worst day.
Storage has its own DC-shaped constraint. Hetzner volumes are created per location and attach only to servers in that same location — a volume in fsn1 cannot move to a server in hel1 — and each volume attaches to exactly one server at a time (volumes FAQ). That is precisely why the reference does not use plain block volumes for shared storage: cross-DC ReadWriteMany needs the DRBD-replicated NFS pair, with a diskless tiebreaker in the third DC so the two storage nodes cannot split-brain during a partition. It works, and it is also the single most old-school piece of the stack — kernel-module replication with fencing semantics, operated by you, at 2 AM, when the partition you designed for finally happens.
Cost #3: the complexity you now own
The infra delta is $0; the complexity delta is the actual price. Name it concretely:
Round-robin ingress. Without the Cloudflare LB upgrade, public failover is "the browser tries the other A record." API clients, webhooks, and anything without browser-style retry semantics will see errors until DNS-level luck or client retries route around the dead LB. The reference is honest about this — the Cloudflare template exists because round-robin is a budget answer, not a complete one.
Cross-DC keepalived. The k8s API VIP floats between DCs on a private-network alias IP. When it works, the API survives a DC loss seamlessly; when keepalived itself disagrees across a partitioned link, you are debugging VRRP elections instead of serving traffic. Single-DC keepalived has the same failure mode in theory, but partitions inside one DC's network are rarer than partitions between countries.
One-DC backoffice. The WireGuard bastion, monitoring stack, and DB-backup scheduler sit on a single box in hel1. Lose Helsinki and you lose your VPN path into the private network at the exact moment you most need it — plan an out-of-band route (Hetzner console, a second bastion, or at minimum documented break-glass steps) before you need one.
Three of everything to upgrade. k0s, Patroni, DRBD, keepalived, and the OS underneath each now roll across DCs with cross-DC replication lag in the loop. The Ansible playbooks automate the mechanics, but the blast radius of a bad upgrade is the whole fleet by construction — staging this stack means staging a second multizone fleet, not a single VM.
None of this is an argument that the reference is badly built. It is an argument that $103/mo buys the servers, while the HA itself is paid for in on-call competence with four distributed systems instead of one.
The honest alternative: single-DC plus backup-and-reprovision
For many small teams the right comparison is not "3-DC versus nothing" but "3-DC versus single-DC with a practiced reprovision runbook": scheduled etcd snapshots and Patroni base backups shipped to object storage, Terraform state intact, and a documented re-apply into a surviving DC. Recovery time is typically measured in hours — provision, restore, repoint DNS — versus seconds to minutes for quorum-held failover. If your business can tolerate a half-day outage once every few years, the reprovision path buys back all of Cost #3 for the price of drill discipline.
This is also where Cluster API changes the shape of the trade. CAPH (the Hetzner provider) uses controlPlaneRegions as the base for the cluster's failure domains — defaulting to a single location, fsn1 (CAPH cluster reference) — which is why most CAPH fleets are single-DC by default. But the failure-domain machinery means a CAPH fleet can declare a multizone spread and, just as importantly, can reprovision an entire replacement fleet from manifests rather than re-running imperative playbooks. Declarative machine lifecycle does not remove the etcd-latency or headroom costs, but it makes the backup-and-reprovision alternative dramatically more credible: the runbook is kubectl apply, not a weekend. (Syself's own HA guidance still says to keep control-plane nodes in the same region and off the cheapest CX shapes — multizone ambition does not excuse undersized controllers.)
One more consideration sharpens the verdict: Hetzner publishes no formal uptime SLA for cloud products — no guaranteed percentage, no credit-backed compensation (Better Stack's 2026 Hetzner review states this explicitly, in contrast to DigitalOcean's 99.99%). You cannot contract your way to uptime on Hetzner; you can only architect it. That cuts both ways: it is the strongest argument for building the multizone topology, and the strongest argument for admitting that the topology's complexity is load-bearing rather than optional.
Verdict: a checklist, not a slogan
Go multizone when most of these are true: a half-day outage costs more than a week of an engineer's time to build and drill this stack; someone on the team has operated (or will learn) etcd, Patroni, DRBD, and keepalived under failure, not just from READMEs; your workloads tolerate ~25ms of extra control-plane latency on every write; and you will actually run the fleet at ≤66% worker utilization instead of letting the headroom get "temporarily" borrowed. Stay single-DC with practiced backups when the outage math does not clear that bar — and if you stay, make the reprovision path declarative (CAPH, or Terraform state you trust) so the runbook survives the engineer who wrote it.
Either way, steal the reference's cheapest insight for free: your failover story is only as real as your last drill. A 3-DC quorum you have never partitioned on purpose and a backup you have never restored are the same thing — a hope with documentation.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



