A typo in your HCloudMachineTemplate used to die at kubectl apply time. Now it dies inside Hetzner's API — after Cluster API has already started reconciling the Machine. That is the quiet consequence of two changes landing on opposite sides of the same seam: Cluster API Provider Hetzner (CAPH) stopped validating hcloudMachine.spec.type in v1.0.7, and Hetzner moved server-type availability from datacenters to per-location signals, with hcloud server-type describe now showing Available and Recommended for every location. The provider dropped its hardcoded list; the cloud gave you a live query to replace it. If your fleet templates still hardcode a type string and assume it exists everywhere, you are flying on the old assumption with the new behavior.
The short version, up front, for the impatient:
hcloud server-type describe cx23
hcloud server-type list -o columns=name,location_available,location_recommendedQuery per-location availability, pin a generation-qualified type per location in each template, and gate applies on that check. The rest of this post is why that practice exists and what breaks if you skip it.
What actually changed, in order
Four events over twelve months, each sensible alone, collectively moved a validation boundary:
| Date | Event | What it means for fleet templates |
|---|---|---|
| Oct 16, 2025 | Hetzner deprecates old server types (cx22, cpx11, …), introduces new types with categories; old types removed Jan 1, 2026 | Any template pinning a first-generation cx/cpx name has a hard expiry date |
| Oct 2025 | CAPH v1.0.7 removes hcloudMachine.spec.type validation (PR #1694) | spec.type becomes a free-form string; typos and retired names pass admission and fail at Hetzner's API |
| Apr 1, 2026 | Availability moves datacenter → location; hcloud CLI v1.63.0 adds per-location Available/Recommended to server-type describe and location_available/location_recommended columns to server-type list | The replacement signal exists — but you have to query it; nothing queries it for you |
| Jul 1 / Oct 1, 2026 | Datacenters removed from the Hetzner API (July); GET /v1/datacenters returns 410 Gone and the hcloud datacenter describe server-type list disappears (October) | Old datacenter-keyed capacity checks and manifests stop working entirely |
The CAPH decision is worth understanding on its own terms, because it looks like a regression and is not one. Hetzner changes the valid type list from time to time — deprecating generations, introducing new ones — and every such change put CAPH's hardcoded enum out of date. The provider was rejecting types Hetzner had just started selling, or accepting types Hetzner had just retired. Dropping the enum was the only way to stop lying. The release note says it plainly: it is up to the user to use a valid type, as not all types are available in all locations.
That last clause is doing the heavy lifting. Validity is now a two-variable function — type × location — and no static list in a CRD can express it. Hence Hetzner's side of the seam: availability as a live, per-location query.
Where the failure moved
Before CAPH v1.0.7, an invalid type failed fast and loud:
# Before: admission-time rejection
kubectl apply -f machinetemplate.yaml
# Error: spec.type: Unsupported value: "cx22x": supported values: "cx22", "cpx11", ...The feedback loop was seconds long and pointed at the exact field. After v1.0.7, the same manifest applies cleanly. The Machine object is created, the controller picks it up, calls Hetzner's server-create API, and only then learns the type does not exist — or exists but not in that location. The symptom is a Machine stuck in provisioning with a provider-side error (ServerCreateFailed, "unsupported location"), which you clear by patching the template and deleting the stuck Machine so the controller reconciles fresh. Operators in the wild have already hit exactly this shape: type valid in one location, rejected in another, discovered only at create time.
The cost difference is not just latency of feedback. An admission error blocks a gitops sync with a clear message before anything is created. An API-time failure leaves half-reconciled state — a Machine with no backing server, possibly a consumed name or IP reservation depending on ordering — and pages whoever watches the fleet, not whoever wrote the manifest. Multiply by every template in every location your fleet spans, and the dropped enum becomes a dropped guardrail on your busiest path: node provisioning.
There is a second, slower failure hiding behind the first. Templates that pinned cx22 or cpx11 kept working right up until January 1, 2026, when Hetzner removed the deprecated types. Downstream projects felt this as CI breakage: Flatcar moved its Hetzner test hosts off cx22, Azure Container Linux and Flatcar's Mantle moved off cpx11, all citing the same October deprecation notice. If your templates were written before the new generations landed and nobody re-audited them against the removal date, "it applied fine last quarter" is not evidence it provisions today.
The replacement practice: poll and pin
The good news is that Hetzner replaced a static fact (the type list) with a queryable one (per-location availability), and the query surface is good. hcloud server-type describe now carries an Available and Recommended value per location, and hcloud server-type list exposes the same pair as location_available and location_recommended columns. For automation, both commands emit JSON:
# Human-readable: is cx23 usable in every location I run?
hcloud server-type list -o columns=name,location_available,location_recommended
# Machine-readable: feed it to a gate
hcloud server-type list -o json | jq '.[] | {name, locations}'The Terraform provider made the same move for the IaC path: hcloud_server_type data sources gained locations[].available and locations[].recommended, replacing the deprecated datacenter-keyed attributes. Whichever tool owns your templates, the signal is one call away.
Poll means: before an apply — in CI, in a pre-sync hook, in a nightly fleet audit — resolve every (type, location) pair your templates reference against live availability. The check is cheap (one CLI call lists all types), and it catches all three failure shapes at once: the typo (no such type anywhere), the retired type (no longer available anywhere), and the location mismatch (available in fsn1, not in ash). That third shape is the one no static list could ever have caught, and it is the common one on a multi-location fleet.
Pin means: write the answer into the template explicitly, per location, using generation-qualified names. Concretely:
- One
HCloudMachineTemplate(or one ClusterClass variable set) per location you provision into, each pinning a type you have verified available in that location — never one template with a type you verified in exactly one place. - Prefer current-generation names (the post-October-2025 families) and treat any surviving
cx11/cx22/cpx11-era string as tech debt with a known removal date. - Remove leftover
datacenterreferences while you are there: the field is gone from server create flags and describe output, and anything still keying capacity logic on datacenter names breaks against the October 2026 API.
A minimal CI gate looks like this — fail the sync when a template's pair is not reported available:
#!/usr/bin/env bash
# fleet-type-gate.sh — run before applying CAPH templates
set -euo pipefail
TYPE="$1" # e.g. cx23
LOCATION="$2" # e.g. fsn1
available=$(hcloud server-type describe "$TYPE" -o json \
| jq -r --arg loc "$LOCATION" '.locations[] | select(.name == $loc) | .available')
if [ "$available" != "true" ]; then
echo "FAIL: server type $TYPE is not available in $LOCATION" >&2
exit 1
fi
echo "OK: $TYPE available in $LOCATION"This is deliberately boring. It re-implements, as a ten-line script you own, the check the provider used to own — except yours queries live data, so it cannot go stale the way the hardcoded enum did. That is the whole point: the validation did not disappear, it changed owners, from CAPH's release cycle to your pipeline.
Gotchas that still bite
Three subtleties survive even after you adopt poll-and-pin.
Available and Recommended are not synonyms. Available answers "can I create this type here right now" — a capacity fact. Recommended is Hetzner's guidance about which types to prefer in that location, which tracks generations, stock depth, and what Hetzner would rather sell you. Gate provisioning on Available; use Recommended when choosing which type to pin during a refresh. Pinning a type that is available-but-not-recommended is not an error, but it is a signal to check whether you are on an older generation with a deprecation notice already written.
Availability is a point-in-time reading, not a reservation. A type available when CI ran at 10:00 can be out of stock when the autoscaler needs it at 10:05 — cheap shared types in popular locations do sell out, and workshop organizers have learned to budget a fallback type for exactly this reason. For control-plane templates, where a failed provision blocks the whole cluster, pin a type with deep stock (usually the current generation's mainstream size) rather than the cheapest one that passed the gate. For worker pools, consider a second template with an alternate type so a stockout degrades to a different shape, not a stalled scale-up.
The location/value split cuts both ways. The same flexibility that lets Hetzner sell different types per location means your fleet's "standard worker shape" may legitimately differ between fsn1 and ash. Resist the urge to force one global type for tidiness: a template that provisions everywhere by pinning the lowest common denominator leaves money or performance on the table in locations where better types are stocked. Per-location pins are not duplication; they are the honest representation of a per-location fact.
Owning the check you used to get for free
Step back and the pattern is a familiar one in platform engineering: a dependency stops guaranteeing something, and the guarantee does not vanish — it becomes your job. CAPH could not keep a static type list truthful against Hetzner's catalog churn, so it handed the job to operators; Hetzner, to its credit, handed operators a live API to do the job with. Fleets that update their practice — poll availability, pin per location, gate the apply — end up with a strictly better check than the old enum, one that understands locations and stock. Fleets that do nothing keep manifests that apply cleanly and provision never, and learn about it from a stuck Machine at the worst hour.
If you run Cluster API on Hetzner, this week's audit is small: list every spec.type in your templates, run each pair through hcloud server-type describe, and fix the ones that fail before the API does it for you.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



