Skip to main content

Metal3 and Ironic for True Bare Metal: What Cluster API Looks Like Without a Cloud API Underneath

9 min readDora NodaDora Noda
Share
On this page

Bare metal is having a moment. Metal3 — the Kubernetes-native bare-metal provisioning project — entered 2026 as one of the CNCF's freshly incubated projects, filed its graduation self-governance review in April, and earlier this month got a CNCF blog walkthrough showing it provisioning KubeVirt virtual machines through virtual Redfish BMCs as if they were physical servers. The message is clear: declarative, Kubernetes-API-driven machine lifecycle is no longer just for cloud APIs. But what does running Cluster API actually change when the machines underneath are servers in a rack — with BMCs, IPMI addresses, and no cloud API to call — instead of Hetzner Cloud VMs or Robot dedicated servers?

This post is the concrete CAPM3-vs-CAPH walkthrough: what changes in image delivery, credential management, and load balancing when your Cluster API provider talks Redfish instead of REST, and why a self-hosted PaaS that already speaks Cluster API should treat Metal3 as the answer to "what if a customer wants the fleet on hardware Hetzner never sees" rather than a replacement for CAPH.

The three moving parts​

Metal3 is three components working as one provisioning pipeline. First, the Bare Metal Operator (BMO), a Kubernetes controller that manages physical servers as BareMetalHost custom resources — each CR is one server, with its BMC address, boot MAC, desired OS image, and a reference to its BMC credentials. Second, Ironic, the OpenStack bare-metal provisioning service, which does the actual work of talking to server management controllers and writing images to disk. Third, Cluster API Provider Metal3 (CAPM3), the CAPI infrastructure provider that lets a MachineDeployment claim BareMetalHost objects the way a cloud provider's Machine claims VMs.

A minimal host registration shows the whole contract:

yaml
apiVersion: v1
kind: Secret
metadata:
  name: rack-u14-bmc
type: Opaque
data:
  username: YWRtaW4=
  password: cGFzc3dvcmQ=

and the host that references it:

yaml
apiVersion: metal3.io/v1alpha1
kind: BareMetalHost
metadata:
  name: rack-u14
spec:
  online: true
  bootMACAddress: 00:c2:fc:3b:e1:01
  bootMode: UEFI
  bmc:
    address: redfish://192.168.100.14/redfish/v1/Systems/1
    credentialsName: rack-u14-bmc
  image:
    url: https://images.internal/ubuntu-24.04-k8s.raw
    checksumType: sha256
    checksum: 9f2c…

That is the entire bottom of the stack: a BMC address, a Secret with its login, a boot NIC, and an image URL. No cloud account, no API token, no region. Everything downstream — inspection, provisioning, deprovisioning — flows from Ironic driving that BMC.

CAPM3 vs CAPH: the same Cluster API, a different bottom​

Here is the core comparison. Above the provider line, nothing changes: same CAPI Cluster, MachineDeployment, and KubeadmControlPlane objects, same GitOps flow, same clusterctl move. Below it, nearly every assumption flips:

ConcernCAPM3 (Metal3 + Ironic)CAPH (Hetzner)
Machine creationClaim from pre-registered inventory: a Metal3Machine binds an available BareMetalHost; scaling out past inventory blocks until someone racks and registers another serverAPI creates capacity: HCloudMachine calls the Cloud API to create a VM; HetznerBareMetalMachine claims a Robot server you already own
Image deliveryIronic boots the Ironic Python Agent (IPA) ramdisk on the target via PXE or virtual media, then IPA writes your image to the local disk and reports backCloud VMs boot a Hetzner snapshot image picked by API parameter; Robot dedicated servers install via rescue-mode provisioning driven through the Robot API
CredentialsOne Kubernetes Secret per host holding that server's BMC username/password, plus fencing for the BMC management networkOne Hetzner API token (Cloud) plus Robot credentials; Hetzner exposes no customer-facing BMC at all
Load balancer / floating IPNothing built in: bring MetalLB (L2 or BGP) or kube-vip plus routable IPs you own; a bare LoadBalancer Service sits in pending otherwiseHCloud Load Balancers and Floating IPs are first-class API objects CAPH-adjacent tooling can allocate
Typical failure modeDead BMC credential, unreachable Redfish endpoint, bad boot MAC, image server unreachable from the provisioning networkAPI error, quota/rate limit, wrong server type or image name, Robot-vs-Cloud mismatch

Two rows deserve emphasis because they surprise cloud-trained operators most. First, Metal3 never creates a server — there is no API that racks hardware. Your BareMetalHost inventory is the capacity ceiling, and "autoscaling" means pre-provisioned spares marked available, not on-demand servers. Second, the BMC is a second network you now own: Ironic must reach every management controller, which means a provisioning/management VLAN, firewall rules, and credential hygiene for dozens of tiny embedded web servers that rarely get patched.

A host's life, step by step​

Watching one host move through its lifecycle makes the model concrete. After you apply the BareMetalHost, BMO hands it to Ironic, which verifies the BMC credentials — a Redfish GET against the management controller — and the host sits in registering. Then comes inspection: Ironic boots the IPA ramdisk on the machine (over PXE or attached virtual media) and the agent phones home with the real hardware inventory — CPU model and count, RAM, disk sizes, NICs. What you declared in YAML meets what is actually screwed into the rack, and mismatches surface here, before anything is installed.

Once inspection succeeds the host becomes available: imaged with nothing, powered to policy, waiting in the pool. Nothing else happens until a CAPM3 Metal3Machine claims it — typically because a CAPI MachineDeployment scaled up. On claim, Ironic boots IPA again, this time to provision: the agent streams your OS image from the URL in the host spec, writes it to disk, verifies the checksum, and reboots into the installed system. The host reports provisioned, kubeadm (or your bootstrap provider) joins it to the workload cluster, and from CAPI's perspective it is just another provisioned Machine.

Two details matter in practice. Deprovisioning reverses the flow — the disk is cleaned and the host returns to available, ready for the next claimant — which is what makes the pool reusable rather than single-use. And the externallyProvisioned flag is the escape hatch: set it and Metal3 manages power and inventory while something else owns the OS install, useful for brownfield racks with an existing image pipeline you are not ready to migrate.

The three bills that come due​

The comparison table is honest only if the CAPM3 column's hidden costs are spelled out. There are three, and each is a system you operate instead of an API you call.

Bill one: the image pipeline is yours now. With CAPH, the image catalog lives at Hetzner: pick a snapshot, pass its name, done. With Metal3, you build the OS images, host them on an HTTP server reachable from the provisioning network, and pin URL plus SHA-256 checksum in every host spec. That includes the IPA ramdisk itself for inspection and deploy. Image builds, signing, retention, and mirror availability become your runbooks — the same work a cloud provider's image team does invisibly.

Bill two: per-host BMC secret lifecycle. One Hetzner token rotates in one place. A Metal3 fleet has one credential Secret per server, each pointing at an embedded controller with its own firmware, default-password history, and network exposure. You need a rotation story per host (or per rack, if your BMCs share credentials — convenient until one leaks), the BMC LAN fenced off from tenant and provisioning traffic, and monitoring for controllers that stop answering Redfish. The September KubeVirtBMC post is a fun proof that "a BMC" can even be virtual — but every virtual BMC still needs its credential managed.

Bill three: DIY ingress and load balancing. On Hetzner Cloud, a LoadBalancer Service gets a real IP from an API call. On rack hardware there is no such API, so you run MetalLB in L2 or BGP mode (or kube-vip for control-plane and service VIPs) and supply the routable address pool yourself — negotiated with whoever owns the rack's uplink. L2 mode is simple and constrains failover to ARP timing; BGP mode is robust and requires a speaking-terms relationship with the top-of-rack router. Either way, the load-balancer control plane is now a component in your cluster, with its own upgrades and failure modes, not a line item on a cloud bill.

None of these bills is a reason to avoid Metal3. Each is the price of hardware no cloud API will ever describe — and each has a mature answer (image servers, external-secrets rotation, MetalLB). But budget them before the first rack, not after.

Addition, not replacement​

So where does this leave a self-hosted PaaS that already runs CAPH against Hetzner? Exactly where the TODO-frame puts it: Metal3 is the answer to "what if a customer wants the fleet on hardware Hetzner never sees," not a CAPH replacement.

The decision rule is physical. If the capacity comes from Hetzner's APIs — Cloud VMs for elastic pools, Robot dedicated servers for heavy nodes — CAPH is the thinner, better-fitting provider: it speaks the vendor's native provisioning (including rescue-mode installs for dedicated servers) and inherits the vendor's load balancers, floating IPs, and image catalog. Nothing about Metal3's graduation changes that math.

Metal3 earns its place the day capacity stops being API-shaped: a customer's own rack in a colo with Redfish-capable BMCs, an air-gapped site where the only "API" is the management LAN, GPU servers from a vendor Hetzner doesn't stock, or edge cabinets where each box was installed by hand. In all of those, CAPH has nothing to call — there is no Hetzner API in the path — while a BareMetalHost per server plus CAPM3 keeps the fleet behind the same CAPI control plane, the same GitOps repo, and the same upgrade runbooks as the Hetzner capacity.

That shared top half is the real prize. The end state is not "Metal3 or CAPH" — it is one management cluster reconciling both: cloud-shaped capacity through CAPH, rack-shaped capacity through CAPM3, with tenants seeing only Machines that become Nodes. Metal3's march toward graduation (Sandbox in 2020, Incubating in 2025, graduation review filed April 2026) is the signal that the rack-shaped half is ready to be a supported leg of that fleet, not an experiment.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex