Skip to main content

Sonobuoy Conformance as a Fleet Gate: Don't Schedule Tenants on an Unproven Node

9 min readDora NodaDora Noda
Share
On this page

Cilium's own documentation admits something most fleet operators learn the hard way: when a new node joins, application pods can be scheduled onto it before the Cilium agent is ready to manage them, leaving those pods with broken networking until someone restarts them. Your Cluster API Machine says Running, your Hetzner server says active, kubelet registered, kubectl get nodes shows Ready — and the first tenant deploy to land there gets a pod that can't talk to anything.

The fix is a gate, not a hope. Every new machine joins the cluster carrying a NoSchedule taint, proves itself against a focused slice of the CNCF conformance suite, and only then gets untainted for tenant workloads: Machine ready → conformance pass → tenant pods. Fail the gate and the machine gets remediated or deleted before a tenant ever notices it existed. The rest of this post is the concrete wiring for that gate on a CAPH-managed Hetzner fleet — what Hetzner's green checkmark doesn't cover, why the full conformance suite can't be the per-machine gate (and what to use instead), and the exact taint, wait, run, and untaint steps for both tiers.

What Hetzner's green checkmark doesn't cover​

CAPH provisioning success means the server exists, cloud-init ran, and kubeadm joined the cluster. None of that exercises the data paths a tenant pod actually needs. Here is the gap, failure by failure:

Failure modeWhat the tenant seesWhat catches it
CNI agent not Ready on the new nodePod networking broken; DNS and east-west traffic fail until pod restartDaemonSet-Ready wait on the node before untaint
hcloud-CSI node plugin failingPVCs stuck Pending or ContainerCreating forever[sig-storage] basics in the focused gate run
CCM never initialized the node (wrong providerID/metadata)Node stuck with node.cloudprovider.kubernetes.io/uninitialized taint; LoadBalancer targets never registerGate pre-check: uninitialized taint gone + providerID set
Kubelet flag or cgroup mismatch vs fleet standardCPU/memory accounting drift; QoS evictions behave differently than on sibling nodessystemd-logs plugin + kubelet config diff in the gate
Kernel module or sysctl drift in the machine imageConformance networking/storage tests fail in ways that look like app bugsFocused [sig-network] + [sig-node] tests
Node can't reach CoreDNS or the API server reliablyIntermittent resolution failures, the classic "works on every node but this one"[sig-network] DNS focus tests scheduled while gated

Every row in that table shares one property: the machine looked healthy to every layer below the workload. Hetzner provisions servers, not working Kubernetes nodes. kubelet Ready means the kubelet is heartbeating, not that a pod placed there can resolve DNS, mount a volume, and reach the cluster network. The gate exists to check the workload's view, from the workload's side of the API.

Two tiers, because the full suite can't gate one machine​

Here is the honest constraint that shapes the whole design: a full certified-conformance run is roughly 460 tests on recent Kubernetes versions — 462 official conformance tests for v1.37, 441 for a v1.35 Kubespray submission — takes an hour or more on a healthy cluster, runs cluster-wide rather than node-pinned, and includes tests Sonobuoy's own docs warn may disrupt other workloads. Sonobuoy's default --timeout is 21,600 seconds (six hours) precisely because conformance runs long. Gating every autoscaler scale-up on that would turn a two-minute node addition into a multi-hour ceremony and blast production tenants with disruptive tests in the process.

So the gate has two tiers with different triggers:

Tier 1: per-machine gateTier 2: template-promotion gate
TriggerEvery new Machine (scale-up, rollout, replacement)New machine image, new Kubernetes minor, new CNI/CSI/CCM version
WhereProduction cluster, node taintedCanary cluster built from the candidate template
Sonobuoy modequick smoke + focused --e2e-focus subset--mode=certified-conformance
RuntimeMinutesOne to several hours
VerdictUntaint the node, or delete/remediate the MachinePromote the template to the fleet, or don't

Tier 1 answers "is this machine a working Kubernetes node." Tier 2 answers "is this template a conformant cluster." Tier 1 runs constantly and must be fast; Tier 2 runs rarely and must be thorough. Confusing the two — running the full suite per machine, or running only a smoke test per template — is how fleets end up with either glacial scaling or untested images.

Wiring Tier 1: the per-machine gate​

Four steps, all automatable from the management cluster: taint at join, wait for node-local readiness, run the focused gate, untaint or remediate.

1. Taint at join. The taint must exist from the moment kubelet registers, not applied afterwards in a race with the scheduler. The standard mechanism is kubeadm's nodeRegistration.taints in the JoinConfiguration rendered by your KubeadmConfigTemplate:

yaml
apiVersion: bootstrap.cluster.x-k8s.io/v1beta1
kind: KubeadmConfigTemplate
metadata:
  name: workers-conformance-gated
spec:
  template:
    spec:
      joinConfiguration:
        nodeRegistration:
          taints:
            - key: fleet.bex.co/conformance-gate
              value: "pending"
              effect: NoSchedule

Nothing without an explicit toleration schedules onto the node now — including the conformance test pods, which matters in step 3. Keep the value machine-readable (pending, then remove the taint entirely on pass) so a stuck gate is visible as pending rather than ambiguous.

2. Wait for node-local readiness. Before spending a single conformance test, check the cheap signals: the node.cloudprovider.kubernetes.io/uninitialized taint is gone (CCM claimed the node), the Cilium agent pod on the new node is Ready, the hcloud-CSI node plugin pod is Ready, and kubelet's cgroup driver and version match the fleet standard. This pre-check catches the majority of bad nodes in seconds. Only nodes that pass it earn a Sonobuoy run.

3. Run the focused gate. Start with the smoke test, then the workload-path subset. Sonobuoy's quick mode runs a single simple, fast test — a reachability check that the cluster responds. Then run a focused, non-disruptive subset aimed at the paths tenants need:

bash
# Smoke: is the cluster (and the new node in it) responsive?
sonobuoy run --mode=quick --wait --timeout 600
sonobuoy results $(sonobuoy retrieve)
 
# Gate: DNS, basic pod networking, and storage provisioning —
# non-disruptive so production tenants are unaffected.
sonobuoy delete --wait
sonobuoy run \
  --e2e-focus='\[sig-network\].*DNS|\[sig-storage\].*provisioning' \
  --wait --timeout 3600
sonobuoy results $(sonobuoy retrieve)
sonobuoy delete --wait

Two notes on that second command. First, Sonobuoy's modes are shorthand for focus/skip values, so the docs say not to combine --mode with --e2e-focus — pass the focus regex alone and the default skip of [Disruptive] tests still applies, which keeps the gate safe to run against production. You pay for minutes of targeted signal instead of hours of full-suite coverage. Second, use sonobuoy e2e --focus <regex> --mode=tagCounts first to preview exactly which tests and tags your regex selects before burning cluster time; it lists the matched tests without launching a single pod.

The scheduling subtlety: your gate taint keeps tenant pods off the new node, but conformance test pods schedule wherever the scheduler puts them. For the DNS and networking tests to exercise the new node, scope the run while the rest of the fleet is otherwise occupied is unreliable — the practical answer is that Tier 1 validates the cluster with this node in it plus the node-local checks from step 2, and node-pinning purists should run the focused subset against a single-node canary (see Tier 2's cluster, reused). Don't overclaim: Tier 1 proves the node is correctly assembled and the cluster still conforms with it present, not that every test pod touched the new kubelet.

4. Untaint or remediate. Parse the result: sonobuoy results exits nonzero on any failure. On pass, remove the taint and the node joins normal scheduling:

bash
kubectl taint nodes "$NODE" fleet.bex.co/conformance-gate-

On fail, the Machine is guilty until proven innocent: cordon it, collect sonobuoy retrieve plus the systemd-logs plugin output for the postmortem, then delete the Machine and let the MachineDeployment replace it. A gate that pages a human at 3am for every flaky e2e test will be disabled within a month; a gate that deletes one bad machine and retries with a bounded budget gets to keep running. Set that budget explicitly — two replacement attempts, then hold the rollout and alert.

Wiring Tier 2: the template gate​

Tier 2 fires when the thing being trusted changes: a rebuilt machine image, a Kubernetes minor bump, a Cilium, hcloud-CSI, or CCM version change. CAPH makes the canary cluster cheap — it's the same clusterctl template with a different name — so stand one up, run the real thing, and promote the template only on a clean pass:

bash
sonobuoy run --mode=certified-conformance --wait --timeout 21600
outfile=$(sonobuoy retrieve)
mv "$outfile" "canary-$TEMPLATE_VERSION.tar.gz"
sonobuoy results "canary-$TEMPLATE_VERSION.tar.gz"
sonobuoy delete --wait

certified-conformance is the mode the CNCF certification program requires: every Conformance-tagged test, including the disruptive ones, which is exactly why it runs on a throwaway canary and never on production. Keep the retrieved tarball per template version; it is your audit trail when someone asks what "tested" meant for the image now running the fleet.

If Sonobuoy's aggregator-plus-plugins architecture feels heavy for a canary job, Hydrophone is the SIG-Testing-maintained lightweight alternative: a single Go binary that starts the official conformance image as a pod, waits, and prints results. The equivalent Tier 2 run is hydrophone --conformance, a single test is hydrophone --focus '...', and the image pins explicitly with --conformance-image 'registry.k8s.io/conformance:v1.37.1'. Same tests, same verdict semantics, less machinery — a reasonable choice when the canary job just needs a pass/fail and a log. Either runner's output also doubles as the artifact for a CNCF conformance certification submission if you ever want the fleet's distribution on the certified list.

What the gate still won't catch​

A gate earns trust by stating its limits. Conformance tests the Kubernetes API's contract, not your fleet's fitness, so several failure classes sail through both tiers. Conformance is cluster-scoped, not node-pinned: Tier 1 can prove the node is correctly assembled, but pinning every test pod onto the candidate machine is not what the suite does. It says nothing about load: a node that passes every functional test can still be a noisy-neighbor disaster under real tenant pressure, which is a PSI-and-benchmark question, not a conformance question. Hetzner-side state — firewall rules, load-balancer targets, private-network routes — lives outside the Kubernetes API the suite exercises, so the CCM-shaped holes need their own checks. And e2e tests flake: version-pin the conformance image to your exact Kubernetes minor, keep a short list of known-flaky tests with focused re-runs rather than blanket --e2e-skip, and treat repeated identical failures across machines as a template bug, not a machine bug.

None of that diminishes the gate. A fleet that deletes machines failing DNS provisioning before tenants schedule is strictly more reliable than one relying on Ready. The gate just isn't a proof — it's a filter with a known mesh size, and the mesh is sized in Tier 1 for speed and in Tier 2 for depth.

The pattern generalizes beyond Hetzner: any Cluster API provider whose "machine provisioned" signal stops at the infrastructure layer leaves the same gap between infrastructure-ready and workload-ready. Taint at join, verify the workload's view, untaint on pass. Your tenants should never be the first thing to discover a node can't run pods.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex