Your fleet's unattended upgrades ran on schedule, the kernel package is new on disk — and every node is still running the old kernel, because nothing ever rebooted. That gap between "patched on disk" and "patched in memory" is the quietest lie in fleet operations: the package manager reports success while the running kernel stays vulnerable to exactly the CVE you thought you fixed. Kured, the Kubernetes Reboot Daemon now maintained by the kubereboot community as a CNCF project, exists to close that gap: a DaemonSet that watches for the OS-upgrade sentinel file and cordons, drains, and reboots each node in turn.
Here is the verdict before the why: run kured on any Cluster API fleet with unattended host patching — but only after four guardrails are in place. A reboot daemon without MachineHealthCheck coordination will get its nodes replaced mid-reboot; without tenant PDBs it drains into an unprotected floor; with the default infinite drain timeout one stuck node halts the whole fleet's patch cycle; and without alert-gated windows it reboots through incidents. This post gives you the concrete before/after: kured's loop, the manual runbook steps it deletes, and the exact guardrail configuration that makes unattended reboots safe on multi-tenant nodes.
| Step | What kured does | Manual step it replaces |
|---|---|---|
| 1. Sense | DaemonSet pod watches /var/run/reboot-required (or a sentinel command) every --period (default 1h) | SSH/orchestrator sweep checking which nodes need reboots |
| 2. Serialize | Takes an API-server lock (weave.works/kured-node-lock annotation); one node at a time by default | Human pacing ("reboot node 3 only after node 2 is back") |
| 3. Evacuate | Cordons and drains the node | kubectl cordon + kubectl drain per node, babysat |
| 4. Reboot | Runs the reboot command (default /bin/systemctl reboot) | SSH in and sudo reboot, then poll for rejoin |
| 5. Restore | Uncordons the node, releases the lock, notifies | kubectl uncordon, update the checklist, next node |
The rest of this post shows the mechanism in one section, the runbook before/after, the four guardrails with a concrete flag set, and the honest reboot-versus-rotate decision for CAPI fleets.
How kured works, in one section
Kured (KUbernetes REboot Daemon) is a DaemonSet — one pod per node — that performs safe automatic node reboots when the underlying OS's package management signals one is needed. The full configuration reference lives at kured.dev; the mechanism fits in five facts.
First, sensing. Each pod watches for a reboot sentinel: by default the file /var/run/reboot-required, which Debian/Ubuntu's unattended-upgrades and manual apt upgrade runs write when an installed update (typically a kernel or core library) only takes effect after a reboot. If your distro or patch pipeline does not write that file, --reboot-sentinel-command accepts any command whose zero exit code means "reboot needed" — the hook that makes kured work with custom image pipelines or updaters that do not write the Debian sentinel path.
Second, serialization. Before touching a node, kured takes a lock recorded as an annotation on its own DaemonSet object. The default allows exactly one node rebooting at a time; --concurrency raises that for large fleets willing to trade speed against reduced headroom. Lock hygiene flags — --lock-ttl, --lock-release-delay — bound how long a crashed or wedged reboot can hold the fleet's patch cycle hostage.
Third, evacuation. Kured cordons the node and drains it like kubectl drain would, honoring PodDisruptionBudgets. Drain behavior is tunable (--drain-grace-period, --drain-delay, --drain-pod-selector, --skip-wait-for-delete-timeout) — and one default deserves a warning label, covered in the guardrails section.
Fourth, reboot and restore. Kured runs the reboot command (default /bin/systemctl reboot; a --reboot-method signal alternative exists), waits for the node to come back, uncordons it, releases the lock, and moves on. --pre-reboot-node-labels and --post-reboot-node-labels let external automation key off the transition, and --annotate-nodes stamps weave.works/kured-reboot-in-progress on the node for dashboards.
Fifth, observability. Each pod exposes a kured_reboot_required gauge on :8080/metrics, designed to power the meta-alert every fleet needs: fire when a node has needed a reboot for longer than the patch cycle tolerates, i.e. "the cluster cannot reboot itself." Notifications go out over --notify-url (Slack, webhooks) at drain, reboot, and successful uncordon.
Two operational details round out the picture. Testing is a one-liner — sudo touch /var/run/reboot-required on a canary node provokes a real reboot cycle — and the emergency stop is taking the lock manually by annotating the DaemonSet, which blocks all further reboots without touching the workload.
The runbook before and after
A manual OS-patch reboot runbook for a fleet of N nodes reads like this: sweep the fleet for the sentinel file; pick node 1; cordon it; drain it and watch the evictions; SSH in and reboot; poll until the kubelet rejoins and the node is Ready; uncordon; verify tenant pods rescheduled; repeat N times, preferably inside a maintenance window, preferably while nobody else is deploying. Every step is a context switch, every node is 10–20 minutes of supervised waiting, and the failure mode is human: the engineer reboots the node holding the only replica of something, or skips the last three nodes because the window closed.
Kured deletes the per-node supervision, not the responsibility. After adoption the runbook is: unattended upgrades install patches and write sentinels on their own schedule; kured notices within one --period, takes the lock, and walks the fleet one node at a time inside the configured reboot window; the on-call engineer watches notifications and the kured_reboot_required alert instead of terminal windows. What used to be N supervised cycles becomes one configured policy plus exception handling.
Be precise about what disappears versus what moves. Gone: per-node cordon/drain/SSH-reboot/poll/uncordon, manual pacing between nodes, the spreadsheet of which nodes are done. Moved into configuration: pacing becomes --concurrency and reboot windows (--reboot-days, --start-time, --end-time, --time-zone); "don't reboot during the incident" becomes the Prometheus alert gate and blocking-pod selectors; "don't strand capacity" becomes the pending-reboot taint. Still yours: the guardrails below. A reboot daemon automates the hands; the judgment about when a reboot is safe has to be encoded, or it has to stay human.
The four guardrails before you trust it on multi-tenant nodes
1. MachineHealthCheck interplay: the timeout inequality. This is the guardrail that bites first. During a kured reboot the node goes NotReady, and a MachineHealthCheck watching Ready=False with a short timeout will declare the machine unhealthy and replace it — deleting a node that was seconds away from rejoining, churning a Hetzner server (or cloud VM) for no reason, and evicting everything kured carefully drained back onto survivors. OpenShift's upgrade documentation states the general rule outright: pause MachineHealthCheck resources before operations that make nodes temporarily unavailable, because MHC cannot distinguish intentional downtime from failure.
Pausing MHC fleet-wide on every patch cycle is operationally heavy, so the practical form is a timeout inequality your MHC must satisfy:
MHC unhealthy timeout > worst-case drain + reboot + kubelet rejoin + margin
For bare metal that typically means timeouts measured in 10+ minutes rather than the 300-second defaults many templates ship — a full drain of a packed multi-tenant node plus a slow firmware POST can consume most of that. And keep --concurrency consistent with maxUnhealthy: kured rebooting one node while MHC tolerates 40% unhealthy is fine; kured at concurrency 3 on a 6-node fleet with a tight maxUnhealthy is a capacity incident wearing a patch cycle as a disguise.
2. PDBs tenants actually set, and the infinite drain trap. Kured's drain honors PodDisruptionBudgets — which protects tenants that set them and says nothing about tenants that did not. On a multi-tenant fleet, "unattended reboots are safe" is only true if single-replica or PDB-less services are either prohibited by policy or explicitly accept eviction. Require PDBs (or multi-replica deployments with topology spread) for anything tenants care about, and enforce it with admission policy rather than documentation.
Then set --drain-timeout. The default is zero, meaning infinite: a drain blocked forever — a pod with a restrictive PDB that can never be satisfied, a finalizer that never completes — holds kured's lock forever, and every other patched node in the fleet queues behind it while running stale kernels. An infinite drain timeout converts one stuck node into a fleet-wide patch freeze. Set an explicit timeout, alert when a drain aborts, and treat an aborted drain as a tenant-configuration bug to fix, not a kured failure.
3. Block on live signal: alerts, pods, and windows. Kured can defer reboots in the presence of active Prometheus alerts (--prometheus-url, with --alert-filter-regexp to ignore the routine ones and --alert-firing-only to skip pending alerts) or selected pods (--blocking-pod-selector — the backup job, the migration runner, the "do not disturb" marker). Combined with day/time windows, this is how "reboot Tuesday at 2 a.m. unless something is on fire" gets encoded. A useful pattern: alerts that notify nobody, existing purely as reboot gates kured reads directly from Prometheus. At minimum, gate on your fleet's "stop all automation" alert set — the same signals that would make a human engineer postpone the runbook.
4. Lock and blast-radius hygiene. Keep --concurrency 1 until the fleet is large enough that serial reboots cannot finish inside the window; each increment trades patch velocity against simultaneously-drained capacity. Set --lock-ttl so a dead kured pod cannot wedge the lock past one cycle. Enable --prefer-no-schedule-taint so nodes awaiting their reboot stop accepting pods evacuated from the node currently rebooting — without it, the fleet can shuffle workloads onto a node about to go down. And treat --force-reboot (reboot even if the drain fails or times out) as the emergency override it is: reaching for it routinely means the drain configuration is wrong, and forcing past a failed drain is how single-replica tenants discover they were single-replica.
A starting flag set for a CAPI-managed Hetzner fleet:
--period=20m
--concurrency=1
--reboot-days=su,mo,tu,we,th
--start-time=01:00 --end-time=05:00 --time-zone=UTC
--prometheus-url=http://prometheus.monitoring.svc:9090
--alert-firing-only
--alert-filter-regexp=^(Watchdog|RebootRequired)
--drain-timeout=15m
--lock-ttl=1h
--prefer-no-schedule-taint=weave.works/kured-node-reboot
--annotate-nodes
--notify-url=<fleet-webhook>Tune the window to your tenants' quiet hours, the drain timeout to your slowest legitimate drain, and the MHC timeout to exceed both with margin — then verify the inequality holds after every change to any of the three.
Reboot in place or rotate the machine?
A CAPI fleet has a second OS-patch strategy kured does not replace: the MachineDeployment rolling update, where new machines boot a fresh image and old ones are deleted. Both close the stale-kernel gap; they differ in cost shape, and the honest answer is that most fleets want both at different cadences.
| Kured in-place reboot | MachineDeployment rotation | |
|---|---|---|
| Mechanism | Patch in place, reboot the same machine | New machines on a new image, cordon/drain/delete old |
| Speed | Minutes per node, serial by default | New-machine boot + image pull per node |
| Provisioning churn | None — same servers, same IPs | New servers provisioned, old released (Hetzner: API churn, re-imaging) |
| State risk | Preserves local disk state (feature and hazard) | Clean slate; local state must be rehydrated or externalized |
| Drift | OS converges by package manager; config drift accumulates | Image rebuild resets drift on every rotation |
| Best for | Weekly security patching between image releases | Monthly/quarterly image releases, Kubernetes minor upgrades |
On owned Hetzner hardware the rotation column costs more than it does in cloud: reprovisioning bare metal is slower than booting a cloud VM, and any workload with local state (caches, build artifacts, node-local volumes) pays a rehydration tax every rotation. That makes kured the right weekly driver — security patches land via unattended upgrades and take effect within one reboot window — with image rotation reserved for the slower cadence where a clean slate earns its cost: Kubernetes minor versions, base-image rebuilds, SSH CA rotations. The two compose: rotate quarterly, reboot weekly, and never run a kernel older than the last window.
One caution for the rotation purists: "immutable infrastructure, therefore no kured" is only true if rotations actually happen on patch cadence. A fleet that rotates quarterly and patches weekly has a two-month stale-kernel window it chose on purpose. If your rotation cadence cannot match your patch cadence — and on bare metal it usually cannot — kured is not a compromise of immutability, it is the mechanism that keeps the immutable story honest between rotations.
Adoption checklist for this week
- Install the DaemonSet from the kubereboot manifests or Helm chart, in
kube-system, with the guardrail flags above adjusted to your windows. - Verify the MHC inequality against your current
MachineHealthChecktimeouts; widen timeouts or narrow the reboot scope until drain-plus-reboot-plus-rejoin fits with margin. - Require tenant PDBs for anything multi-replica that matters, enforced by admission policy — then set
--drain-timeoutto a finite value and alert on aborted drains. - Wire the gates: Prometheus URL plus firing-only alert filters, blocking-pod selectors for backup/migration workloads, and reboot windows in your tenants' timezone.
- Canary it:
sudo touch /var/run/reboot-requiredon one node, watch drain → reboot → uncordon → notify, and confirm MHC stayed quiet. - Alert on the meta-signal:
kured_reboot_requiredheld past one full patch cycle means the cluster cannot reboot itself — page on that, not on individual reboots.
Unattended patching without automated reboots is a patch pipeline that stops one step short of patching. Kured is a small, boring daemon that finishes the job — one node at a time, behind a lock, inside your windows — provided you encode the four judgments above instead of assuming them. Do that, and the next kernel CVE lands, reboots, and clears without a human touching a checklist.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



