Your cluster-autoscaler taints a node ToBeDeletedByClusterAutoscaler and starts draining it. At the same moment, your MachineHealthCheck notices the node's pods disappearing, the DaemonSet controller notices a daemon pod going missing and helpfully reschedules it — onto the very node being deleted. Three controllers, three different stories about the same machine, all of them locally reasonable and globally wrong.
Kubernetes 1.37, Garhwal, shipped August 26, 2026 with the first step toward ending that comedy of errors: five well-known Node Lifecycle Conditions, introduced on the Kubernetes blog on September 9 by Ryan Hallisey (NVIDIA). They give the Node API a shared, Kubernetes-owned vocabulary for saying what a machine is actually doing:
| Condition | What it reports |
|---|---|
DrainInProgress | The node is actively being drained per the administrator's chosen drain criteria |
Drained | The node has reached the administrator's selected drain criteria |
MaintenancePlanned | The node is expected to undergo a change in the future |
MaintenanceInProgress | The node is actively undergoing maintenance |
GracefulNodeShutdownInProgress | Graceful Node Shutdown is determined to be in progress on the node |
If you run a Cluster API fleet — CAPH on Hetzner, CAPD in a lab, or anything in between — this is the future shape of every pool-scaling and health-check signal you currently hand-wire. Here is what shipped, what it fixes, and what to actually do about it this quarter.
What v1.37 actually ships (read the fine print)
The honest version first: in v1.37, almost nothing consumes these conditions. The release reserves the five names as well-known NodeConditionType constants and adds an alpha NodeLifecycleConditions feature gate, disabled by default. The blog post says it plainly — in this release the gate is effectively a no-op. It does not restrict who can set the conditions, and no core component reads them.
That sounds like a letdown until you see the shape of the plan. This is KEP-5683, led by SIG Node with the Node Lifecycle Working Group and SIG Apps, and v1.37 is deliberately the vocabulary release: agree on the words first, then teach controllers to read them. Two practical consequences fall out of that design:
- You do not need to enable the gate to start publishing the conditions today. Any administrator or administrator-authorized controller can set and clear them on any 1.37 node right now. The gate exists so the built-in consumer behavior planned for future releases can be opted into when it arrives.
- Publishing is all you get. No core workload controller changes its behavior based on these conditions yet. Their immediate value is operational clarity: a common status channel for lifecycle work that already happens today.
Like every other Node condition, each one carries status (True while the state is active, False or removed when it is not, Unknown when Kubernetes cannot tell), a stable machine-readable reason, and a human-readable message. The recommended pattern is explicit: use lifecycle conditions to report status, and keep managing operations through the existing mechanisms — kubectl cordon, kubectl drain, taints, and workload-specific controls. An authorized maintenance controller publishing a planned window looks like this:
status:
conditions:
- type: MaintenancePlanned
status: "True"
reason: MaintenanceWindow
lastTransitionTime: "2026-12-09T12:00:00Z"
message: "Hardware maintenance is scheduled for this Node"One more rule from the announcement worth adopting on day one: decide which component owns each condition type, so two writers never fight over the same condition. On a CAPI fleet, that ownership question is the whole game — more on that below.
The problem: everybody infers, nobody agrees
To appreciate why a shared signal matters, inventory how your fleet answers "what is happening to this node?" today. Readiness comes from kubelet heartbeats. Scheduling intent comes from taints. Pod state comes from the workload controllers. Provider truth comes from your infrastructure provider's API. Your own automation probably adds labels or annotations on top.
Each signal is useful for its purpose, but none of them answers the lifecycle question — and the announcement names three real failure modes where independently correct components make conflicting decisions from indirect signals:
- The DaemonSet that fights the kubelet. During graceful node shutdown, the kubelet intentionally terminates pods. The DaemonSet controller sees a missing daemon pod and replaces it — rescheduling work onto a node that is trying to shut down. Neither component is buggy; they just disagree about what the node is doing.
- The Job that waits forever. A Job controller waits for a terminal pod phase on a node an administrator is already removing. Nothing tells the controller the wait is futile, so it waits.
- The storage operator that learns too late. Maintenance begins, drain starts, and only then does the storage layer discover the node was going away — after the ordering decision it needed to participate in has already been made.
The same class of conflict lives inside every CAPI fleet's scaling loop. A NotReady node does not explain whether the cause is an unexpected failure, a graceful shutdown, or planned maintenance — yet MachineHealthCheck remediation and the autoscaler's scale-down logic must each guess, from different angles, right now.
There is also a subtler, long-standing edge case the announcement calls out as the motivating future work: a node that is broken or under maintenance still consumes a DaemonSet rollout's availability budget, which can slow or block the rollout on healthy nodes. The DaemonSet controller knows a pod is unavailable but cannot tell whether the new revision failed or an administrator intentionally took the node out of service. MaintenanceInProgress creates the Kubernetes-owned place to publish that context, so future work can define rollout ordering, availability accounting, and status reporting around it — no more manually adjusting rollouts during maintenance windows.
The before/after for a CAPI fleet's signals
Here is the concrete mapping for a Cluster API operator: what you wire by hand today, and which lifecycle condition each signal grows into.
| Today's bespoke signal | What it actually tells you | Lifecycle condition to publish alongside it |
|---|---|---|
MHC unhealthyConditions (Ready Unknown/False + timeout, typically 300s) | The kubelet stopped reporting; cause unknown | MaintenanceInProgress (set by your remediation automation when it decides to fix the machine, so other controllers stop guessing) |
Cluster-autoscaler ToBeDeletedByClusterAutoscaler taint | The autoscaler intends to remove this node | DrainInProgress, then Drained when your drain criteria are met |
| Drain scripts watching pod termination | Pods are gone; drain probably finished | Drained with a stable reason naming the criteria (e.g. all non-daemon pods evicted) |
| Out-of-band maintenance calendar / runbook | A human knows a window is coming; the cluster does not | MaintenancePlanned when the window is scheduled, flipped to MaintenanceInProgress when work starts |
| Kubelet graceful-shutdown config firing | Visible only in kubelet logs and pod churn | GracefulNodeShutdownInProgress |
Three things to notice about this table.
First, nothing in the left column goes away. Your MHC timeouts still decide remediation. The autoscaler's taint still drives scale-down. Cordon, drain, and taints still change behavior. The right column is a reporting layer, not a replacement control plane. If anyone on your team reads the announcement as "1.37 replaces MachineHealthCheck," correct them early.
Second, the value lands the moment two consumers read the same condition. Today your drain automation and your remediation automation each maintain a private theory of node state — taint scraping here, provider-ID matching there, a label your install script added two years ago. When both publish and read DrainInProgress / Drained, the "is it safe to delete this Machine yet?" question stops being a bespoke query against three APIs and becomes a condition read. On a small CAPH fleet where one engineer is the platform team, that consolidation is the difference between a runbook and a folk tale.
Third, ownership discipline is the price of admission. The announcement explicitly tells administrators to decide which component owns each condition. A sane first split for a CAPI fleet: your maintenance/upgrade automation owns MaintenancePlanned and MaintenanceInProgress; your drain wrapper owns DrainInProgress and Drained; node-shutdown detection owns GracefulNodeShutdownInProgress. Write that mapping down before you publish your first condition, or you will recreate the same conflicting-writers problem at a new layer.
What to actually do this quarter
This is a vocabulary release, so the rational response is to start speaking the vocabulary — cheaply, without betting behavior on it:
- Publish conditions from automation you already run. The lowest-cost adoption is a few lines in your existing drain wrapper and maintenance scripts: set
MaintenancePlannedwhen a window is scheduled,MaintenanceInProgresswhen work starts,DrainInProgresswhen eviction begins,Drainedwhen your criteria are met. No feature gate needed. Dashboards, alerts, and on-call humans get a shared signal immediately. - Keep every behavior decision on the old mechanisms. Remediation still keys off MHC conditions and timeouts. Scale-down still follows the autoscaler. Nothing in 1.37 reads the new conditions, so gating behavior on them today means gating behavior on a signal only you write — a single point of confusion, not a safety improvement.
- Write down the ownership map. One writer per condition type, documented where your team will find it during an incident. Revisit it when you adopt any tool that also publishes lifecycle conditions.
- Track KEP-5683 follow-ups, especially DaemonSet rollout accounting. That is the first consumer behavior on the roadmap that directly changes fleet operations: maintenance-aware availability budgets. When it lands behind the
NodeLifecycleConditionsgate in a future release, you will want a staging fleet ready to opt in. - Do not wait for GA to standardize your
reasonstrings. The machine-readable value of these conditions comes from stable reasons. Pick yours now (MaintenanceWindow,KernelUpgrade,ScaleDownDrain) so that when consumers arrive, your history is already parseable.
Longer term, the announcement hints the coordination problem may outgrow conditions entirely — explicit ownership, locking, and potentially a dedicated API are all on the table. That is a healthy sign: the working group is scoping the mechanism to the problem rather than overloading conditions forever. But conditions are the interop layer every later mechanism will build on, so publishing them now is not throwaway work.
The fleet that speaks the same language
Step back and notice the direction of travel. Kubernetes spent the last few years making node state observable — PSI metrics for real contention signals, DRA for device-aware scheduling, workload-aware placement. Node Lifecycle Conditions make node state legible: not another metric to alert on, but a shared sentence about what a machine is doing that every controller can read without reconstructing it from shadows.
For a self-hosted CAPI fleet, that legibility compounds. Your remediation loop, your autoscaler, your upgrade automation, and your future GPU node pools all stop maintaining private theories of node state and start reading the same five words. The 1.37 release only teaches the cluster to say the words — teaching it to act on them comes next, through KEP-5683 follow-ups you can already follow and shape in the Node Lifecycle Working Group. Start publishing now, and your fleet will be fluent by the time anything is listening.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



