Skip to main content

Kubernetes 1.37 Locks In On-Demand PLEG Relist: What Cost-Sized Nodes Actually Save

9 min readDora NodaDora Noda
Share
On this page

Every second, on every node, your kubelet asks the container runtime the same question about every pod: anything changed? It has done this since the PLEG — the Pod Lifecycle Event Generator — was designed, one global sweep per second whether anything happened or not. Kubernetes 1.37, Garhwal, released August 26, 2026, finally retires that polling loop as the only way kubelet learns what it just did itself: PLEGOnDemandRelist is now GA, locked on, with the off switch slated for removal in v1.40.

The headline numbers, measured by SIG Node on kind and published in the original change (averages over 1,000 runs, 10 operations in parallel):

OperationBefore (global relist)After (on-demand relist)
Pod create, request → Running1.798s1.136s
Pod resize, total1.778s0.796s
Pod resize, complete phase alone1.004s0.075s

Same cluster, same runtime — the only difference is kubelet re-inspecting the pod it just acted on instead of waiting for the next one-second sweep. If you run cost-sized nodes with no headroom to spare, this is the kind of upstream efficiency fix that lands directly on your deploy latency and your 3 a.m. pager. Here is what changed, why it nearly shipped turned off, and how to verify it during your 1.37 upgrade.

What PLEG does, in sixty seconds​

PLEG converts container-runtime state into pod-level events. When a container starts, dies, or changes state, something has to notice and wake the right pod worker — that something is PLEG, and for most of Kubernetes' life it has noticed by relisting: calling the CRI ListContainers endpoint, diffing the result against its cache, and emitting events for whatever moved.

Two properties of that loop matter for operators. First, the sweep is global and periodic: the default relist period is one second (genericPlegRelistPeriod), and every sweep covers every container on the node. A pod the kubelet just created itself still waits up to a full period before its Running state is observed — pure dead time injected into every pod transition.

Second, the loop is load-bearing for node health: if too long passes without a successful relist (the threshold is three minutes), kubelet reports PLEG is not healthy and the node flips NotReady, taking every pod on it out of service rotation. On an overloaded node, the polling loop that detects trouble becomes the thing that causes it.

That is the tax on-demand relist cuts: not the safety-net sweep, which stays, but the wait for it on the path where kubelet already knows something changed — because kubelet changed it.

What on-demand relist actually changes​

The insight, stated plainly in tallclair's change (merged March 13, 2026, as a kind/bug): most container runtime operations are synchronous. When kubelet takes an action in SyncPod — start a container, resize a pod — the result is visible as soon as the CRI call returns. So instead of waiting for the next global relist, kubelet can immediately re-inspect just that pod.

Mechanically, the change adds a RequestRelist(podUID) method to the PodLifecycleEventGenerator interface: pod workers queue the pod they acted on, and PLEG relists that pod out of band. The periodic global sweep keeps running underneath as the safety net for changes kubelet did not initiate (a container OOMing, a runtime-side restart). The 1.36 release note put it in one line: "kubelet: relist pods on-demand for lower latency operations."

Operators also get a new instrument: kubelet_pleg_pod_relist_duration_seconds, an ALPHA histogram tracking how long a single-pod relist takes — the per-pod counterpart to the existing kubelet_pleg_relist_duration_seconds (whole-sweep duration), kubelet_pleg_relist_interval_seconds, and kubelet_pleg_last_seen_seconds. If you graph PLEG health today, that new histogram is the series that tells you the on-demand path is actually firing.

The bumpy road to GA: on, off, on again, locked​

This feature nearly shipped dark in 1.36, and the timeline is worth knowing because it is a worked example of how SIG Node earns a GA:

DateEvent
Mar 13, 2026#137362 merges: on-demand relist ships in 1.36 as beta, default-on, with the gate as an emergency-off switch
Mar 19, 2026#137909: default flipped to off — the new path sharply increased failures in the should never report container start when an init container fails e2e test
Mar 21, 2026#137749: the flake is root-caused — a pre-existing issue (#135713, open since December 2025) where runc reports exit code 2 instead of 1 when signaled before execve handoff; the test is fixed to assert the real invariant
Mar 25, 2026#137946: the flip is reverted, gate re-enabled — 1.36 ships with on-demand relist on by default
Jul 22, 2026#140805: graduated to GA for 1.37 — "We haven't received any major issue reports with the new functionality"
v1.40 (planned)Gate removed entirely (LockToDefault: true in the feature registry)

Two things to take from this. First, the scare was a test bug surfacing a real runtime quirk, not a design flaw — the faster kubelet observed container state, the more often it tripped over a pre-existing exit-code inconsistency. That is precisely the class of issue a beta period exists to shake out.

Second, and operationally more important: in 1.37 there is no opt-out. The gate is locked to true, and one downstream gate already depends on it — EventedPLEG now lists PLEGOnDemandRelist as a prerequisite in the feature registry. You are getting this behavior on upgrade; the only rollback is a kubelet downgrade. Which makes the verification section below not optional.

What it saves on a cost-sized node pool​

Here is the honest accounting. Upstream published latency numbers, not CPU numbers — nobody has published a "this many millicores saved per node" figure, and this post will not invent one. What the numbers and the mechanism together support is a latency story with a clear sensitivity variable: pod churn rate. The saving per pod operation is roughly two-thirds of a second on the create path (1.798s → 1.136s) and roughly a second on the resize completion phase — and it multiplies by how often your nodes do pod operations.

Churn scenarioPod opsWhere the saving lands (extrapolated)
Quiet node, steady stateA handful per hourSeconds per day — negligible; the global sweep was never your bottleneck
Rolling deploy of a 20-pod service~20 creates + reads per rolloutRoughly 10–15 seconds of status latency off the rollout wall clock
Node drain, HPA storm, or crashloop recovery100+ pod restarts in minutesA minute or more of container-status lag eliminated — exactly when stale status hurts most

Two caveats, stated plainly. These extrapolations scale the PR author's kind-cluster benchmarks (64 CPUs, cached images) to operation counts; on a 2-vCPU Hetzner node your image pulls, storage, and CNI dominate absolute times, so measure your own Scheduled → Running distributions rather than quoting this table at your team. And the CPU story is mechanism, not measurement: per-pod relists replace whole-node sweeps on the action path, so a churning node does less redundant ListContainers work per event — but an idle node's global sweep still runs every second, so expect no change at rest. The metric pair that settles it on your fleet is kubelet_pleg_relist_duration_seconds before and after the upgrade, plus the new kubelet_pleg_pod_relist_duration_seconds to confirm the fast path is firing.

Why does any of this matter disproportionately to cost-sized fleets? Because every one of these savings spends where headroom is thinnest. On a 2–4 vCPU node, kubelet and runtime overhead compete directly with tenant workloads — there is no spare core absorbing a status-latency spike during a rollout.

Stale pod status delays readiness gates, which stretches rollouts, which holds old and new ReplicaSets alive simultaneously — precisely the memory doubling a packed node cannot afford. And PLEG is not healthy NotReady flaps get less likely when each pod event costs a single-pod re-inspection instead of contributing to sweep pile-up. Hyperscaler node pools absorb all of this with slack; a fleet sized for cost feels each piece.

Verify it during your 1.37 upgrade​

Do this in order, ideally with one canary node pool on 1.36 as the control while the rest moves to 1.37:

  1. Baseline before the upgrade. Capture a week of kubelet_pleg_relist_duration_seconds (p50/p99), kubelet_pleg_relist_interval_seconds, and kubelet_pleg_last_seen_seconds per node, plus your deploy pipeline's Scheduled → Running p50. You cannot claim a win you did not measure.
  2. Confirm the fast path fires after the upgrade. kubelet_pleg_pod_relist_duration_seconds should appear and accumulate observations on every 1.37 kubelet. No series, no on-demand relists — investigate before declaring victory.
  3. Compare the global sweep. Whole-sweep duration should hold steady or drop under churn; what must not happen is a rise in kubelet_pleg_relist_interval_seconds (sweeps spacing out under load) or gaps in pleg_last_seen_seconds. Those are the early warnings for the old pile-up pathology.
  4. Watch rollout latency, not just node metrics. The user-visible win is First Sync → Running shrinking toward the ~1.1s regime from ~1.8s on cache-warm creates. If your p50 does not move, your bottleneck is image pulls or CNI attachment — useful knowledge either way.
  5. Grep for the old failure signature. PLEG is not healthy in kubelet logs plus NotReady flaps correlated with pod churn should decrease across the upgrade boundary, not increase. An increase means something in your runtime version disagrees with the faster observation cadence — check your containerd version against the 1.37 skew policy before blaming the feature.
  6. Remember there is no per-node toggle. The gate is locked on; a canary that misbehaves gets a kubelet version rollback, not a flag flip. If you run EventedPLEG, note it now depends on this gate — watch evented_pleg_connection_* alongside the PLEG series.

Small loops, owned hardware, compounding wins​

Nobody migrates to Kubernetes 1.37 for PLEG relist latency. But a self-hosted fleet on owned machines is a stack of small loops — scheduler, kubelet, runtime, CNI — each charging its own per-operation tax, and nobody hands you headroom to cover the total. Upstream efficiency fixes like this one are the rare changes that cut the tax without asking anything of you: no migration, no new CRD, no operator to install. The price is attention — knowing which loop changed, what it saves, and how to prove it on your own graphs.

That is also the argument for tracking the freeze calendar, not just the feature list: this GA was knowable the day #140805 merged on July 22, five weeks before 1.37 shipped. A pre-flight audit against the changelog beats a post-upgrade mystery every time — especially when the change, like this one, ships without an off switch.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex