On August 26, 2026, Kubernetes 1.37 — codenamed Garhwal — shipped with 67 enhancements, and one quiet graduation changes the idle-cost math for every multi-tenant platform: HPA scale-to-zero is now beta and enabled by default. A single line, minReplicas: 0, and your idle queue workers, preview environments, and off-hours batch consumers vanish completely instead of burning a full pod's worth of CPU and RAM around the clock.
Here is the punchline the release notes bury: scaling to zero was never the hard part. Scaling back from zero is where workloads go to die. Two deadlocks strand a scaled-to-zero workload at zero replicas forever — a metric rule that can never fire without running pods, and a status condition that only exists if the HPA itself did the parking. If you run a git-push PaaS and plan to promise scale-to-zero to tenants, this post is the pre-flight checklist that keeps that promise from becoming a 3 a.m. page.
What 1.37 actually changed
The HPAScaleToZero feature gate (KEP #2021, SIG Autoscaling) spent 1.36 in alpha, off by default — you had to flip it on the controller manager, which most managed offerings never exposed. In 1.37 it graduated to beta and turned on by default, and the official Kubernetes blog followed with the full specification on September 2. On any 1.37 cluster, including GKE 1.37 and later, minReplicas: 0 just works.
Three rules govern the whole feature:
| Rule | What it means |
|---|---|
| Only object or external metrics can wake a workload | Queue length, a Prometheus value, a cloud pub/sub backlog — anything that exists independently of your pods |
| CPU and memory can never scale from zero | Resource metrics come from running pods; at zero pods there is nothing to measure |
The ScaledToZero condition records who parked it | ScaledToZero=True means the HPA drove the workload to zero and can drive it back; anything else at zero replicas is just… at zero |
The canonical shape is a queue worker:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: queue-worker
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: queue-worker
minReplicas: 0
maxReplicas: 10
metrics:
- type: External
external:
metric:
name: queue_length
target:
type: AverageValue
averageValue: 5Queue drains, workers evaporate. Tasks arrive, the external metric is still readable with nobody running, and the HPA scales back up. That is the happy path, and it genuinely works. Everything below is about the three ways the happy path silently stops being happy.
The two deadlocks that strand you at zero
Deadlock 1: the metric rule that can never fire
This one is pure logic, and it bites anyone who ports an existing CPU-based HPA to zero by changing one number. The HPA computes desired replicas from metric values. CPU and memory utilization are measured per running pod. Zero pods means zero measurements, which means no signal that could ever justify scaling up. The workload sits at zero with an HPA that is technically healthy, faithfully reading metrics that do not exist.
Google's GKE troubleshooting docs state it bluntly: the HPA cannot scale to zero on CPU or memory alone — you must configure at least one External or Object metric. Note the asymmetry that makes this a deadlock rather than an error: the HPA will happily scale down to zero on resource metrics (zero utilization looks like zero demand), and then can never come back. Down works; up is impossible. That is the exact shape of a trap.
Deadlock 2: the condition that was never set
This is the subtler one, and it catches platforms rather than individuals. The HPA controller only scales a workload from zero if it was the one that scaled it to zero — it tracks that memory in the ScaledToZero condition. If a workload is sitting at zero replicas without that condition, the controller treats it as someone else's decision and refuses to override it.
Three everyday paths lead here:
- Deploying at zero. You ship a manifest with
replicas: 0and attach a fresh HPA. The controller sees zero replicas, noScaledToZerocondition, and never activates. Field reports flag this as the least-documented gotcha in the beta. - A platform reaper doing its old job. Your PaaS already has idle-reaping logic that parks unused tenant apps by setting replicas to zero directly. The moment you also add an HPA, the reaper's parking looks identical to a manual scale-down — no condition, no wake-up.
- A human with kubectl.
kubectl scale deployment/foo --replicas=0during an incident, followed by everyone assuming the HPA will bring it back when traffic returns. It will not.
The third failure mode: a blind HPA freezes
External metrics are a dependency, and dependencies fail. When the metric source stops answering — the Prometheus adapter is down, the cloud monitoring API is throttled, the queue exporter wedged — the HPA shows the metric as <unknown> and freezes scaling decisions rather than guessing. (1.36 added an alpha fallback on retrieval failure under KEP #5679, but holding position is still the default behavior.) A workload parked at zero with a dead metric source is locked there until telemetry recovers, and nothing in the default HPA status screams about it.
| Failure mode | Symptom | Detection | Fix |
|---|---|---|---|
| Resource-metric-only HPA | At zero forever despite demand, HPA looks healthy | kubectl describe hpa shows only Resource metrics | Add an object/external metric; never ship CPU-only with minReplicas: 0 |
| Zero without the condition | Fresh or externally-zeroed workload never activates | ScaledToZero condition absent at zero replicas | Start at ≥1 replica and let the HPA park it; route all parking through the HPA |
| Metric source outage | Metric reads <unknown>, decisions frozen | HPA conditions + metric staleness alerts | Recover the telemetry path; treat it as control-plane-critical |
Which metric wakes which workload
The workload table is the heart of the tenant-facing promise. Not every workload type has a natural wake-up signal, and the honest answer for one of them is "not yet":
| Workload | Safe wake-up metric | Notes |
|---|---|---|
| Queue / batch worker | External queue length or backlog depth | The ideal case — the metric literally counts pending work |
| Scheduled / cron-like | Object metric on the schedule, or a time-based external metric | Cron triggers are still KEDA's home turf; evaluate before hand-rolling |
| HTTP web service | Requests-per-second via gateway or Prometheus adapter | Works, but read the cold-start caveat below before promising it |
| Idle preview environments | Activity signal (recent deploys, recent requests) published as an external metric | The platform must own publishing this signal reliably |
The HTTP caveat deserves its own paragraph because it is the one tenants will trip over first. As the official 1.37 blog notes, Kubernetes Services do not buffer requests while no pods are ready. Scale a web deployment to zero and the next request arrives to find no endpoints: it waits or fails until a pod boots, pulls, passes readiness, and joins the endpoints list. That is a cold start measured in seconds at best, and no HPA setting shortens it. Your options are honest ones: accept the cold start for internal and preview workloads, keep latency-sensitive production services at a minimum of one replica, or adopt a request-buffering activator in the Knative style. What you cannot do is promise both zero idle cost and warm latency on a bare Service. Put that sentence in your tenant docs verbatim and you will save yourself a quarter of support tickets.
Operate it: catch a zero-lock before the tenant does
A zero-locked workload looks exactly like an idle one from the outside — that is what makes it dangerous. Build detection into the platform before you enable the feature, not after the first incident.
Start with the HPA conditions, which are the controller telling you what it believes. kubectl describe hpa <name> shows the Conditions block: scaling ability, metric validity, and the ScaledToZero state. A healthy parked workload reads ScaledToZero=True with live metric values. The two alarm patterns are a metric stuck at <unknown> (telemetry path broken) and zero replicas with no ScaledToZero condition at all (someone else parked it; the HPA will never wake it).
kubectl describe hpa queue-worker
kubectl get hpa queue-worker -o jsonpath='{.status.conditions}'
kubectl get --raw /apis/external.metrics.k8s.io/v1beta1 | head -c 300The first two show what the HPA believes; the third checks whether the external-metrics API it depends on is even answering. Then encode the dangerous states as alerts, not runbook entries: fire when an HPA sits at zero desired replicas while its wake-up metric reports pending work, and when any minReplicas: 0 HPA reports <unknown> metrics for more than a few minutes. The tenant should never be the first monitoring system to notice.
One owner for the replica count
This is the architectural rule the whole feature turns on: exactly one writer may own spec.replicas on any workload with an HPA, and that writer is the HPA. Every deadlock in this post is some version of two controllers fighting over one number.
Concretely, that means three changes for a platform with existing idle management. First, the platform's idle reaper must stand down wherever an HPA exists — parking by direct replica edits creates deadlock 2 by construction. If the platform still wants a say in idling, it expresses that say through the HPA's own inputs: publish the demand metric, adjust minReplicas, or suspend the HPA — never write spec.replicas behind its back. Second, audit GitOps the other direction: a reconciler that "corrects" replicas back to the manifest's value will fight the HPA upward just as steadily as a reaper fights it downward, so HPA-managed targets need spec.replicas ignored in the diff. Third, document the handoff for humans — after an incident scale-down, the recovery step is restoring replicas manually or re-creating the HPA's parking cycle, not assuming autoscaling resumes on its own.
Do you still need KEDA or Knative?
Fair question, and the honest answer is narrower than either camp admits. A large fraction of "HPA vs KEDA" guidance still states flatly that native HPA cannot scale to zero — that was true for a decade and is now outdated as of 1.37. For queue workers with a Prometheus or cloud-queue metric, native HPA with minReplicas: 0 covers the case with one fewer operator to run.
KEDA keeps three things native HPA does not have: an activation threshold distinct from the scaling target (scale 0→1 on any activity, then 1→N on depth), more than sixty event-source scalers without hand-wiring metric adapters, and fallback behavior when telemetry dies. Knative keeps the one thing neither HPA flavor has: a request-buffering activator that holds HTTP traffic during cold start. If your tenants are event-driven workloads on exotic sources, keep KEDA. If they are HTTP services that must wake without dropping requests, you still need Knative's activator pattern. If they are queue workers and preview apps on metrics you already export, 1.37 lets you delete an operator. Choose per workload, not per platform.
The pre-flight checklist
Before your PaaS promises scale-to-zero to a single tenant, every line below should be true:
- Every
minReplicas: 0HPA carries at least one object or external metric. CPU/memory-only plus zero is a documented trap, not a configuration. - No workload starts at zero. Deploy at one or more replicas and let the HPA do the first parking run so the
ScaledToZerocondition exists from day one. - The HPA is the sole writer of
spec.replicas. Idle reapers, GitOps reconcilers, and incident runbooks all route through HPA inputs — never direct replica edits. - The metric pipeline is monitored as control plane.
<unknown>on a wake-up metric pages like an apiserver issue, because for a parked workload it is one. - Zero-with-demand alerts exist. Desired replicas at zero while the wake-up metric shows pending work fires before the tenant notices, with the
describe hpacommands above attached to the alert. - HTTP cold start is disclosed, not discovered. Tenant docs state plainly that zero means cold starts on bare Services, with the keep-one-warm and activator options named.
- KEDA/Knative scope is decided per workload. Event-exotic workloads stay on KEDA, request-buffered HTTP stays on Knative patterns, and plain queue workers move to native HPA — deliberately, not by default.
Kubernetes 1.37 turned a decade-old "you need an add-on for that" into a one-line beta default — a genuine maturation moment for the autoscaling API. But defaults are not guarantees. The metric rule and the status condition are small, legible mechanics, and a platform that respects them gets idle-to-zero as a quiet cost win. A platform that discovers them via incident gets the same education at 3 a.m., priced per tenant. Run the checklist; keep the quiet win.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



