Skip to main content

HPA Scales to Zero Now. Should Your Fleet Retire KEDA or Knative?

10 min readDora NodaDora Noda
Share
On this page

After seven years behind an alpha feature gate, the Kubernetes Horizontal Pod Autoscaler can finally scale your workloads to zero replicas — and back — with no add-ons. Kubernetes v1.37, released August 26, 2026, graduated HPA scale-to-zero to beta and enabled it by default. So here is the verdict up front: queue consumers and batch workers can consolidate onto the native primitive today, but request-driven HTTP services and event-diverse workloads should stay on Knative and KEDA. The native HPA still cannot buffer the request that wakes a service, and it still has no catalog of event sources — the two things those projects exist to provide.

That sentence is worth unpacking carefully, because "scale to zero" sounds like one feature and is actually three different ones wearing a trench coat. What shipped in v1.37, what KEDA owns, and what Knative owns overlap far less than the marketing suggests. This post compares the three concretely and ends with a per-workload decision table you can apply to your own fleet.

What v1.37 actually shipped​

The feature is KEP-2021, led by SIG Autoscaling and announced on the Kubernetes blog on September 2, 2026. The headline: an HPA using a suitable object metric or external metric can now set minReplicas: 0, scale a workload down to nothing when idle, and bring it back when the metric moves. The HPAScaleToZero gate is on by default in both kube-apiserver and kube-controller-manager.

The history matters because it explains the shape of the API. The first alpha shipped in v1.16, back in 2019. Kubernetes v1.36 added the missing semantic piece — a ScaledToZero status condition — and v1.37 turned the whole thing on by default after adding integration and end-to-end coverage. Graduation to GA now waits on operational feedback from fleets like yours.

Three mechanics decide whether the feature fits your workload:

Only object and external metrics can wake a workload. CPU and memory percentages come from running pods. At zero replicas there are no pods left to measure, so there is no signal to scale back up on. A queue length, by contrast, exists independently of the workers consuming it. The API server enforces this: it rejects minReplicas: 0 on an HPA that carries only resource metrics.

In practice you expose something like a Prometheus queue_consumer_lag series through the External Metrics API — typically via prometheus-adapter — and point the HPA at it.

Zero is now two different states. A replica count of zero can mean "the HPA scaled this down" or "an operator paused this." The controller distinguishes them with the ScaledToZero condition: when it scales a workload from one or more replicas to zero itself, it records ScaledToZero=True and keeps evaluating metrics; a workload sitting at zero without that condition stays paused. The sharp edge: manually setting a Deployment to zero has always paused autoscaling, and v1.37 preserves that behavior — the HPA will not wake a workload it did not scale down. Start every candidate Deployment with at least one replica.

The defaults are conservative. Normal HPA behavior still applies, including the five-minute default downscale stabilization window, which stops a momentary dip in queue length from deleting your whole worker pool. Tune it through spec.behavior.scaleDown if your workload needs faster draining. And mind version skew during upgrades: wait until both the API server and the controller manager support the feature before creating zero-capable HPAs, and before any downgrade, restore minReplicas to at least one and scale zero-state workloads back up.

What each one actually owns​

With the native behavior pinned down, the three-way comparison gets concrete. Each tool answers a different question.

HPA scale-to-zero (v1.37+)KEDAKnative Serving
Core questionHow many replicas does this metric justify?Which events should wake this workload?Who serves the request while pods start?
Zero triggerObject or external metric value60+ built-in event scalers plus activation thresholdIncoming HTTP request after idle window (default 60s)
0-to-1 pathController loop observes metric, sets replicasScaledObject activates, drives generated HPA from zeroActivator buffers the first request, holds it until a pod is ready
Request bufferingNone — Services do not hold requests with no endpointsOnly via the separate KEDA HTTP add-on (a proxy in the request path)Built in — the request that wakes the service is the one that gets served
RolloutsWhatever your Deployment doesWhatever your Deployment doesRevision-based: traffic splitting, canary, blue-green
Extra componentsA metrics adapter you wire yourselfKEDA operator, per-source trigger authKnative Serving control plane, replaces Service with Route
MaturityBeta, default-on, pre-GA feedback phaseCNCF Graduated (2023), 2.x seriesCNCF Graduated (October 2025)

The table's first row is the whole argument. HPA computes replica counts from metrics. KEDA translates the event-driven world — Kafka consumer lag, SQS depth, RabbitMQ queues, cron schedules, Prometheus queries — into something the autoscaler can act on, and it has done scale-to-zero for years through its ScaledObject CRD and per-trigger activation thresholds. Knative owns the request path itself: concurrency-based autoscaling, the Activator that queues the wake-up request, and revision snapshots that make every rollout a traffic split. A git-push PaaS that consolidated "idle handling" onto HPA alone would get the replica math and lose the other two jobs.

The two things HPA still doesn't do​

First: no request buffering. The Kubernetes blog states it plainly — Services do not buffer requests while no pods are ready, so HTTP and other request-driven workloads need a separate buffering layer. Picture a zero-replica web service behind a ClusterIP Service: the first request arrives, there are no endpoints, and the request fails while the HPA notices the metric, schedules a pod, and waits for readiness.

Somebody has to hold that connection open. In Knative that somebody is the Activator, which receives the first request itself, triggers the cold start, and forwards the buffered request once a pod passes health checks — the caller sees a slow response, not an error. KEDA's answer is its HTTP add-on, an opt-in reverse proxy in the request path. Native HPA has no equivalent, and none is in scope for the beta.

Cold-start latency is the price either way, and it varies by orders of magnitude: seconds for a typical web pod, minutes for GPU model loads (recent LLM-serving benchmarks put 70B-parameter cold starts at 2–8 minutes on an H100). Buffering does not make the start faster; it decides whether the wake-up request survives it. For queue workers that distinction is irrelevant — the work waits in the queue. For synchronous HTTP, it is the entire feature.

Second: no event-source catalog. KEDA ships more than 60 built-in scalers, each bundling the connection, authentication, and polling semantics of one event source. With native HPA you assemble that plumbing per source: instrument the metric, expose it through the External Metrics API, verify it with a raw API read, then write the HPA.

For one Prometheus series that is a morning's work. For Kafka lag plus SQS depth plus a cron schedule plus a cloud queue, it is a small platform project — one KEDA already finished, with trigger authentication (TriggerAuthentication, file-based auth in recent releases) included. The ecosystem gap compounds: every new source your tenants adopt is new adapter config on HPA versus a new ScaledObject trigger on KEDA.

The decision, per workload type​

Consolidation is not all-or-nothing. Most self-hosted fleets run all four of these shapes, and the right answer differs by row:

WorkloadVerdictWhy
Queue consumers, batch processorsConsolidate onto HPAWork waits in a durable queue; one external metric per consumer; drops an operator from the critical path
Cron and multi-source event workloadsKeep KEDACron plus 60+ source scalers with bundled auth beat hand-wired adapter config per source
Request-driven HTTP, serverless-style servicesKeep KnativeOnly the Activator pattern survives the wake-up request; revisions give you canary and blue-green for free
GPU inference endpointsKeep Knative (via KServe) or KEDAMinute-scale model cold starts demand buffering or queue semantics; HPA's metric loop alone strands wake-up requests

Two nuances before you act on the table. First, KEDA and the native HPA are not rivals architecturally — KEDA generates and manages HPA objects under the hood, so "keep KEDA" still means your scaling decisions flow through the standard autoscaler. Retiring KEDA for queue workers that already expose a clean metric is a genuine simplification; ripping it out of a fleet with a dozen event sources is trading a solved problem for a wiring project. Second, Knative's cost is real: it replaces your Service with a Route and adds a serving control plane to operate. Pay it where request-driven scale-to-zero is the product — preview environments, per-tenant inference, idle customer sandboxes — not where a queue already holds the work.

How to trial HPA scale-to-zero without an outage​

Beta means "enabled by default," not "proven in your fleet." If the table above points a workload class at the native primitive, roll it out in this order:

  1. Pick a queue worker, not a web service. Durable queues absorb the cold-start gap that synchronous HTTP cannot.
  2. Verify the metric endpoint first. Read the external metric through the API (kubectl get --raw against external.metrics.k8s.io) before creating the HPA. An HPA whose metric is unavailable cannot scale from zero — debug the pipeline, not the autoscaler.
  3. Start at one replica, set minReplicas: 0, watch the condition. kubectl describe hpa shows ScaledToZero=True once the controller owns the zero state. If a workload sits at zero without it, someone paused it manually.
  4. Tune spec.behavior.scaleDown deliberately. The five-minute default stabilization is usually right for queues; shorten it only if you have measured idle cost that justifies the churn risk.
  5. Mind the control plane. During upgrades, confirm both API server and controller manager support the feature before relying on it — a skewed controller can mistake controller-owned zero for a manual pause. And never downgrade past it without first restoring minReplicas to one and waking zero-state workloads.

Run that loop for a full business cycle — including your quietest weekend and your noisiest deploy day — before promoting the pattern from trial to default. SIG Autoscaling is explicitly gathering operational feedback ahead of GA; production evidence from small fleets counts as much as hyperscaler data here.

The native primitive wins the queue; the specialists keep the request path​

Step back and the shape of the answer is satisfying. Kubernetes spent seven years deciding what "zero" means — an owned, conditioned, metric-driven state rather than an absence the controller ignores — and v1.37 is the release where that semantic becomes the default. Queue-driven autoscaling no longer needs an add-on, and every self-hosted fleet should take that simplification where it fits.

But scale-to-zero was never just replica math. Somebody has to speak the event source's protocol, and somebody has to hold the wake-up request. Those are KEDA's and Knative's jobs, both CNCF-graduated, both still without a native substitute. Consolidate the metric loop; keep the specialists on the request path.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex