Skip to main content

Kubernetes 1.37 Takes HPA Scale-to-Zero to Beta: What Idle Finally Costs vs Fly.io, Render, and Railway

8 min readDora NodaDora Noda
Share
On this page

On September 2, 2026, Kubernetes quietly retired one of the oldest lines in every autoscaling comparison: "HPA can't scale to zero." The v1.37 announcement graduates HPA scale-to-zero to beta, enabled by default, so a plain in-tree HorizontalPodAutoscaler can now take a workload from zero replicas back to serving without Knative, KEDA, or any second control plane on the queue path. If you run a self-hosted fleet, the idle bill just changed shape. Here is the new math, up front.

The idle bill, up front​

One small service (smallest comparable size per platform) at four duty cycles, September 2026 list prices:

Duty cycleSelf-hosted K8s + HPA to zeroFly.io (min_machines_running = 0)Render Starter ($7)Railway Hobby ($5 floor)
Near-idle (~5%)~$0.04*~$0.10$7~$5
Side project (~10%)~$0.08*~$0.20$7~$5
Business hours (~50%)~$0.38*~$1.01$7~$5–12
Always-on (100%)~$0.75*~$2.02$7~$20

* Share of one ~$6/mo shared node (2 vCPU / 4 GB class), bin-packed with other tenants, assuming idle stretches longer than the 5-minute default downscale stabilization window. A single service alone on a dedicated box still pays the full node floor — scale-to-zero returns capacity to the pool, it does not unplug the machine.

Three things jump out. First, idle was already cheap on Fly.io — per-second billing plus auto-stop meant a 10%-duty side project cost about twenty cents — and self-hosted Kubernetes now matches that shape with in-tree primitives. Second, Render and Railway don't move at any duty cycle: Render's $7 Starter never sleeps on a paid plan, and Railway's meter runs warm regardless of traffic. Third, the gap is widest exactly where hobbyists and internal tools live: at 10% duty, the always-on platforms cost 25–90x the scale-to-zero options for the same service.

The rest of this post earns that table: what actually shipped, why the managed columns look frozen, and the one honest gap that keeps "idle costs zero" from being a promise you can make on day one.

What actually shipped on September 2​

The feature is KEP-2021, and its headline properties fit in two short paragraphs. In v1.37 the HPAScaleToZero gate is on by default in both the API server and the controller manager. An HPA driven by an object or external metric — queue depth, pending-work count, anything that exists independently of running pods — may set minReplicas: 0. The API server rejects minReplicas: 0 on CPU/memory-only HPAs, because at zero pods there is no CPU signal left to wake up on.

A new ScaledToZero status condition records whether the controller owns the zero state, so automatic scale-down and an operator's manual pause are no longer ambiguous. Normal HPA behavior still applies, including the five-minute default downscale stabilization window that stops a momentary dip in queue length from deleting every worker.

"Native" deserves one explicit sentence, because it is the title's load-bearing word. Before v1.37, every scale-to-zero story on Kubernetes ran through an add-on: Knative Serving with its activator, KEDA with its 60+ scalers, or the alpha gate almost nobody enabled in production. Now the queue-consumer path — external metric in, replicas out — is served by the HPA controller you already run. KEDA is still the right answer for trigger variety and for HTTP (more on that below); it is no longer mandatory for "wake my workers when the queue fills."

The minimal shape, adapted from the announcement's queue-worker example:

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: queue-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: queue-worker
  minReplicas: 0
  maxReplicas: 10
  metrics:
    - type: External
      external:
        metric:
          name: queue_consumer_lag
        target:
          type: Value
          value: "30"

One replica per 30 queued tasks, zero when the queue is empty, capped at ten. Start the Deployment at one or more replicas: a workload you scale to zero by hand is a pause, and the HPA will not wake a workload it did not scale down itself.

Two guardrails before you apply this anywhere real. First, during a version-skewed control-plane upgrade, wait until both the API server and the controller manager are on 1.37 with the gate enabled — an old controller treats replicas: 0 as a manual pause and may leave your workload at zero. Second, the metrics pipeline must exist before the HPA does: verify the adapter serves your series (kubectl get --raw against the external metrics API) first, because an HPA whose metric is unavailable reports ScalingActive=False and cannot scale from zero at all.

Why the managed columns look frozen​

Each vendor's idle posture is a pricing-architecture decision, not an oversight, and none of them moved this month.

Fly.io is the only managed PaaS of the four with true scale-to-zero: auto_stop_machines = "stop" with min_machines_running = 0 parks a Machine when traffic leaves and wakes it on the next request, with cold starts in the 300ms–2s band. A 256 MB shared machine lists around $1.94–2.02/mo always-on depending on region, bills by the second, and costs essentially nothing stopped — which is exactly how the table's twenty-cent row is computed. The asterisk is that organizations created after October 2024 get no ongoing free allowance, so those cents are real money from day one.

Render splits the difference by tier. The $0 free service spins down after 15 minutes idle but wakes in 30–60 seconds — fine for a demo, miserable for a daily driver. Every paid tier, starting at the $7 Starter (0.5 CPU / 512 MB), stays warm and bills the flat rate whether traffic exists or not. There is no dial between "free and sleepy" and "paid and always on."

Railway is the most predictable and the least idle-aware: the $5/mo Hobby floor plus metered compute ($20/vCPU, $10/GB-month) that accrues for every second a service exists. A small service hovers near $5 while quiet and climbs toward ~$20 always-on at half a vCPU — but nothing in the meter knows what "idle" means, so the duty-cycle rows in the table barely move.

If you want the full invoice-level audit behind these rows — egress tripwires, database moats, and the Hetzner fourth column — our September pricing audit re-ran every verdict against September list prices. This post's contribution is narrower: the self-hosted column just gained a native primitive, and the managed columns' shape explains why that matters.

The honest gap: HTTP doesn't wake itself​

Here is the sentence that keeps the table honest: Kubernetes Services do not buffer requests while no pods are ready. The announcement says it plainly, and it draws the feature's boundary in one stroke. A queue consumer scaled to zero wakes when the queue metric moves, because the queue holds the work. An HTTP service scaled to zero has nowhere to park an inbound request while the HPA observes the metric, schedules a pod, pulls the image, and starts the app. Request-driven workloads need a separate buffering layer — a KEDA HTTP add-on, a Knative-style activator, or a gateway that holds connections — before minReplicas: 0 is safe for them.

That makes the adoption map refreshingly clear. Adopt now for anything with a durable queue in front of it: background workers, batch processors, cron-driven jobs, staging environments that only need to exist during the workday, and GPU inference workers fed by a queue — the announcement explicitly calls out expensive reserved resources as where the savings land hardest, and a GPU node reclaimed for 20 hours a day dwarfs every CPU row in the table above. Wait — or budget a buffering component — for latency-sensitive HTTP you serve directly off a Service.

Either way, define the wake-latency SLO before you promise anything. The cold-start chain has four links (metric observation on the HPA sync loop, pod scheduling, image pull, application start), and only the last two behave like the Fly.io or Render cold starts you may have measured before. Size the stabilization window and maxReplicas against the burst you actually get: a queue that refills in seconds wants a short scale-up reaction and a generous cap, while a nightly batch wants the opposite. "Idle costs zero" is a billing statement; "wakes in under X seconds at the 99th percentile" is the operability statement that makes it a promise instead of a footgun.

What to do Monday​

If you run queue consumers, batch jobs, or daytime-only staging on a 1.37-capable fleet, this is a small, reversible experiment: pick one Deployment with an existing Prometheus-series signal, expose it through an adapter, verify it with a raw metrics read, and set minReplicas: 0 with the default five-minute stabilization. Watch one full idle-to-burst cycle — including the ScaledToZero condition transitions — before rolling it wider. If your fleet is HTTP-heavy behind plain Services, the correct Monday move is a buffering-layer spike, not a replica-count edit: the beta gives you the scaler, but the request still needs somewhere to wait.

The strategic read is that Kubernetes just closed the last idle-pricing gap that used to require leaving the platform. Fly.io earned its scale-to-zero reputation by building wake-on-request into the product; the HPA beta puts the queue-driven half of that story into every 1.37 cluster by default, and the HTTP half is now a single well-understood component away rather than a second platform to adopt. Teams pricing the self-hosted row no longer have to add a standing idle tax per service — they have to design one wake path per workload class, which is a better problem to have.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex