Skip to main content

Kubernetes 1.36 Schedules Whole Workloads, Not Pods: Bin-Packing Tenant Apps on Fixed Hardware

10 min readDora NodaDora Noda
Share
On this page

Every Cluster API fleet on owned hardware runs two schedulers and pretends it runs one. The first is kube-scheduler, placing pods one at a time, each decision blind to its siblings. The second is the descheduler, evicting pods after the fact to repair the spread and co-location mistakes the first one made.

On a cloud cluster the autoscaler papers over this place-then-repair loop with fresh nodes. On three owned Hetzner machines there are no fresh nodes, so the loop burns the only currency a fixed fleet has: headroom.

Kubernetes 1.36, Haru, released April 22, 2026, attacks the root cause with its workload-aware scheduling suite: a PodGroup scheduling cycle that evaluates a whole group against one snapshot of cluster state and binds it atomically, group-level topology constraints, workload-aware preemption, and first-phase Job controller integration. The promise is scheduling decisions that see the whole workload. The honest question for a self-hosted PaaS is narrower: what does any of that change about bin-packing tenants' ordinary multi-pod services onto nodes you already paid for?

Short answer, up front: for Deployments, almost nothing yet in 1.36 — and the reasons are specific, documented, and worth knowing before you plan around this track. For batch Jobs, something real. Here is the worked math.

The fixed pool, before and after, with numbers​

Take a concrete small fleet: three identical nodes, 8 CPU each, 24 CPU total. Two tenants run long-lived services. Tenant A has an API Deployment, 6 replicas at 1 CPU each, with a topologySpreadConstraints maxSkew of 1 across nodes (wants 2/2/2). Tenant B has a worker Deployment, 3 replicas at 2 CPU each, with preferred pod anti-affinity.

One nightly indexed Job runs 4 pods at 1 CPU each, and system reservations take about 1 CPU per node. Steady state allocates roughly 15 of 24 CPUs, leaving 9 free — until a rolling update or a node reboot lands pods lopsided and the descheduler starts evicting skew violators to fix it.

That evict-and-retry churn is the "corrected after the fact" tax. It scales with fragmentation at schedule time, which is why one number would lie — so here is the range, modeled on the pool above with stated assumptions (cached images, default descheduler cadence, no HPA stampede):

Fragmentation at schedule timeWhat happens (status quo)Wasted capacity per incidentReschedule window
Low (pool under 60% allocated)1 skew violator evicted, lands immediatelyabout 0.01 CPU-hours, one replica flapunder a minute
Medium (60–80% allocated)3 evictions, one pod misses twice before fittingabout 0.1 CPU-hours, degraded spread for minutes2–5 minutes
High (over 80% allocated)5 or more evictions colliding with a rollout; one 2-CPU worker finds no single node with room and sits pending0.5-plus CPU-hours, one tenant visibly degraded10 minutes or a page

Two things matter in that table. First, the absolute CPU-hours look small because the pool is small — but on owned hardware the cost was never metered CPU, it is headroom: every incident like the "high" row is the pool telling you the next tenant does not fit without another machine. Second, the driver is pool fullness, and a PaaS packing tenants tightly to keep per-tenant unit costs down lives in the medium row by design and visits the high row weekly.

Now the honest 1.36 "after" table, split the way the release forces you to split it:

WorkloadWhat 1.36 changesNumbers
Tenant Deployments (A and B)Approximately nothing automatic. No controller stamps PodGroups for Deployments, and hand-wired groups lose the placement guarantee the moment they carry inter-pod affinity or spread constraints — which these services do. The descheduler stays.Same churn as the "before" table. Zero reclaimed.
Nightly indexed Job (parallelism 4, completions 4)With the WorkloadWithJob gate, the Job controller auto-creates the Workload and PodGroup and the gang binds atomically: no more two pods running-but-blocked while two sit pending.In the high-fragmentation case, roughly 2 CPU times a 20-minute blocked window — about 0.7 CPU-hours — reclaimed per bad night instead of burned on rendezvous retries.

That split — services unchanged, Jobs improved — is the whole post in one table. Everything below explains why the line falls exactly there, and what sits on each side of it.

What 1.36 actually shipped, as scheduling decisions​

The upstream May 13 deep-dive frames the release as five mechanisms. Each is easier to evaluate as the placement decision it makes that pod-at-a-time scheduling cannot:

Mechanism (KEP, gate)The decision it addsWhy pod-at-a-time cannot do it
PodGroup scheduling cycle (KEP-4671, GangScheduling)Evaluate every member against one snapshot; bind all or bind none, with already-bound pods never unassigned by a later cycleSequential binding commits pod N before knowing pod N+1 fits — the half-placed gang
Group topology constraints (KEP-5732, TopologyAwareWorkloadScheduling)Co-locate a whole group inside a declared domain (for example topology.kubernetes.io/rack) via generated, evaluated, scored candidate placementsPer-pod affinity sees siblings already bound, not siblings still queued
Workload-aware preemption (KEP-5710, WorkloadAwarePreemption)Preempt as one unit across the whole cluster to make room for a group, with group priority overriding pod priority and disruptionMode choosing all-or-nothing versus independent victimhoodDefault preemption evicts per node and can tear one pod out of a running gang
DRA claims per PodGroup (KEP-5729, DRAWorkloadResourceClaims)One ResourceClaim generated for the entire group; one PodGroup reference in status.reservedFor replaces up to hundreds of pod entries past the 256-item capPer-pod claims cannot express "these pods share this device"
Job controller integration (KEP-5547, WorkloadWithJob)Auto-create and garbage-collect the Workload plus PodGroup for qualifying JobsWithout it, every gang needs hand-stamped objects and manual schedulingGroup wiring

The architectural move underneath is the API split: scheduling.k8s.io/v1alpha2 replaces v1alpha1 wholesale, Workload becomes a static template of podGroupTemplates, PodGroup becomes the runtime object the scheduler reads directly (with per-replica-sharded status), and pods link via schedulingGroup instead of workloadRef. All five mechanisms sit behind alpha gates, with GenericWorkload on apiserver and scheduler as the prerequisite.

Qualifying for the Job integration is deliberately narrow: parallelism above 1, completionMode: Indexed, completions equal to parallelism, and no schedulingGroup already set. That describes exactly the nightly Job in the worked scenario — and almost nothing a git-push PaaS runs for tenants during the day.

Why tenant services still mostly sit out​

Each limit below is documented upstream, and each has a concrete consequence for the three-node pool:

Only the Job controller stamps groups. Deployments, StatefulSets, and everything else get no integration in 1.36; elastic Jobs and other controllers are explicitly future work under KEP-5547. Consequence: tenant A's API replicas never become a PodGroup unless the platform hand-wires one per rollout — toil no small team should accept for an alpha API that already broke compatibility once (v1alpha1 to v1alpha2 in a single release).

The placement guarantee excludes the constraints services actually use. Upstream states it plainly: the cycle is expected to find a placement only for homogeneous groups without inter-pod affinity, anti-affinity, or topology-spread constraints. Heterogeneous groups and dependency-carrying groups may fail to place even when capacity exists, and intra-group dependencies can defeat the deterministic processing order regardless of cluster state. Consequence: the two mechanisms tenant services would need atomicity for — co-location affinity and spread — are precisely the ones that void the warranty. A hand-wired PodGroup around tenant B's anti-affinity workers is the worst of both worlds: new API surface, no guarantee.

Topology-aware scheduling does not preempt in 1.36. The placement engine picks the best feasible domain but will not evict anything to satisfy a group constraint; preemption integration is slated for a later release. Consequence: in the "high" fragmentation row, where a rack-pinned group most needs muscle, the feature declines to use any.

Group priority and disruption modes are ignored by default preemption. priority and disruptionMode are respected only inside the workload-aware preemption path, not by the pod-by-pod preemption that still governs everything else. Consequence: until the whole fleet's preemption flows through the group path, a latency-sensitive scale-up can still orphan half a gang exactly as before.

It is all alpha, and the alpha already churned. Five gates, a replaced API version, renamed pod fields. Consequence for a CAPI fleet: every one of those flags has to ride your control-plane machine templates and scheduler configuration through upgrades, and any pilot manifests get rewritten at least once more before Beta.

None of this is a criticism of the direction — atomic group placement is genuinely the fix for the place-then-repair loop. It is a statement about where 1.36 sits on the road: the machinery for batch gangs is real, while the services story is scaffolding with the tenant-facing floors still missing.

Against the alternatives​

Three layers compete to own "which pods land where," and they own different verbs. The status quo — pod affinity plus topologySpreadConstraints plus the descheduler — decides placement approximately now and repairs it later; its cost is the churn table from section one. Kueue and Volcano decide whether and when a workload is admitted, with quota, fair-share, and queueing the in-tree work deliberately does not replicate; as one GPU scheduling survey puts it, Kueue governs the quota question while atomic placement needs something underneath it, which is why the two are frequently deployed together. The native track decides how an admitted group's pods bind — all or nothing, in one snapshot.

That split dictates the verdict, tied back to the scenario numbers. If your pain is the churn table — services landing lopsided and getting repaired — no 1.36 gate reduces it; keep the descheduler and save your budget for the release where controller integration reaches Deployments. If your pain is half-placed batch gangs burning headroom on fixed nodes, the WorkloadWithJob pilot is cheap and the 0.7-CPU-hour bad nights go away.

If your pain is fair sharing across tenants — who gets the GPUs this afternoon — you still want Kueue or Volcano in front; Google's 130,000-node GKE experiment showed Kueue preempting and re-admitting batch at a speed kube-scheduler's per-pod path cannot match, and nothing in 1.36 closes that gap. And if you are tempted to build a custom scheduler for the services case: the upstream limitations section is your spec review — heterogeneous shapes plus inter-pod dependencies with guaranteed placement is the problem SIG Scheduling itself has not solved yet.

What a small fleet should actually do​

Stay on the upgrade path that gets you to 1.36 or later, but treat workload-aware scheduling as tracked, not adopted: enable nothing in production, pilot WorkloadWithJob on staging with a homogeneous indexed Job shaped like the nightly batch in this post, and keep the descheduler as the services answer until a release integrates the controllers tenants actually use. When the v1.37 Beta notes land, re-read the limitations section first — the day heterogeneous groups with affinity constraints get a placement guarantee is the day this story starts being about tenant apps instead of batch jobs.

The deeper lesson survives whatever the next release renames: on owned hardware, scheduling quality is headroom, and headroom is the next machine you do not have to buy. Measure the churn table for your own pool before adopting anything — then adopt the layer that moves your worst row.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex