A queued batch job asks for 8 CPUs and 32 GB of RAM. The node pool has 4 CPUs and 16 GB free. Until recently, Kubernetes gave the queue controller exactly two bad options: leave the job waiting until a big enough hole opens up, or delete the job and recreate it with a smaller ask — losing its metadata, status, and history in the process. Kubernetes 1.36 adds the obvious third option: edit the queued job's resource request in place, then admit it now.
That is what MutablePodResourcesForSuspendedJobs does. Introduced as alpha in 1.35 and promoted to beta — enabled by default — in the 1.36 "Haru" release that shipped April 22, 2026, it lets a controller or an operator change CPU, memory, GPU, and extended-resource requests and limits on a Job's pod template while the Job is suspended, then unsuspend it so the new pods are created with the adjusted shape. No new API types, no replacement objects: the API server simply relaxes the old immutability constraint on four fields. If you run bursty, queue-shaped workloads — AI-agent sandboxes, training jobs, ETL — on nodes you own, this is the primitive that lets a queue fit the work to the capacity instead of the other way around.
The one-patch flow: suspend, resize, resume
Here is the whole feature in three commands. A sandbox job was submitted asking for the full developer-workstation shape, but the pool is half-drained by the afternoon agent rush:
kubectl patch job agent-sandbox-7f3c -p '{"spec":{"suspend":true}}'
kubectl patch job agent-sandbox-7f3c --type strategic -p '{"spec":{"template":{"spec":{"containers":[{"name":"sandbox","resources":{"requests":{"cpu":"4","memory":"16Gi"},"limits":{"cpu":"4","memory":"16Gi"}}}]}}}}'
kubectl patch job agent-sandbox-7f3c -p '{"spec":{"suspend":false}}'Suspend the Job, patch the pod template's resources down to what is actually free, unsuspend it. The Job controller then creates pods with the new shape. The job object keeps its name, labels, annotations, creation timestamp, and status history — everything a delete-and-recreate would have thrown away.
The canonical example from the upstream 1.36 announcement is a training job: submitted asking for 4 GPUs, resized to 2 while suspended because only 2 are free, then resumed to make progress on half the accelerators rather than blocking the queue. The same pattern maps directly onto agent-sandbox fleets, where each sandbox is short-lived and the queue depth swings wildly with user activity. A controller watching free capacity can shrink queued sandboxes to fit the holes instead of holding them until a full-sized hole appears.
On any cluster running 1.36 or later this works with no configuration — the feature gate is on by default. On 1.35 it is alpha and needs MutablePodResourcesForSuspendedJobs enabled on the API server.
What exactly is mutable (and the guardrails)
The mutable surface is deliberately small. The API server permits changes to exactly four field families, and only while the Job is suspended:
| Field | Mutable while suspended? |
|---|---|
containers[*].resources.requests | Yes |
containers[*].resources.limits | Yes |
initContainers[*].resources.requests | Yes |
initContainers[*].resources.limits | Yes |
| Image, command, env, volumes | No |
| Scheduling directives (nodeSelector, affinity, tolerations) | Yes, but under a separate gate |
DRA resourceClaimTemplates | No — still immutable |
Two hard conditions gate every mutation. First, the Job must have spec.suspend: true — this is strictly a queue-time primitive, not a way to resize running pods (that is the separate in-place pod resize feature, more below). Second, if the Job was previously running and got suspended, all of its active pods must have fully terminated first: the API server rejects the patch while status.active is greater than zero, so a running pod can never disagree with the template it came from.
Standard validation still applies to whatever you write: limits must be greater than or equal to requests, and extended resources such as GPUs must be whole numbers where the resource type requires it. Two practical notes from the upstream guidance are worth keeping. If the Job may have failed pods around, set podReplacementPolicy: Failed so replacements only start after the previous pods terminate, avoiding resource contention from overlapping pods. And if the workload uses Dynamic Resource Allocation, its resourceClaimTemplates stay immutable — DRA-backed GPU claims must still be recreated to change, which bounds how far this feature reaches into the newest GPU-scheduling stack.
The sibling worth knowing about is MutableSchedulingDirectivesForSuspendedJobs, which graduated alongside it and is also on by default in 1.36: it allows mutating placement directives — node selectors, affinity, tolerations — on a suspended Job. Resources answer "how big," scheduling directives answer "where," and a queue controller doing serious bin-packing wants both knobs.
The bin-packing payoff for bursty sandbox queues
To see why this matters on owned hardware, compare the old decision tree with the new one. Before 1.36, a queue controller holding a suspended job that no longer fits free capacity had three moves, all costly:
- Wait. Keep the job suspended until enough capacity frees up. On a bursty sandbox pool this means queue latency spikes exactly when demand is highest, and freed capacity often arrives in fragments no full-sized job can use.
- Delete and recreate. Rewrite the job with a smaller ask as a brand-new object. This works, but the new object loses the original's metadata, status conditions, and event history — the audit trail that tells you why a sandbox ran small — and every controller watching the old object must re-resolve the new one.
- Over-provision. Keep permanent headroom so full-sized jobs almost always fit. On rented cloud instances this is a line item; on bare metal you bought, it is machines you paid for sitting idle to absorb bursts.
The 1.36 primitive replaces all three with "shrink the queued job to the hole." Walk through a concrete pool. Say a two-node sandbox pool has 4 CPUs and 16 GB free on one node and the next queued sandbox asks for 8 CPUs and 32 GB — the full workstation template. The old controller waits or recreates. The new controller patches the suspended job to 4 CPUs and 16 GB and unsuspends it immediately. The sandbox starts now, on hardware that would otherwise have idled, and the job object records exactly what happened.
Multiply that by a queue. Sandbox demand is spiky — agent frameworks fan out dozens of sandboxes per task, then go quiet — so a pool sized for the average sees constant fragmentation: many small free fragments, few big holes. A controller that can only admit fixed-size jobs either leaves fragments idle or over-provisions to make fragmentation rare. A controller that can right-size each queued job to the largest currently-free fragment converts idle fragments into admitted work. That is bin-packing in the literal sense: the item changes size to fit the bin.
This is also the feature Kueue-shaped controllers were waiting for. Kueue, the reference Kubernetes-native queueing layer, already suspends jobs behind ClusterQueues and LocalQueues, admits them when quota and capacity allow, and supports cohort borrowing, preemption, and gang admission for multi-pod jobs. Until now its admission decision was binary — the job fits as specified, or it waits. Mutable suspended-job resources turn that into a spectrum: admit now at a reduced shape, with the reduction recorded on the job itself. The ecosystem is already extending the pattern: JobSet, the multi-pod-job API, has KEP-1150 tracking the same mutability for JobSets, which matters the moment sandboxes graduate from single pods to pod groups.
The honest limits
A resize primitive is not a free-lunch primitive. Four limits should shape how you use it.
Smaller is not free. A sandbox admitted at half its requested CPUs still runs its workload — just slower, and with less headroom before OOM or CPU throttle. Right-sizing trades queue latency for execution slowdown, and the trade only wins when the job's work degrades gracefully. CPU-bound agent loops mostly do; memory-threshold workloads (a model that needs 24 GB resident or it does not load) do not. A controller should know which kind it is shrinking, or at minimum record the reduction so the slowdown is explainable.
Previously-running jobs must drain fully. The status.active == 0 rule means suspend-then-shrink is not instantaneous for a job that already started: every running pod must terminate before the patch is accepted. For short-lived sandboxes that drain in seconds this is a non-issue; for a long training job it means the resize path is "stop the world, then resize," which is correct but not cheap.
DRA claims are outside the boundary. Workloads scheduled through Dynamic Resource Allocation — the structured GPU/device API that went GA across the 1.34–1.36 window — keep immutable resourceClaimTemplates. If your GPU sandbox fleet has moved to DRA claims, changing the device ask still means recreating the claim. The classic nvidia.com/gpu resource-count path is fully covered; the newest device API is not yet.
Multi-pod jobs are still catching up. As noted, the JobSet equivalent is a tracked KEP, not a shipped feature. Single-Job, single-pod-shape queues get the full benefit today; gang-scheduled pod groups still wait or recreate.
None of these is an argument against the feature — they are the map of where it applies. Fixed-shape, single-pod, queue-held batch work: full win. DRA-device-shaped or gang-scheduled work: partial, improving.
What to do on a self-hosted fleet
Adoption is refreshingly undramatic. If your nodes already run Kubernetes 1.36, the gate is on and there is nothing to enable — verify with a suspended test job and the three-command flow above. If you are on 1.35 and want it early, enable MutablePodResourcesForSuspendedJobs (and its scheduling-directives sibling) on the API server. Both gates are beta-on-by-default in 1.36, so the upgrade path is the adoption path.
The larger decision is the queue layer that wields it. Raw suspend plus kubectl patch proves the mechanic, but production bin-packing wants a controller making admit-now-smaller versus wait-for-full decisions against live capacity and per-tenant quotas. That is Kueue's job description: quotas, admission order, preemption, and borrowing across tenants, with the job's suspended state as its control surface. Pair it with the complementary runtime primitive — in-place pod vertical resize, GA for containers in 1.35 with pod-level resize in beta in 1.36 — and the fleet gets both halves of elasticity: reshape work before it starts, reshape it while it runs.
One operational habit to build early: treat every controller-initiated shrink as an auditable event. The job object survives the resize, which is precisely what makes the audit possible — annotate the job with the original ask, the admitted shape, and the capacity snapshot that motivated it. Six months from now, when someone asks why sandbox p95 slowed down the week of the big launch, that annotation is the answer.
The direction of travel is clear. Kubernetes spent a decade treating a job's resource ask as a fixed declaration made at submit time. Between mutable suspended-job resources, mutable scheduling directives, and in-place resize, the ask is becoming a live negotiation between the work and the capacity — exactly the negotiation a self-hosted fleet, with no cloud autoscaler to hide behind, has to win on its own nodes.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



