Your inference workload asked for an H100. All four H100s are busy, so the pod sits in Pending — while two A100 nodes on the same fleet sit completely idle. Nothing is broken. The scheduler did exactly what the manifest said: H100 or nothing. The manifest just had no words for "or."
Kubernetes 1.36 ("Haru," released April 22, 2026) graduated the fix to Stable: DRA prioritized lists let a ResourceClaim declare an ordered cascade of acceptable devices instead of one exact match. The pod above gets twenty lines of YAML and stops pending:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: inference-gpu
spec:
spec:
devices:
requests:
- name: gpu
firstAvailable:
- name: h100
deviceClassName: gpu-h100
- name: a100
deviceClassName: gpu-a100
- name: t4
deviceClassName: gpu-t4The scheduler tries h100 first, falls back to a100, then t4 — "give me an H100, or an A100, or a T4," expressed declaratively, in one claim, with no duplicate Deployments and no node-affinity gymnastics. The rest of this post is the mechanism behind those twenty lines, the claim-template design that makes fallback safe to promise tenants, and the four gotchas that bite if you skip the design step.
How firstAvailable actually works
Dynamic Resource Allocation lets workloads request hardware through ResourceClaim objects instead of the old opaque nvidia.com/gpu: 1 resource-count hack. Before prioritized lists, each DeviceRequest inside a claim described exactly one acceptable device: one class, one set of selectors, one count. If no device matched, the pod pended — even when a slightly different device would have run the workload fine.
KEP-4816 restructures DeviceRequest into a one-of: exactly for a single non-negotiable request, or firstAvailable for a prioritized list of subrequests. Exactly one subrequest is satisfied; the scheduler walks the list in order and uses the first entry it can place. The list caps at 8 subrequests, and a DeviceSubRequest carries everything a normal request does — class, CEL selectors, count, allocation mode — except AdminAccess.
Two properties make this cheap to adopt. First, scheduling stays entirely control-plane-local: no kubelet or DRA-driver changes are needed for the scheduler to evaluate alternatives (there is one driver-side caveat, covered in the gotchas). Second, the maturity ramp is already behind us — alpha in 1.33, beta enabled by default in 1.34, Stable in 1.36. If your fleet is on a supported 1.36, the feature gate DRAPrioritizedList is on and the API is resource.k8s.io/v1.
The subtle part is scoring. Trying alternatives "in order" answers which device a pod gets on a given node, but nodes usually carry only one GPU type — so many nodes can each satisfy different alternatives, and the scheduler needs a reason to prefer the node holding your first choice over the node holding your third. DRA implements limited preference scoring for exactly this: each request scores 8 when its first alternative wins down to 1 for the eighth, scores sum across the pod's requests, normalize to 0–100, and the plugin carries weight 2.
In practice: when both an H100 node and a T4 node are free, the H100 node wins the pod. Without that scoring (the 1.33 alpha behavior), the scheduler could place you on the T4 while H100s sat idle — first fit is not best fit.
The claim-template design for a mixed-card fleet
The YAML in the hook works, but a platform needs a design around it before fallback becomes a primitive tenants rely on. Three pieces:
One DeviceClass per GPU tier. DeviceClass is the admin-controlled vocabulary of what may be requested — its selectors decide which physical devices belong to the class. Give each card generation its own class (gpu-h100, gpu-a100, gpu-t4), optionally refined with CEL selectors on attributes your DRA driver publishes (memory size is the canonical example: one 80GB card versus two 40GB cards is a real fallback pair, as Google's GPU fungibility tutorial demonstrates with vLLM). Tenants never invent class names; they pick from the menu you installed. When a new card generation joins the fleet, you add one class — existing claim templates keep working untouched.
One fallback template per workload shape, not per tenant. Publish a ResourceClaimTemplate like inference-gpu above and let tenant pods consume it by reference:
spec:
resourceClaims:
- name: gpu
resourceClaimTemplateName: inference-gpu
containers:
- name: server
resources:
claims:
- name: gpuThe pod says "I need the inference GPU claim" and the scheduler resolves which tier actually backs it. This is the portability story from KEP-4816's second user story: a workload author ships one manifest that runs on a wide range of clusters, with preference order expressing "best here" rather than "only here."
Runtime discovery inside the container. Fallback only helps if the workload can actually run on whatever it gets. The established pattern is detecting allocated GPUs at container start — nvidia-smi -L in an entrypoint wrapper — and deriving configuration from what shows up, like vLLM's tensor-parallel-size following the visible GPU count. A claim that can resolve to one 80GB card or two 40GB cards needs a container that adapts to either; otherwise you've built a scheduler that says yes and an application that says no. Make the discovery wrapper part of the same blessed image or Helm chart as the claim template, so tenants get both or neither.
Four gotchas before you promise this to tenants
1. Quota is charged for every alternative, not just the winner. ResourceQuota enforcement requires quota headroom for each DeviceSubRequest under every firstAvailable — the KEP is explicit that the pick-one behavior must not become a quota bypass. A three-tier fallback claim consumes quota in all three classes at admission, even though only one device is ever allocated. Size tenant quotas for the menu, not the meal, and document that fallback breadth has a quota cost. Queueing systems like Kueue follow the same rule: every mentioned device class counts against quota.
2. Replicas of one Deployment can land on different GPU types. Cross-claim consistency is an explicit non-goal: nothing guarantees that all claims in a Deployment resolve to the same alternative. Replica 0 gets an H100, replica 1 gets a T4, and your p99 latency graph now has two personalities. For batch and queue-drained inference this is usually fine — arguably the point. For latency-sensitive serving behind one Service, either pin the claim to exactly, or split tiers into separate Deployments with separate claims so each pool's performance is legible.
3. Your DRA driver must parse the subrequest result format. The "no driver changes" claim has one footnote: allocation results reference the winning subrequest as <main-request>/<subrequest> (e.g. gpu/a100) in DeviceRequestAllocationResult.request, and the driver must understand that format to bind the device into the pod. Any driver updated for the 1.36 DRA API handles it — but if you run a pinned or vendored driver that predates the format, claims will allocate in the scheduler and then fail at the kubelet. Verify the driver version before advertising fallback, not after the first tenant ticket.
4. The autoscaler reads your preferences too. Cluster-autoscaler support for DRA evaluates the same ordered list when deciding what to provision — which means a pending pod with an H100-first claim can trigger scale-up of an H100 node pool even when a T4 pool has room to schedule it after fallback... or conversely, satisfy itself from existing capacity you'd rather reserve. Walk through each fallback template against your autoscaler configuration and pool priorities once, deliberately, before tenants do it accidentally at 3am.
One pool per card, one claim across them
On a Cluster-API-managed fleet, the natural topology falls out of how CAPI machine sets work: one MachineDeployment per GPU card type, each producing homogeneous nodes (same card, same driver version, same DRA ResourceSlice shape), with a single firstAvailable claim spanning the tiers. Pools stay boring and uniform — exactly what you want from machine lifecycle automation — while fungibility lives one layer up, in the claim.
This maps well to how owned-hardware fleets actually grow. Nobody buys forty identical GPUs on day one; you buy what fits the budget, then add a newer generation eighteen months later, and suddenly the fleet is two card types with different memory and different performance. Under exact-match scheduling, every generation boundary is a scheduling wall: workloads pinned to the old card can't spill onto the new one and vice versa. A prioritized claim turns the wall into a slope — new capacity absorbs overflow from old pools automatically, ordered by your preference list.
One honest calibration on card names: the H100/A100/T4 cascade in this post follows upstream's own examples, but a Hetzner-backed fleet shops from a different catalog — dedicated GPU servers there are RTX-class cards (RTX 4000 Ada, RTX PRO 6000, and similar), roughly €200/month per bare-metal box with one GPU each. The mechanism doesn't care: gpu-rtx4000 and gpu-rtxpro6000 classes with a firstAvailable spanning them behave identically to the datacenter-card version. Name your tiers after what you own, keep one pool per card, and let the claim do the spanning.
The upgrade path is gentle. Prioritized lists were beta and default-on from 1.34, so a 1.34/1.35 fleet may already carry firstAvailable claims in v1beta1 form; 1.36 graduates the feature itself to GA in the v1 API. And know when not to use it: a latency-SLO serving path that genuinely needs one card should keep exactly — fallback is for workloads where running sooner on lesser hardware beats waiting for the best hardware. That tradeoff is a per-workload call, which is precisely why it belongs in the claim, not in the scheduler's defaults.
The pending pod from the opening paragraph was never a capacity problem. The fleet had free GPUs; the API just couldn't express "any of these will do." DRA prioritized lists close that gap with twenty lines of YAML, preference-aware scoring, and an upgrade path that needs no kubelet or driver rollout. Design the classes, publish the templates, teach containers to discover what they got — and watch how many Pending pods were really just exact-match artifacts.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



