On July 25, former Mesosphere co-founder Tobi Knaup published an essay arguing that open-weight AI is having its Kubernetes moment — and Hacker News agreed, pushing it past 400 points and 300 comments. His thesis, forged from losing the container-orchestration war to Kubernetes: once an open platform that people can customize becomes the industry's center of gravity, no single vendor can match the combined rate of innovation around it.
The evidence for the inflection is hard to wave away. Hugging Face now hosts more than two million public models, and Chinese models accounted for 41% of downloads over the past year. The frontier gap is narrowing fast: Z.ai's GLM-5.2 ships under an MIT license with a reported 62.1% on SWE-bench Pro versus 58.6% for GPT-5.5, and Moonshot's Kimi K3 scores alongside Opus 4.8 and GPT-5.5 in Artificial Analysis's independent evaluation. Kubernetes didn't win because its repository was public, Knaup writes — it won by becoming a neutral substrate with common interfaces and vendor-neutral governance that everyone could build on.
But here is the question the essay leaves for operators: what does "standardized" concretely mean for a self-hosted fleet? If you run a git-push PaaS on machines you own, serving an open-weight model next to tenant web apps is not just another Deployment. This post maps the standardized inference stack layer by layer, names what your platform is missing, and answers whether inference arrives as a tenant workload you schedule or a platform primitive you operate.
The standardized inference stack, layer by layer
A July 2026 survey of Kubernetes model serving frames the shift well: in 2024 the ecosystem was isolated features; by 2026 it is an explicit stack of control-plane and data-plane components. Here is that stack as a requirements checklist for a fleet, with a maturity verdict per layer.
| Concern | The 2026 standard piece | Maturity | What it replaces |
|---|---|---|---|
| Model distribution + versioning | OCI image volumes (stable in K8s 1.36); tags/digests for versions, KServe Local Model Cache for warm nodes | Stable API; caching still yours to operate | Init-container downloads from object storage per replica, KServe ModelCar sidecars |
| Accelerator allocation | Dynamic Resource Allocation: DeviceClass, ResourceClaim, ResourceClaimTemplate | Core API stable since 1.34; 1.36 adds partitionable devices (beta), health, drivers | nvidia.com/gpu: N device-plugin integer counting |
| Request routing | Gateway API Inference Extension: InferencePool + endpoint picker (EPP) | Extension GA; picker policies vary by implementation | Round-robin Service load balancing |
| Serving API | KServe LLMInferenceService (generative) vs InferenceService (predictive) | KServe 0.18; API still v1alpha2 — plan for schema churn | Engine-specific manifests, hand-rolled controllers |
| Multi-node execution | LeaderWorkerSet: a group of pods replicating as one unit | Available; test failure modes before production | Custom multi-node deployments, often tied to Ray |
| Prefill/decode split | llm-d: disaggregated prefill/decode pools, prefix-cache-aware routing, distributed KV cache | CNCF Sandbox (March 2026); v0.7 line current | Colocated prefill+decode in every replica |
| Autoscaling signals | Queue depth, tokens/sec, active sequences, KV-cache pressure via KEDA/custom metrics + llm-d workload-aware autoscaling | Workable; no single standard signal | CPU/memory HPA thresholds |
| Tenant policy | AI Gateway Working Group patterns: token rate limits, quotas, guardrails, provider failover | WG formed March 2026; proposals, not stable APIs | Product-specific gateway middleware |
Three rows deserve a closer look, because each corrects a plausible wrong assumption.
First, DRA went GA in Kubernetes 1.34 — not 1.36, which continued the maturation with device health, partitionable devices, and more drivers. The stable API does not make every accelerator stack production-ready: you still need a supported vendor driver (NVIDIA's DRA GPU driver moved to CNCF governance at KubeCon EU 2026), a compatibility story with existing device-plugin workloads, and an answer for MIG/time-slicing representation. For many clusters, device plugins remain the production default; adopt DRA when its richer selection semantics solve a concrete placement problem, like requesting two GPUs with NVLink and minimum VRAM instead of any two cards.
Second, OCI image volumes solve packaging, not cold starts. A pod can now mount weights read-only straight from a registry — no sidecar needed just to expose the filesystem — but a 100–500 GB model still has to reach the node, decompress, and load. Startup behavior depends on registry throughput, layer structure, disk capacity, and whether the model is already cached. Versioning comes free with the artifact format (immutable digests, rollback by tag), but warm capacity is an operational property: KServe's Local Model Cache pre-downloads to NVMe so replicas share a warmed copy, and autoscaling cannot create warm GPU capacity instantly no matter what the API looks like.
Third, the Gateway API Inference Extension being GA does not mean every routing policy is mature. The extension owns the InferencePool API and conformance; advanced endpoint-selection and model-rewrite logic has moved to llm-d's router. That split is exactly the Kubernetes-moment pattern — a stable narrow waist with competing implementations above it — but it means "GA" answers which API to build against, not which picker policy is production-proven for your traffic shape.
What your git-push PaaS doesn't have yet
A platform built for stateless web services runs on three comfortable assumptions: replicas are interchangeable, CPU and memory describe load, and a cold start takes seconds. Inference breaks all three, and the breakage is structural rather than tunable.
Replicas are not interchangeable. Two healthy vLLM replicas can have wildly different costs for the same request: one already holds the prompt prefix in its KV cache, one sits behind a long queue, one is near KV-cache exhaustion, or only one has the requested LoRA adapter loaded. Round-robin routing — the default behind every Kubernetes Service — ignores all of this. Prefix-cache affinity alone can eliminate seconds of re-prefill per request, which is why the endpoint picker exists as a separate component with live engine metrics rather than as a Service annotation.
CPU and memory do not describe inference load. Demand is measured in prompts, tokens, context lengths, and active sequences. A replica at 30% GPU utilization can still be saturated if its KV cache is full; a replica at 90% can still accept a short decode cheaply. Autoscaling on GPU utilization is the load-average heuristic of inference: directionally right, wrong at exactly the moments that matter. The signals that work — queue depth, prompt/output length mix, concurrency, cache pressure — come from the engine and the router, not the kubelet, which is why KEDA-style custom-metric scaling and llm-d's workload-aware autoscaling sit in the stack where HPA used to suffice.
Cold starts are measured in minutes and gigabytes. A web deploy pulls a hundred-megabyte image and passes readiness in seconds. An inference replica pulls up to half a terabyte of weights, initializes the engine, and warms caches — and scale-to-zero, the thing that makes GPU economics work for bursty tenants, makes every wake-up pay that bill. KServe's LLMInferenceService supports scaling to zero, but the platform has to decide what a cold wake-up is allowed to cost in latency before offering it, because no routing trick hides a multi-minute weight load.
Finally, placement is topological, not bin-packing. Web pods land wherever CPU and RAM fit. Inference placement cares about which GPUs share NVLink, which nodes share InfiniBand for tensor parallelism and KV transfer, and whether prefill and decode pools are sized independently. On owned hardware this is fully your problem: there is no managed node pool abstracting the interconnect away.
Tenant workload or platform primitive?
This is the decision the stack table cannot make for you. Standardized model deployment can arrive two ways: as an ordinary tenant workload the platform schedules (a tenant ships a vLLM container like any other app), or as a platform primitive the platform operates (shared model pools with platform-owned routing, caching, and policy). The right answer depends on where your fleet sits on four axes.
Start with tenant workloads when GPUs are scarce and tenants are few. If the fleet has one or two GPU nodes and a handful of tenants, a tested model-server image behind a standard Service and Gateway API route — weights from an OCI image volume or pre-warmed cache, Prometheus metrics at gateway, engine, GPU, and node layers — is the whole stack. Don't install an inference-aware router until replicas are busy enough for routing quality to matter. At this scale, inference is just a stateful-ish app with a big image and slow starts, and every controller you skip is reconciliation behavior, RBAC surface, and upgrade-compat testing you don't owe.
Promote to a platform primitive when sharing begins. The moment two tenants want the same model, or one tenant's bursty traffic needs scale-to-zero on shared GPUs, the tenant-workload model starts duplicating the expensive parts: each tenant pulls its own 500 GB copy, warms its own caches, and idles its own cards. That's when InferencePool plus a production endpoint picker earns its keep, when queue-aware autoscaling replaces per-tenant HPA, and when token rate limits and quotas move into gateway policy so one tenant's long-context job doesn't starve another's interactive session. Multi-node (LeaderWorkerSet) and disaggregated prefill/decode (llm-d) enter only when measurements show the simpler topology is the bottleneck — disaggregation on a bandwidth-constrained interconnect moves the bottleneck rather than removing it.
The adoption sequence, then, is a ladder, not a shopping list: single model on plain primitives → shared service with inference-aware routing and queue-driven autoscaling → multi-node replicas → heterogeneous pools with DRA selection and llm-d distribution. Each rung adds controllers, and each controller adds status to monitor and desired-versus-actual state that can diverge — so climb only when the current rung's SLOs say so.
For a small owned-hardware fleet, the honest default today is tenant workloads with a platform-operated on-ramp: standardize the OCI weight format, pre-warm the cache, and expose the engine metrics — then let demand pull the routing and pooling layers in. You get portability from the standards without operating a distributed inference system before you have distributed inference load.
Standards are the moat
Knaup's policy section reaches for Kubernetes conformance as a governance model, and the analogy already has teeth: CNCF's Certified Kubernetes AI Conformance Program, launched at KubeCon North America in late 2025, sets testable requirements for GPU/TPU scheduling, telemetry, and orchestration, with stricter Kubernetes AI Requirements following — stable in-place pod resizing and workload-aware scheduling are now mandatory. GKE, EKS, and AKS are certifying against it. A self-hosted fleet that builds on the same narrow waist — OCI weights, DRA claims, InferencePool, LLMInferenceService — inherits portability with every hyperscaler that certifies, which is precisely the dynamic that made betting on Kubernetes safe in 2016.
The gaps are real and worth naming plainly: predictable cold starts, safe autoscaling that doesn't destroy cache state, portable accelerator behavior across drivers, and several APIs still in alpha. Standardization is a direction, not a destination. But the direction is unmistakable — model weights are becoming the portable artifact that container images were, and the control plane around them is converging on shared Kubernetes APIs instead of per-vendor snowflakes.
That is what a Kubernetes moment actually looks like from inside: not a single release, but the quarter-by-quarter accumulation of boring, interoperable primitives until building on the open substrate is the obvious choice. For fleets on owned hardware, the move is to adopt the artifact format and the stable APIs now, operate the smallest stack that meets today's SLOs, and let the ecosystem compound around you.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



