Sixty-six percent of organizations hosting generative AI models run some or all of their inference on Kubernetes. Only seven percent deploy models daily.
Both numbers come from the same survey — the CNCF Annual Cloud Native Survey, 2025 edition, published in January 2026 — and the gap between them is the most important number in platform engineering this year. Kubernetes has clearly won the infrastructure argument: 82% of container users run it in production, and it is now the default place a model goes to serve traffic. But winning the infrastructure argument is not the same as winning the operating-model argument. Nearly half of organizations (47%) deploy models only occasionally, and 44% still don't run AI/ML workloads on Kubernetes at all. Running a model and shipping a model are two different bars, and most teams have cleared only the first one.
This post turns that gap into a narrow, concrete golden path for a self-hosted, git-push PaaS. Not a GPU node pool with a marketing page — four mechanisms that together make model deploys as repeatable as git push:
| # | Capability | Concrete mechanism | What it replaces |
|---|---|---|---|
| 1 | Versioned code, model, and configuration | Weights, tokenizer, and serving config shipped as one versioned artifact (e.g. OCI), deployed and rolled back like app code | Hand-copied weights on a PVC nobody can reproduce |
| 2 | Explicit accelerator requests | DRA ResourceClaim/DeviceClass describing the device a workload needs, not an opaque nvidia.com/gpu: 1 count | Integer-count scheduling that can't express model, memory, or topology |
| 3 | Evaluation before promotion | KServe canary/shadow rollout with eval gates between candidate and production traffic | Pushing new weights straight at 100% of users |
| 4 | Correlated deployment, queue, GPU, and inference telemetry | Deploy events joined with queue depth, GPU counters, and per-request inference metrics (vLLM /metrics, Gateway API Inference Extension signals) | Four dashboards that never agree during an incident |
If your platform has all four, it can call itself AI-ready. If it has a GPU node pool and ambitions, it has cleared the 66% bar — congratulations, you're average. The rest of this post is about the 7% bar, and why each row of that table is load-bearing.
Running a model is not shipping a model
Start with what the 66% actually proves. Getting inference running on Kubernetes in 2026 is genuinely easy: the NVIDIA device plugin (or the newer DRA driver), a vLLM or Ollama container, a Service in front of it, and prompts get completions. The CNCF's own commentary calls Kubernetes the de facto operating system for AI, and on that evidence, fair enough.
But notice what that setup can't answer. Which exact weights are serving production right now, and can you get back to last Tuesday's set in one command? What happens to the request queue when the new checkpoint loads — does the scheduler even know a model swap is happening? Did the new model get better or worse before it took live traffic, and by whose measurement? When p99 latency spikes at 3am, can you see the deploy event, the queue depth, the GPU memory pressure, and the per-request timings on one timeline, or are you stitching four tools together while users wait?
Each unanswered question is a reason teams deploy "occasionally" instead of daily. Daily deployment isn't a virtue teams lack; it's an outcome a platform earns by making every one of those answers cheap. The 7% aren't braver. They're better tooled — the survey's own report notes that the teams with full Kubernetes adoption for inference are the ones that implemented GitOps workflows for model deployment and folded model metrics into their existing Prometheus and Grafana stacks. Tooling first, cadence follows.
That reframes the golden path correctly: it is not "deploy daily," it is "be able to deploy daily without fear." Four mechanisms, each removing one category of fear.
Mechanism 1: version everything, including the weights
Application developers solved this a decade ago: the deployable unit is a versioned, immutable artifact, and rollback means pointing at the previous version. Model serving on most clusters still lives in the pre-artifact era — weights rsynced onto a shared volume, a config edited in place, a container tag of latest doing load-bearing work.
The fix is to treat model weights, tokenizer files, and serving configuration as one versioned artifact with the same lifecycle as app code. In practice that means packaging the model as an OCI artifact: content-addressed, signed, pulled through the same registry and supply chain as container images, pinned by digest in the serving manifest. A promotion is a manifest change pointing at a new digest; a rollback is a manifest change pointing at the old one. The previous model stays available until the new one proves itself, exactly like a blue-green app deploy.
This is also where GitOps earns its keep for AI workloads. When the desired model version lives in git next to the app revision, "which weights are in production" has exactly one answer, and every promotion leaves an audit trail. The platform primitive a git-push PaaS already owns — push code, get a running service — extends naturally: push a model revision, get a serving endpoint. Same shape, heavier artifact.
Mechanism 2: ask for accelerators explicitly
The legacy GPU contract on Kubernetes is an integer: nvidia.com/gpu: 1. One what? Which model, how much memory, which NVLink domain, what MIG partition? The scheduler can't see any of it, so bin-packing inference workloads onto shared GPU nodes is guesswork, and "it fit on the staging node" is not a promise it fits in production.
Dynamic Resource Allocation, GA since Kubernetes 1.34, replaces the integer with a structured claim. The NVIDIA DRA driver publishes ResourceSlice objects describing every GPU's real attributes — model, memory, topology, partitioning capability — and workloads request what they need through ResourceClaim and DeviceClass objects with CEL selectors over those attributes. The scheduler reasons about devices the way it already reasons about CPU and memory: as resources with properties and constraints, not as counts.
For a PaaS, DRA matters twice. First, it makes multi-tenant GPU pools schedulable at all: fractional sharing, topology-aware placement, and driver-level isolation become expressible instead of hand-rolled. Second, it future-proofs the golden path — KServe has an open proposal for first-class DRA support, NVIDIA donated its DRA GPU driver to the CNCF at KubeCon Europe 2026, and the Kubernetes AI Conformance program (launched April 2026 as a testable superset of standard conformance) effectively requires DRA-style scheduling for any cluster claiming to be AI-capable. Building the accelerator story on device-plugin integers in 2026 is building on the API the ecosystem is walking away from.
The honest caveat: DRA is GA as an API, not as an operator experience. Moving a GPU workload off device plugins means new scheduler config, claim CRDs, driver installation, and in NVIDIA's case a minimum stack (recent Kubernetes, CDI-enabled runtime, new driver) plus known operational limitations. Adopt it as the declared direction, migrate node pools deliberately, and don't pretend "GA" means "turn it on Friday afternoon."
Mechanism 3: evaluate before you promote
Here is the deepest difference between shipping code and shipping models: a code change is a behavior change you wrote and tested; a model change is a behavior change you discovered. New weights can be strictly better on every benchmark you track and still worse on the traffic pattern that pays your bills. Pushing them straight to 100% of users is roulette, and teams that know it deploy occasionally — not from laziness, but from correctly pricing the risk.
The golden path prices the risk down with two patterns, both available in stock KServe today. Canary rollout shifts a fraction of live traffic (10%, then 50%, then 100%) to the new model version while the stable predictor keeps serving the rest, with canaryTrafficPercent controlling the split. Shadow deployment goes further: the candidate model receives a mirrored copy of production traffic and serves no user, existing purely to be measured. Either way the promotion decision is gated on evaluation — accuracy or task metrics, latency distributions, error rates — computed against real production-shaped traffic, not a static eval set run once in a notebook.
Note what evaluation-before-promotion assumes: mechanism 1. You cannot canary between two model versions unless both versions exist as addressable artifacts you can run side by side and roll back between. The golden path composes — versioning is the foundation eval gates stand on, which is why "occasional deploys with hand-copied weights" can't be fixed by adding a canary alone. Order matters, and versioning is first for a reason.
Mechanism 4: one timeline for deploys, queues, GPUs, and requests
When inference degrades in production, the guilty signal could live in four places: a deploy just changed the model, the request queue just deepened, the GPU just ran out of KV-cache headroom, or the request mix just shifted to longer prompts. A platform that shows these on four unrelated dashboards — deploy history over here, vLLM queue metrics over there, DCGM GPU counters somewhere else, request logs in a fourth tool — turns every incident into archaeology. Teams that have lived through that archaeology deploy occasionally, because each deploy multiplies the timelines they might have to correlate by hand.
The golden path joins them. The building blocks all exist: vLLM exposes queue depth, running requests, KV-cache utilization, and per-request timings on a Prometheus /metrics endpoint; the Gateway API Inference Extension (with llm-d's Endpoint Picker, a CNCF Sandbox project since March 2026) adds routing-level signals like prefix-cache-aware scheduling decisions and per-pool queue state; GPU counters come from the standard exporter stack; and deploy events come from the GitOps pipeline that mechanism 1 established. The platform's job is to correlate them onto one timeline keyed by model revision — this latency spike started with that promotion, on these GPUs, at this queue depth — so the incident review begins with the answer instead of the search.
This is also the mechanism that pays for the other three. Canary analysis is just correlated telemetry with a decision attached: same dashboard, comparing revisions instead of debugging one. Once the join exists, eval-gated promotion stops being a separate system and becomes a view.
Daily is a capability, not a quota
A necessary corrective before the conclusion: 47% of organizations deploy models occasionally, and for many of them that is exactly right. A fraud model retrained monthly, a support chatbot pinned to a quarterly checkpoint, an embedding model that changes twice a year — none of these teams are failing. Cadence should follow the cost of being wrong and the value of being current, and "occasionally, safely" beats "daily, fearfully" every time.
So read the 7% figure as a capability indicator, not a quota. The question the golden path answers is not "do you deploy daily" but "could you deploy today, safely, if the new checkpoint demanded it — and roll back by lunch if it misbehaved?" Teams with the four mechanisms can; teams without them schedule a maintenance window and hope.
For a small platform team, the adoption order that respects limited headcount is the table order: version artifacts first (it unlocks everything downstream), explicit accelerator requests second (it's the 2026 API migration anyway), eval-gated promotion third (start with canary percentages, add shadow later), and the correlated telemetry join last (it compounds the value of the first three). Each step is independently useful; none requires the next to pay off.
What "AI-ready" means for a git-push PaaS
Here is the narrowing the title promised. A git-push PaaS already owns the shape of the answer: the developer pushes, the platform builds a versioned artifact, the platform rolls it out with a safe strategy and observable behavior. The AI-ready bar is that same shape applied to models — the deploy primitive treats weights like code, the scheduler understands accelerators, promotions pass through eval gates, and the telemetry join covers inference signals natively. A GPU node pool with kubectl apply instructions is infrastructure; the four mechanisms are the product.
The ecosystem is converging on this bar faster than most PaaS roadmaps assume. The AI Conformance program gives "AI-capable cluster" a testable definition. Inference-aware routing (Gateway API Inference Extension, llm-d, Envoy AI Gateway) is moving request routing from round-robin to cache- and queue-aware within a single year. DRA turned GPU scheduling from counting into reasoning. Each of these started 2026 as early-stage and will end it as expected — which means the window where "we have GPUs" counts as an AI story is closing.
The 66% built the runway. The 7% fly daily. The golden path is what turns one into the other — and it's four mechanisms wide, not forty.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. The same git-push primitive this post argues models need is the one bex already gives your apps. Star the repo on GitHub or deploy your first app today.



