On March 24, 2026, in front of a KubeCon Europe crowd in Amsterdam, the CNCF announced it had nearly doubled the number of certified Kubernetes AI platforms — from 18 to 31 in five months — and one of the new names on the list was OVHcloud. The French bare-metal giant now ships a Managed Kubernetes Service certified against the program's stricter v1.35 rules, with validation for agentic workflows baked in. Meanwhile the other European bare-metal giant, Hetzner — the provider behind half the self-hosted fleet math on the internet — is nowhere on the list, and shows no sign of wanting to be.
That split is the most useful hardware signal a self-hosted platform team has gotten all year. Two providers that look interchangeable on a price sheet are placing opposite bets: certified AI-grade Kubernetes versus the cheapest reliable box in Europe. This post works out what each bet actually buys, with real numbers, and reduces the choice to a workload-shape rule you can apply to your own fleet.
TL;DR — the verdict up front:
- The badge is real engineering, not marketing. AI conformance certifies DRA-grade GPU scheduling, in-place pod resizing, workload-aware scheduling, and agentic-workflow validation against Kubernetes v1.35 primitives.
- The price gap is just as real. A Hetzner AX41-class box starts around €47/month; OVHcloud's entry bare metal runs roughly $80/month, and its H100 GPU time starts at €2.80/hour.
- The decision rule: if your fleet will schedule GPUs for agent sandboxes or inference this year, the certified lane is worth pricing. If it runs CPU-only git-push web services, the badge buys you nothing and the cheap box wins.
- The likely answer for a small PaaS is both: a cheap CPU default fleet plus a certified GPU lane, not one provider for everything.
What the AI-conformance badge actually certifies
The Kubernetes AI Conformance Program launched in November 2025 as a superset of standard Kubernetes conformance: same portability promise, plus testable requirements for the things AI workloads need that a generic conformant cluster never had to prove. The March 2026 update did two things at once. It grew the certified roster from 18 to 31 platforms — a surge of over 70% — adding OVHcloud, SpectroCloud, JD Cloud, and China Unicom Cloud. And it codified the new bar as Kubernetes AI Requirements (KARs): stricter v1.35-aligned rules with official validation for agentic workflows, aimed at what the CNCF calls "industrial-scale" AI deployment without infrastructure fragmentation.
Strip away the press-release language and the KARs test roughly four things a self-hosted operator should care about:
| Requirement area | What it proves | Why it matters on owned hardware |
|---|---|---|
| Structured GPU scheduling (DRA) | GPUs allocated through claim-based Dynamic Resource Allocation, not integer-count device plugins | Fractional sharing, topology-aware placement, and driver versioning across tenants sharing a few cards |
| Stable in-place pod resizing | CPU/memory envelopes resize without a pod restart | Vertical scaling for inference and agent workers without a bounce |
| Workload-aware scheduling | The scheduler places a workload's pods with sibling context | Gang-style co-location for multi-pod inference and training jobs |
| Agentic-workflow validation | End-to-end agent workload patterns pass conformance tests | Proof the platform runs sandbox-per-task agent loops, not just static deployments |
The DRA line is the load-bearing one. Dynamic Resource Allocation graduated to GA in Kubernetes v1.34 with stable resource.k8s.io APIs, and at the same KubeCon where OVHcloud got certified, NVIDIA donated its DRA GPU driver to the CNCF while Google open-sourced a TPU driver. The old nvidia.com/gpu integer-count device-plugin model still works, but it is a dead end for anything fancier than one-whole-GPU-per-pod: it cannot express fractional sharing, MIG partitions, or topology. Certification is, in practice, a promise that somebody else has already wired DRA, a conformant driver, and the scheduler integration together and proven it with tests.
What the badge does not prove matters too. It does not promise a price, a region near your users, or an SLA on any particular GPU model. It is a compatibility and capability floor, not a bill of materials. Keep that in mind when the price tables arrive — the badge and the invoice answer different questions.
Why Hetzner is still the default — and what that costs
Hetzner's claim to "default" was never a certification. It is three mundane things: the lowest per-box price in Europe from a reputable operator, a server auction with real liquidity in the €30–50/month band for 4–8-core boxes with 32–64 GB of RAM, and a decade of community gravity that made it the reference hardware behind guides, cost comparisons, and the Cluster API Provider Hetzner (CAPH) ecosystem.
When a self-hosting tutorial says "provision a node," the node it pictures is usually an AX41-class Hetzner box. After the mid-2026 repricing, entry dedicated starts around €47/month — still roughly half of what comparable single-tenant iron costs elsewhere in Europe.
OVHcloud, meanwhile, spent 2026 visibly chasing the AI buyer. Its Managed Kubernetes control plane has a free tier (with a paid Standard tier around $0.099 per cluster per hour for stricter workloads), public-instance traffic is free — no egress meter running while you serve inference responses — and its GPU catalog now spans H100 PCIe instances from €2.80/hour alongside newer serverless and GPU-equipped bare-metal offerings.
Entry bare metal starts around $80/month for an Advance-class box. That is not expensive in absolute terms. It is simply a different posture: Hetzner sells you the cheapest box that runs Kubernetes well, while OVHcloud sells you a certified platform for the workloads Kubernetes is becoming.
Put a typical small self-hosted fleet on each side and total it:
| Fleet shape | Hetzner-anchored | OVHcloud-anchored |
|---|---|---|
| 3× entry CPU nodes | ~€141/mo (3× AX41-class at ~€47) | ~$240/mo (3× Advance-class at ~$80) |
| Control plane | Self-managed on the same boxes (your ops time) | Free MKS tier (or ~$72/mo Standard at 730 hrs) |
| Occasional GPU (say 40 hrs/mo) | Monthly-only pre-configured GPU box — no hourly lane | ~€112/mo (H100 PCIe at €2.80/hr, hourly) |
| Egress for served traffic | Generous included traffic, metered beyond | Free/unlimited instance traffic |
| AI-conformance badge | None — you wire DRA yourself if you ever need it | Certified KAR platform |
Two honest caveats keep this from being a Hetzner advertisement. First, Hetzner's GPU story is monthly-only, pre-configured boxes: fine if you need one card full-time, useless if you need forty hours of H100 this month and zero next month. Per comparison pricing tracked through 2026, even Hetzner's on-demand GPU lane (where it exists, e.g. RTX 4000 SFF Ada around $0.43/GPU/hour) is a narrow catalog next to OVHcloud's hourly H100s.
Second, Hetzner's footprint is roughly half a dozen EU/US regions plus Singapore, against OVHcloud's ~40. If your users or your data-residency story need a region Hetzner does not have, no per-box saving fixes that.
Even with those caveats, the default status holds for one reason: most self-hosted fleets are still CPU-only, and for CPU-only fleets the badge certifies capabilities they will never invoke. Paying more per box to certify GPU scheduling you do not use is not future-proofing. It is a donation.
Which workload shape decides
Here is the rule stated plainly: the provider bet follows the GPU bet, and the GPU bet follows the workload shape.
Shape 1: CPU-only git-push web services and workers. Stateless apps, background queues, scheduled jobs, maybe Postgres on a dedicated box. Nothing here touches a GPU, a device claim, or a gang scheduler. DRA could not buy you anything if it were free, because there is no device to allocate. In-place resize and workload-aware scheduling are nice, but they are upstream Kubernetes features you inherit on any recent cluster — Hetzner boxes run the same kubelet as anyone else's. For this shape, Hetzner-plus-CAPH remains the rational default, and the AI-conformance badge should carry zero weight in procurement.
Shape 2: agent sandboxes and inference on shared GPUs. The moment tenants run AI-agent sandboxes with model inference — one agent per task, each needing a slice of VRAM, scheduled and torn down in minutes — integer-count GPU allocation falls over. You need fractional sharing so four sandboxes can split one card, topology-aware placement so the scheduler stops scattering gang-scheduled inference replicas across NUMA nodes, and driver-level isolation so one tenant's CUDA crash does not take the neighbor down. That is exactly the DRA claim/class model the KARs test, and it is exactly what a from-scratch Hetzner fleet does not give you: no hourly GPU lane, no certified driver stack, no validated scheduler integration. For this shape, price the certified lane — whether that is OVHcloud's MKS, another KAR platform, or your own DRA wiring with open eyes about the engineering cost.
The sensitivity case: your first GPU workload. Most fleets do not start in Shape 2; they drift into it when one tenant asks for inference. The cost of that drift is asymmetric. Adding a first GPU workload to a certified platform is a scheduling exercise. Adding it to a monthly-box fleet means either a full-time GPU server for part-time need or a second provider relationship under time pressure.
So the real question is not "are we an AI platform" but "will anyone ask for GPU time in the next twelve months." If the honest answer is yes — agent sandboxes are on your roadmap, a tenant is already asking — start the certified lane now, while it is a calm procurement decision. If the honest answer is no, bank the per-box savings and re-ask every quarter. The badge is not going anywhere; there will be more than 31 certified platforms by the time you need one.
What a self-hosted PaaS fleet should actually do
For a small platform team running its own machines, the split suggests a two-lane fleet rather than a single-provider oath:
- Keep the CPU default cheap. Run web, worker, and control-plane capacity on the lowest-cost reliable iron that your provisioning story supports — today, that is still Hetzner-class boxes under CAPH, with the auction as your capacity buffer. Revisit only if your region needs or the repricing trend breaks the math.
- Stand up a certified GPU lane when Shape 2 appears on the roadmap. Favor a KAR-certified managed Kubernetes for the GPU pool specifically, so DRA drivers, device claims, and agentic-workflow validation arrive tested instead of assembled. Hourly GPU billing matters more than the badge here — forty hours of H100 should cost forty hours, not a month.
- Treat conformance as a procurement shortcut, not a religion. The KAR list is most valuable as a shortlist filter: any vendor on it has proven the primitives, so your evaluation can skip "does DRA work here" and start at "what does an hour of VRAM cost and which regions have it." Re-run that filter yearly; the March 2026 class of 31 will look quaint by 2027.
- Track the primitives, not just the badge. DRA is GA, partitionable devices are maturing through 1.36/1.37, and the CNCF-conformant floor keeps moving with each Kubernetes release. Even on the cheap lane, keep your cluster within one minor of current so the day you need device claims, the API is already under you.
None of this requires abandoning the machines you own. Both lanes are compatible with a Cluster-API-managed fleet; the GPU lane just rents its control-plane certainty from someone who already passed the test suite. Owning the hardware and outsourcing the certification homework are not contradictory — they are how a two-person platform team gets both the invoice and the capability model it needs.
The divergence is the point
It is tempting to read OVHcloud's certification as a verdict on Hetzner, or Hetzner's prices as a verdict on the badge. Both readings miss what actually happened in Amsterdam. The CNCF did not declare uncertified iron obsolete; it published a test suite for a workload class that barely existed three years ago, and 31 vendors lined up to pass it. OVHcloud decided its future tenants run agent loops against GPU pools and bought the proof. Hetzner decided its present tenants run web services on cheap boxes and kept the price floor. Both decisions are coherent because they serve different workload shapes — and your fleet almost certainly contains exactly one of those shapes today, with the other one somewhere on the roadmap.
Place the bet that matches the shape you have, keep the door open for the shape you are growing into, and let the invoice follow the workload instead of the other way around.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.
Sources
- CNCF press release, "CNCF Nearly Doubles Certified Kubernetes AI Platforms," March 24, 2026 (18 to 31 certified; OVHcloud, SpectroCloud, JD Cloud, China Unicom Cloud; KARs; agentic workflows)
- Cloud Native Now, "CNCF Nearly Doubles Certified Kubernetes AI Platforms, Adds Agentic Workflow Validation," March 2026 (v1.35 primitives: in-place pod resizing, workload-aware scheduling)
- OVHcloud blog, "KubeCon + CloudNativeCon Europe 2026 in Amsterdam" recap (MKS certification)
- Qovery blog, "Which EU Cloud Providers Support Kubernetes? The 2026 List, Prices and Sovereignty Checklist," September 2026 (MKS free tier, ~$0.099/hr Standard, free instance traffic, v1.35 conformance)
- Techerati / OVHcloud AI expansion announcement (GPU instances from €2.80/hour for H100 PCIe)
- Atal Networks, "7 Best Dedicated Server Providers in Germany 2026," July 2026 (Hetzner entry dedicated ~€47/month post-June-2026 pricing)
- Mid-2026 bare-metal pricing comparisons (Hetzner AX41-class ~€40, auction €30–50; OVHcloud Advance ~$80; Hetzner GPU monthly-only)
- Kubernetes documentation and release notes (DRA GA in v1.34,
resource.k8s.ioAPIs; NVIDIA DRA driver donation and Google TPU driver at KubeCon EU 2026)



