In 2008, McKinsey measured average utilization of enterprise bare-metal servers and landed on a number that embarrassed an industry: 10%. Eighteen years, one container revolution, and several trillion dollars of cloud spend later, Kubernetes — the platform that was supposed to fix exactly this — has managed to go lower. According to Cast AI's 2026 State of Kubernetes Optimization Report, average CPU utilization across production clusters now sits at 8%, down from 10% a year earlier. We built the most sophisticated bin-packing scheduler in infrastructure history, and the bins are emptier than the racks they replaced.
This post does the thing most coverage of the report doesn't: it puts a dollar figure on that 8%. Below is a worked, line-by-line comparison of what an 8%-utilized fleet costs under metered cloud billing versus consolidated onto flat-rate dedicated hardware — plus the sensitivity analysis that keeps the comparison honest, including the ops-cost row where self-hosting stops winning.
8%, 20%, 5%: the numbers, and what they actually measure
The headline figures from the 2026 report, drawn from roughly 23,000 production clusters across AWS, Azure, and Google Cloud run by more than 2,100 organizations, are bleak in every dimension (press release, covered by DataCenterKnowledge, ITOps Times, SDxCentral, and Cloud Native Now):
- CPU utilization: 8%, down from 10% the prior year.
- Memory utilization: 20%, down from 23%.
- GPU utilization: 5% — measured for the first time, and already the worst number in the set.
- CPU overprovisioning: 69% of requested CPU goes entirely unused, up from 40% year over year. Memory overprovisioning hit 79%.
Two methodology notes matter before anyone builds a budget on these. First, the utilization window covers January through December 2025 (GPU through April 2026), so this is a trailing picture, not a live meter. Second, Cast AI sells Kubernetes optimization software, and the data was collected before the measured organizations enabled its automation — a genuine baseline of unoptimized clusters, but a baseline published by a vendor with a stake in the gap looking large. The direction of travel is corroborated independently, though: the prior year's edition found 99% of clusters overprovisioned with only 13% of provisioned CPUs actually used, and FinOps Weekly noted that the industry has been staring at the same ~10% figure since the McKinsey bare-metal days. The trend is real even if any single vendor's decimal point deserves a raised eyebrow.
The report's own summary line is the one to sit with: "The efficiency gains that Kubernetes was designed to unlock are not emerging naturally with scale, and the gap between what organizations are paying for and what they are actually using is widening." Scale was supposed to fix utilization. It is making it worse.
The worked comparison: a typical 4-node fleet at 8%
Averages only bite when they attach to a bill. So let's price a concrete, typical small fleet: 4 general-purpose cloud nodes, 8 vCPU / 32 GB each (think AWS m7i.2xlarge class at roughly $0.48/hour, or about $353/month per node), running a standard mix of web services, workers, and datastores under a managed Kubernetes control plane. Total provisioned capacity: 32 vCPU and 128 GB of RAM.
At the survey median of 8% CPU and 20% memory, that fleet's real consumption is about 2.6 vCPU and 26 GB of RAM — roughly one developer laptop's worth of compute, spread across four rented servers.
Here is the monthly math, line by line, at on-demand list prices:
| Line item | Metered cloud (4 nodes) | Flat-rate owned (2 dedicated boxes) |
|---|---|---|
| Compute | 4 × $353 = $1,412 | 2 × ~$75 (8c/64 GB class, e.g. Hetzner AX52) = $150 |
| Control plane | Managed K8s fee ≈ $73 (EKS/GKE class) | Self-managed, $0 |
| Storage | 4 × 100 GB gp3 ≈ $32 | 4 TB NVMe included, $0 |
| Monthly total | ≈ $1,517 | ≈ $150 |
| Annual total | ≈ $18,204 | ≈ $1,800 |
| Effective cost per used vCPU | ≈ $552 | ≈ $59 |
The owned side needs one justification: the real load (2.6 vCPU, 26 GB) fits on a single 8-core/64 GB box with room to spare. The second box is redundancy and headroom, not capacity — lose a machine and the fleet still holds. So this is not a stripped-to-the-bone comparison; the $150 side carries an N+1 spare the $1,517 side doesn't even price in.
The delta is ≈ $1,367/month, or about $16,400/year — a 10× multiple — for running the same workloads. That is what 8% utilization costs when the meter runs per provisioned second: you don't pay for the laptop's worth of compute you use; you pay for the four servers you reserved to feel safe.
Sensitivity: at what utilization does metered cloud win back?
A single-point comparison at 8% would be cherry-picking if it stopped there. The honest question is how the delta behaves as utilization rises — and, more importantly, what happens when the cloud side fights back with rightsizing instead of sitting static at 4 nodes.
The table below recomputes the matchup at four utilization levels. The "static cloud" column keeps all 4 nodes (the survey-median reality: overprovisioned and untouched). The "rightsized cloud" column sheds nodes to match real load plus 30% headroom — the outcome of actually doing FinOps. The owned column stays at 2 boxes throughout, since even 60% of this fleet's load (19 vCPU-equivalents, well under 128 GB) fits comfortably.
| Utilization | Real load | Static cloud | Rightsized cloud | Owned (2 boxes) |
|---|---|---|---|---|
| 8% | 2.6 vCPU | $1,517 | 1 node ≈ $434 | $150 |
| 20% | 6.4 vCPU | $1,517 | 2 nodes ≈ $795 | $150 |
| 40% | 12.8 vCPU | $1,517 | 3 nodes ≈ $1,156 | $150 |
| 60% | 19.2 vCPU | $1,517 | 4 nodes ≈ $1,517 | $150 |
Three conclusions survive contact with this table:
- Utilization alone never flips the match at this fleet size. The owned side is flat at $150 while both cloud columns only grow. There is no crossover row — the meter charges for provisioned capacity, and provisioned capacity is what utilization measures the waste of.
- Rightsizing closes most of the gap — if you actually do it. A disciplined team running one node instead of four pays $434, not $1,517. The uncomfortable corollary is the report's real message: almost nobody does. Overprovisioning rose from 40% to 69% in a year when every cost tool on the market screams about it. The waste isn't a knowledge problem; it's a doing problem.
- The row that flips the math is ops, not utilization. Add a fully-loaded platform cost of, say, 10 hours a month at $100/hour ($1,000) for patching, upgrades, and on-call on machines you own, and the effective owned total becomes ~$1,150 — at which point the rightsized-cloud column wins at every utilization level. Self-hosting arbitrages the compute meter, but it bills you in engineering hours instead. A team with no spare platform capacity can burn the entire $16k/year savings in labor before March.
So the representative verdict is conditional, and both conditions matter: if your fleet looks like the survey median — static, 8%, nobody's rightsizing — the meter is the problem and consolidation wins by an order of magnitude. If you have the discipline to run lean on autoscaled cloud (or no platform hours to spare), the cloud column is fixable without moving a single workload.
Why utilization keeps falling
The report documents the what; the why is structural, and it explains why the number keeps sliding despite a decade of cost tooling:
- Requests-as-insurance. Developers set CPU requests for the worst minute of the worst day, because the penalty for under-requesting (throttling, OOMKills, a 3 a.m. page) falls on them while the penalty for over-requesting (a line item) falls on nobody in particular. Across hundreds of services, the safety buffers compound into the 69% overprovisioning figure.
- The autoscaler headroom ratchet. Horizontal pod autoscaling and cluster autoscaling add capacity fast and shed it slowly — cooldown timers, disruption budgets, and minimum-replica policies all bias toward holding nodes. Fleets grow into peaks within minutes and decay out of them over hours, if ever.
- The GPU land-grab. At 5% average utilization, GPUs are the purest example of reserve-now-use-later: teams claim scarce accelerator capacity the moment it's available because re-acquiring it later is uncertain, then park inference workloads that idle between requests on whole cards.
- Idle-but-reserved agent workloads. AI-agent sandboxes hold inference connections open and wait on tool calls while registering as near-zero CPU to the scheduler — one analysis of the report's findings notes workloads changed faster than the tooling built to manage them. The fastest-growing workload class on Kubernetes is structurally invisible to CPU-based bin-packing.
None of these yield to a dashboard. They yield to policy: enforced VPA recommendations, Karpenter-style just-in-time node provisioning, GPU time-slicing or MIG partitioning, and request quotas with chargeback that make the buffer's cost visible to whoever sets it.
What to do on either side
Staying on metered cloud? The report is your budget justification for the unglamorous checklist:
- Run VPA in recommendation mode, then enforce its output on the top 20 fattest workloads.
- Replace static node groups with Karpenter or cluster-autoscaler profiles that bin-pack and shed nodes within minutes.
- Put HPA on business metrics (queue depth, latency) instead of CPU, so replicas track demand rather than headroom.
- Partition GPUs with time-slicing or MIG before buying more cards.
- Give every namespace a quota plus a cost report its owner actually reads.
The rightsized column in the table above is achievable — it just isn't the default.
Considering owned hardware? Price it with eyes open about the three things the $150/month figure doesn't include. First, the control plane is yours: upgrades, etcd backups, CVE patching, and the on-call rotation for all of it. Second, elasticity goes away: no scale-to-zero, no burst-to-40-nodes for a launch — capacity planning becomes a quarterly exercise again. Third, managed services unbundle: the RDS, ElastiCache, and load-balancer bills don't vanish, they move — either to self-managed equivalents on your boxes or to standalone price tags you must add back to the comparison.
The teams for whom the math works best are the ones with steady-state workloads, existing platform skills, and fleets sitting near the survey median — which, at 8% average utilization, is most teams. The teams for whom it works worst are spiky, bursty, or staffed with zero platform engineers. Know which one you are before either migrating or dismissing the idea.
The bottom line
An 8%-utilized fleet is not a cost problem with a tooling solution — it is a pricing-model problem. Metered billing charges you for the capacity you provisioned to feel safe; flat-rate hardware charges you for the capacity you actually need plus a spare. On a typical small fleet, that difference is roughly $16,000 a year, and no utilization level the industry has ever measured makes it disappear. What can erase it is discipline (real rightsizing on cloud) or labor (real ops hours on owned boxes) — so the decision was never really about hardware. It is about which of those two bills your team is actually willing to pay.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



