Skip to main content

H100 Rentals Crashed 70%. Should You Still Buy the GPU?

9 min readDora NodaDora Noda
Share
On this page

Two years ago, renting a single Nvidia H100 cost roughly $8 to $10 per GPU-hour. Today you can get one for about $2.50. That is a 70% price collapse on the most important chip in AI infrastructure, and it quietly rewrote one of the standard arguments in platform engineering: at $8 an hour, buying your own GPU was obviously smart. At $2.50, the answer is genuinely unclear — and the breakeven point moved from “run it a quarter of the time” to “run it nearly around the clock.”

Here is the one-line verdict this post will substantiate: against a $2.50/hour rental, a $31,000 H100 you own only pays for itself if you keep it busy more than about 85% of the time on a two-year amortization, or about 65% on a three-year one. Anything below that — bursty inference, experiments, a product still finding traffic — and renting wins by hundreds of dollars per card per month. Ownership went from the default to a bet you have to justify with a utilization number.

The crash in one chart

The price history is unusually well documented, because GPU-market trackers spent 2023–2026 watching the shortage unwind in public. Reconstructed from provider pricing and market indices:

PeriodTypical H100 rental (per GPU-hour)What was happening
Late 2023$7–10, spiking above $12 on major cloudsAcute shortage; hyperscaler-only supply
Q4 2024$8–10Peak scarcity pricing
Q1 2025$5.50–7.00Supply improves as datacenter buildouts land
Q2 2025$3.50–4.50Neocloud capacity floods in
Late 2025$2.85–3.50Market stabilizes at a third of peak
2026~$2–3 on specialized cloudsMedian on-demand about $2.99

Today's spread tells its own story. Budget marketplaces list H100s from around $1.49 (Vast.ai) to $1.99 (RunPod); managed specialists like Lambda sit near $2.49–3.99; CoreWeave is around $4.25; hyperscalers still charge up to $6.88–6.98 per GPU-hour. The Financial Times noted smaller-provider H100 rates falling roughly another 22% year to date, and AWS itself cut P5 H100 prices 44% back in June 2025. The direction has been down for two straight years.

What crashed — and what didn't

Three things are worth separating, because the headline number hides them.

First, the crash is concentrated in on-demand rental on specialized clouds and marketplaces. Hyperscaler on-demand barely budged in comparison: AWS and Azure still charge roughly $7 per H100-hour, nearly triple the neocloud rate. If your workloads live inside a hyperscaler VPC with no exit plan, you have not experienced a 70% price cut. The glut pooled where new capacity was actually built — the neoclouds — and hyperscaler pricing lagged it by the usual 6–12 months, if it follows at all.

Second, the purchase price of the card did not crash. A new H100 80GB still costs around $31,000, and an 8-GPU HGX baseboard runs $250,000 to $320,000. Meanwhile the input everyone hoped would get cheaper — memory — went the other way: HBM scarcity pushed H100/H200 contract pricing up roughly 40% in a six-month window, and the same memory pressure now shows up as $3,500-plus street prices on consumer RTX 5090 cards against a $1,999 sticker. The rental price of compute fell while the purchase price of the hardware stayed flat or rose. That divergence is the entire story.

Third, demand did not collapse — supply caught up and overshot. Nvidia posted a record $75.2 billion quarterly datacenter quarter, the big hyperscalers still plan a combined $630 billion in 2026 capex, and Blackwell chips carry their own wait times into mid-2026. There are even flickers of firming: some trackers show H100 rates ticking back up toward $2.35 after the lows. This looks less like the end of GPU demand and more like the standard commodity cycle — shortage, overbuild, glut — playing out at unprecedented speed.

The worked math: renting at 2.50vsowningat2.50 vs owning at 31,000

Take a single H100 rented at a representative $2.50 per GPU-hour, running inference 24/7 for a 730-hour month:

  • Rent, 100% utilization: 2.50 × 730 = $1,825/month, or $21,900/year.

Now buy the card: $31,000 for the GPU plus a realistic $300/month for hosting and power (a 700W card draws real electricity, and it has to live in a rack or a GPU-bearing dedicated box with cooling and bandwidth):

  • Own, 24-month amortization: 31,000 ÷ 24 + 300 = ~$1,592/month all-in.
  • Own, 36-month amortization: 31,000 ÷ 36 + 300 = ~$1,161/month all-in.

At full utilization, ownership still wins: about $233/month cheaper on a two-year horizon, $664/month on a three-year one. But utilization is doing all the work, because the owned card costs the same whether it computes or idles. Here is the breakeven table against $2.50/hour rent:

GPU utilizationRent/monthOwn/month (24-mo)Own/month (36-mo)Winner
100% (always on)$1,825$1,592$1,161Own
65%$1,186$1,592$1,161Rent (24-mo) / tie (36-mo)
50%$913$1,592$1,161Rent, by ~$250–680
25% (nights/weekends idle)$456$1,592$1,161Rent, by 3x or more

The breakeven sits near 87% utilization on a two-year amortization and 64% on a three-year one. Now compare with the old world: at $8/hour, a month of 24/7 rental was $5,840, and ownership broke even at roughly 27% utilization. Almost any serious workload cleared that bar, which is why “just buy the card” was the obvious advice in 2024. The crash moved the goalposts from “run it a quarter of the time” to “barely ever let it idle.”

And the amortization itself got riskier. A two-year payback assumes the card is still worth running in year three. But you are now amortizing a fixed $31,000 asset against a rental price that fell 70% in two years and could fall further as Blackwell supply lands. If H100 rentals drop to $1.50, your owned card's effective hourly cost roughly doubles relative to the market — and its resale value falls with it. Buying hardware into a falling spot market is catching a falling knife with extra steps.

Two honest caveats in the other direction. Spot and interruptible listings (down to ~$1.19 on RunPod, pennies on Vast.ai) are not numbers to budget production on — preemption makes them a batch-workload price, not an inference price. And hyperscaler-bound teams paying ~$7/hour face completely different math: at that rate, ownership breaks even near 30% utilization and buying remains obviously correct.

When renting wins now, and when owning still does

Renting at $2–3/hour wins for everything bursty, experimental, or still scaling: fine-tuning runs, eval harnesses, agent sandboxes that need a GPU for minutes at a time, inference for a product whose traffic curve is still a guess. At 25–50% utilization the rental bill is a third to a half of ownership, with zero capex, zero resale risk, and the option to switch to Blackwell pricing the month it gets cheap. The flexibility itself has a price, and at current rates it is negative — you get paid to stay flexible.

Ownership still wins in exactly one shape: sustained, predictable, high-utilization inference. A model serving steady traffic at 80%-plus utilization, a speech or embedding endpoint with a flat load curve, a batch pipeline that genuinely runs all night. That is a load profile you can only claim with measurements, not hopes — which makes GPU utilization monitoring the prerequisite purchase, not the accessory.

Note what is untouched by all of this: CPU and RAM. A Hetzner-class dedicated box still costs tens of euros a month flat, with no capex, no depreciation cliff, and bandwidth included. The “own the hardware” pitch never depended on a volatile spot market for CPUs, because there isn't one. GPUs are now a genuinely different capital bet than CPU boxes — a depreciating asset priced against a falling rental curve — and platform economics should treat them that way instead of bundling both under one “self-host everything” slogan.

What this means for a self-hosted GPU pool

For a platform running tenant GPU workloads on its own fleet, the rational response is a barbell, not a side: a small owned base sized to measured steady-state load, plus rented burst for everything above it. Size the owned pool to the trough, not the peak — the trough is the only utilization you can guarantee — and let the neocloud absorb spikes at $2.50/hour rather than buying cards that idle at 30%.

Two mechanisms make the owned base cheaper to justify. Fractional sharing — HAMi-style vGPU slicing today, Dynamic Resource Allocation as it matures — packs multiple tenants onto one card and pushes its utilization toward the 85%-plus zone where ownership wins. And a proper autoscaler with scale-to-zero for GPU pools (KEDA on queue depth, not CPU) stops the owned fleet from silently becoming the always-on bill the rental market just undercut. The platforms that lose in this market are the ones holding fixed GPU capacity at 40% utilization while quoting 2024 amortization math.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Owned hardware still wins outright for CPU-bound workloads; for GPUs, the honest answer is now “measure first, then decide.” Star the repo on GitHub or deploy your first app today.

Sources

  • CloudZero, “H100 GPU Cost in 2026: Buy, Rent, and Cloud Pricing Compared” (verified August 2026) — $31k new card, $1.49–6.98/hr rental spread
  • Silicon Data, “Compute Market Q1 2026 Outlook” and “H100 Rental Price Over Time (2023–2025)” — $7–10/hr 2023 history, peaks above $12
  • Jarvis Labs, “NVIDIA H100 Price Guide 2026” — quarterly timeline Q4 2024–2026, $2.99 median
  • DataCenterKnowledge / FT reporting — ~22% YTD decline among smaller providers, pricing compression
  • Spheron / Thunder Compute blogs — AWS 44% P5 cut (June 2025), spot dynamics
  • IntuitionLabs “Data Center GPU Pricing 2026” — HBM-driven 40% contract price pressure, $630B hyperscaler capex
  • Hetzner press room — GEX45 (€214/mo) and GEX130 (€838/mo) dedicated GPU servers

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex