Two years ago, renting a single Nvidia H100 cost roughly $8 to $10 per GPU-hour. Today you can get one for about $2.50. That is a 70% price collapse on the most important chip in AI infrastructure, and it quietly rewrote one of the standard arguments in platform engineering: at $8 an hour, buying your own GPU was obviously smart. At $2.50, the answer is genuinely unclear — and the breakeven point moved from “run it a quarter of the time” to “run it nearly around the clock.”
Here is the one-line verdict this post will substantiate: against a $2.50/hour rental, a $31,000 H100 you own only pays for itself if you keep it busy more than about 85% of the time on a two-year amortization, or about 65% on a three-year one. Anything below that — bursty inference, experiments, a product still finding traffic — and renting wins by hundreds of dollars per card per month. Ownership went from the default to a bet you have to justify with a utilization number.
The crash in one chart
The price history is unusually well documented, because GPU-market trackers spent 2023–2026 watching the shortage unwind in public. Reconstructed from provider pricing and market indices:
| Period | Typical H100 rental (per GPU-hour) | What was happening |
|---|---|---|
| Late 2023 | $7–10, spiking above $12 on major clouds | Acute shortage; hyperscaler-only supply |
| Q4 2024 | $8–10 | Peak scarcity pricing |
| Q1 2025 | $5.50–7.00 | Supply improves as datacenter buildouts land |
| Q2 2025 | $3.50–4.50 | Neocloud capacity floods in |
| Late 2025 | $2.85–3.50 | Market stabilizes at a third of peak |
| 2026 | ~$2–3 on specialized clouds | Median on-demand about $2.99 |
Today's spread tells its own story. Budget marketplaces list H100s from around $1.49 (Vast.ai) to $1.99 (RunPod); managed specialists like Lambda sit near $2.49–3.99; CoreWeave is around $4.25; hyperscalers still charge up to $6.88–6.98 per GPU-hour. The Financial Times noted smaller-provider H100 rates falling roughly another 22% year to date, and AWS itself cut P5 H100 prices 44% back in June 2025. The direction has been down for two straight years.
What crashed — and what didn't
Three things are worth separating, because the headline number hides them.
First, the crash is concentrated in on-demand rental on specialized clouds and marketplaces. Hyperscaler on-demand barely budged in comparison: AWS and Azure still charge roughly $7 per H100-hour, nearly triple the neocloud rate. If your workloads live inside a hyperscaler VPC with no exit plan, you have not experienced a 70% price cut. The glut pooled where new capacity was actually built — the neoclouds — and hyperscaler pricing lagged it by the usual 6–12 months, if it follows at all.
Second, the purchase price of the card did not crash. A new H100 80GB still costs around $31,000, and an 8-GPU HGX baseboard runs $250,000 to $320,000. Meanwhile the input everyone hoped would get cheaper — memory — went the other way: HBM scarcity pushed H100/H200 contract pricing up roughly 40% in a six-month window, and the same memory pressure now shows up as $3,500-plus street prices on consumer RTX 5090 cards against a $1,999 sticker. The rental price of compute fell while the purchase price of the hardware stayed flat or rose. That divergence is the entire story.
Third, demand did not collapse — supply caught up and overshot. Nvidia posted a record $75.2 billion quarterly datacenter quarter, the big hyperscalers still plan a combined $630 billion in 2026 capex, and Blackwell chips carry their own wait times into mid-2026. There are even flickers of firming: some trackers show H100 rates ticking back up toward $2.35 after the lows. This looks less like the end of GPU demand and more like the standard commodity cycle — shortage, overbuild, glut — playing out at unprecedented speed.
The worked math: renting at 31,000
Take a single H100 rented at a representative $2.50 per GPU-hour, running inference 24/7 for a 730-hour month:
- Rent, 100% utilization: 2.50 × 730 = $1,825/month, or $21,900/year.
Now buy the card: $31,000 for the GPU plus a realistic $300/month for hosting and power (a 700W card draws real electricity, and it has to live in a rack or a GPU-bearing dedicated box with cooling and bandwidth):
- Own, 24-month amortization: 31,000 ÷ 24 + 300 = ~$1,592/month all-in.
- Own, 36-month amortization: 31,000 ÷ 36 + 300 = ~$1,161/month all-in.
At full utilization, ownership still wins: about $233/month cheaper on a two-year horizon, $664/month on a three-year one. But utilization is doing all the work, because the owned card costs the same whether it computes or idles. Here is the breakeven table against $2.50/hour rent:
| GPU utilization | Rent/month | Own/month (24-mo) | Own/month (36-mo) | Winner |
|---|---|---|---|---|
| 100% (always on) | $1,825 | $1,592 | $1,161 | Own |
| 65% | $1,186 | $1,592 | $1,161 | Rent (24-mo) / tie (36-mo) |
| 50% | $913 | $1,592 | $1,161 | Rent, by ~$250–680 |
| 25% (nights/weekends idle) | $456 | $1,592 | $1,161 | Rent, by 3x or more |
The breakeven sits near 87% utilization on a two-year amortization and 64% on a three-year one. Now compare with the old world: at $8/hour, a month of 24/7 rental was $5,840, and ownership broke even at roughly 27% utilization. Almost any serious workload cleared that bar, which is why “just buy the card” was the obvious advice in 2024. The crash moved the goalposts from “run it a quarter of the time” to “barely ever let it idle.”
And the amortization itself got riskier. A two-year payback assumes the card is still worth running in year three. But you are now amortizing a fixed $31,000 asset against a rental price that fell 70% in two years and could fall further as Blackwell supply lands. If H100 rentals drop to $1.50, your owned card's effective hourly cost roughly doubles relative to the market — and its resale value falls with it. Buying hardware into a falling spot market is catching a falling knife with extra steps.
Two honest caveats in the other direction. Spot and interruptible listings (down to ~$1.19 on RunPod, pennies on Vast.ai) are not numbers to budget production on — preemption makes them a batch-workload price, not an inference price. And hyperscaler-bound teams paying ~$7/hour face completely different math: at that rate, ownership breaks even near 30% utilization and buying remains obviously correct.
When renting wins now, and when owning still does
Renting at $2–3/hour wins for everything bursty, experimental, or still scaling: fine-tuning runs, eval harnesses, agent sandboxes that need a GPU for minutes at a time, inference for a product whose traffic curve is still a guess. At 25–50% utilization the rental bill is a third to a half of ownership, with zero capex, zero resale risk, and the option to switch to Blackwell pricing the month it gets cheap. The flexibility itself has a price, and at current rates it is negative — you get paid to stay flexible.
Ownership still wins in exactly one shape: sustained, predictable, high-utilization inference. A model serving steady traffic at 80%-plus utilization, a speech or embedding endpoint with a flat load curve, a batch pipeline that genuinely runs all night. That is a load profile you can only claim with measurements, not hopes — which makes GPU utilization monitoring the prerequisite purchase, not the accessory.
Note what is untouched by all of this: CPU and RAM. A Hetzner-class dedicated box still costs tens of euros a month flat, with no capex, no depreciation cliff, and bandwidth included. The “own the hardware” pitch never depended on a volatile spot market for CPUs, because there isn't one. GPUs are now a genuinely different capital bet than CPU boxes — a depreciating asset priced against a falling rental curve — and platform economics should treat them that way instead of bundling both under one “self-host everything” slogan.
What this means for a self-hosted GPU pool
For a platform running tenant GPU workloads on its own fleet, the rational response is a barbell, not a side: a small owned base sized to measured steady-state load, plus rented burst for everything above it. Size the owned pool to the trough, not the peak — the trough is the only utilization you can guarantee — and let the neocloud absorb spikes at $2.50/hour rather than buying cards that idle at 30%.
Two mechanisms make the owned base cheaper to justify. Fractional sharing — HAMi-style vGPU slicing today, Dynamic Resource Allocation as it matures — packs multiple tenants onto one card and pushes its utilization toward the 85%-plus zone where ownership wins. And a proper autoscaler with scale-to-zero for GPU pools (KEDA on queue depth, not CPU) stops the owned fleet from silently becoming the always-on bill the rental market just undercut. The platforms that lose in this market are the ones holding fixed GPU capacity at 40% utilization while quoting 2024 amortization math.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Owned hardware still wins outright for CPU-bound workloads; for GPUs, the honest answer is now “measure first, then decide.” Star the repo on GitHub or deploy your first app today.
Sources
- CloudZero, “H100 GPU Cost in 2026: Buy, Rent, and Cloud Pricing Compared” (verified August 2026) — $31k new card, $1.49–6.98/hr rental spread
- Silicon Data, “Compute Market Q1 2026 Outlook” and “H100 Rental Price Over Time (2023–2025)” — $7–10/hr 2023 history, peaks above $12
- Jarvis Labs, “NVIDIA H100 Price Guide 2026” — quarterly timeline Q4 2024–2026, $2.99 median
- DataCenterKnowledge / FT reporting — ~22% YTD decline among smaller providers, pricing compression
- Spheron / Thunder Compute blogs — AWS 44% P5 cut (June 2025), spot dynamics
- IntuitionLabs “Data Center GPU Pricing 2026” — HBM-driven 40% contract price pressure, $630B hyperscaler capex
- Hetzner press room — GEX45 (€214/mo) and GEX130 (€838/mo) dedicated GPU servers



