Paying today's street price for an RTX 5090 adds roughly $49–75 a month to the amortized cost of a self-hosted inference node — real money, but not the catastrophe the 2.3x sticker multiple suggests. The bigger shock is on the other side of the ledger: H100 rentals have fallen roughly 70% from their peak, so the own-vs-rent gap that used to be a 21x monthly-cost canyon is now about a 5x ditch. Two GPU price curves moved in opposite directions at once, and most "just buy a 5090 instead of renting" math circulating online was built when both curves sat somewhere else entirely.
That is the short version. The table below is the slightly longer one — a single-GPU inference node, three purchase prices, one honest throughput assumption:
| Card purchase price | Node cost (36-mo amortized + host/power) | Effective cost at 60% utilization* |
|---|---|---|
| $1,999 MSRP | ~$272/mo | ~$2.18/M tokens |
| $3,750 street (spring 2026) | ~$320/mo | ~$2.57/M tokens |
| $4,700 street (Aug 2026 median) | ~$347/mo | ~$2.79/M tokens |
*Serving a 27–32B-class open-weight model at ~80 sustained tok/s (~124M tokens/mo at 60%). Assumptions boxed in the next section — argue with them there, not here.
The premium stings, but look at the last column sideways: dropping from 60% to 30% utilization more than doubles every row, while the MSRP-to-street jump adds about 28% worst case. Utilization still decides the economics; the street premium just moves the goalposts. This post prices both effects honestly, then checks them against what renting looks like now that H100 time is cheap.
Why a 4,700
The mechanism is a memory-market reallocation, and this list has tracked its CPU/RAM side already: AI data centers are vacuuming up fab capacity for High Bandwidth Memory, which consumes roughly four times the wafer area per gigabyte of standard DRAM. Samsung, SK Hynix, and Micron increasingly sell to whoever buys HBM in bulk on multi-year contracts — and that buyer is never a graphics-card factory. Consumer GDDR7 is what's left over.
The numbers trace the squeeze precisely. Standard 2GB GDDR7 modules now cost around $20 each, per TrendForce data cited in Tom's Hardware's August reporting — and a 32GB 5090 needs sixteen of them, about $320 in memory alone, up from roughly half that at the start of the year. Bulk 16GB GDDR7 procurement reportedly surged from $80–90 to $200–300 in the same span, and memory has gone from 30–40% of a card's bill of materials to as much as 80%.
The shortage is spreading, not easing. Reported gaming-GPU output was cut 30–40% in the first half of 2026, SK Hynix's chairman has warned it could persist toward 2030, and AMD pushed at least a 10% increase across its Radeon lineup on memory costs. NVIDIA itself raised RTX PRO 6000 pricing 55% to $13,250 in June while adding $300 to board-partner costs in May.
But keep the BOM math in proportion: the memory-cost increase explains maybe $150 of a $2,700 premium — a single-digit share. The rest is scarcity rent captured somewhere between the fab and the checkout page: constrained supply meeting inelastic AI-driven demand, with retailers and board partners pricing to clear. Street prices tell that story month by month — $3,500–4,000 in March (75–100% over MSRP), roughly 65% over global MSRP by Q1's end, $5,000–6,000 in April snapshots, an August median near $4,700 with the cheapest US retail listings around $4,399 and South Korea touching $5,100. TechPowerUp recorded 50-series medians jumping 41% in August alone. This is not a memory-cost pass-through; it is a shortage market with memory as the trigger.
That distinction matters for the amortization question, because it determines what you're actually betting on when you buy at street: not that memory gets cheaper (that moves the total ~5%), but that the scarcity rent unwinds — or that your utilization is high enough that it doesn't need to.
The Amortization Model (Assumptions Up Front)
Every own-vs-rent post hides its assumptions; here are mine, all chosen conservatively:
- Card price scenarios: $1,999 (MSRP, the fantasy baseline), $3,750 (spring-2026 street midpoint), $4,700 (August 2026 median tracking).
- Rest of node: $1,300 for CPU, board, 64GB RAM, 1000W+ PSU, case — a 575W card plus system overhead demands serious power delivery and cooling.
- Host + power: $180/mo combined (rack space or colo plus ~725W at typical EU/US commercial rates running 24/7).
- Lifespan: 36 months, zero residual value. This flatters renting — used 5090s currently trade strongly — so treat the own-side totals as ceilings, not floors.
- Throughput: ~80 sustained tokens/second serving a 27–32B-class instruct model (FP8/Q4) under concurrent load on one 32GB card. Sanity anchors: 4x5090 rigs sustain 80–120 tok/s on 70B-Q4 (20–30 tok/s per card on a much bigger model), and dual-5090 setups post ~59 tok/s on a 32B reasoning model. Eighty is the realistic middle for a smaller non-reasoning model with continuous batching — not a peak benchmark.
Monthly node cost is (card + build) / 36 + host/power:
| Scenario | Hardware amortized | + Host/power | Node total |
|---|---|---|---|
| $1,999 MSRP | $92/mo | $180/mo | ~$272/mo |
| $3,750 street | $140/mo | $180/mo | ~$320/mo |
| $4,700 median | $167/mo | $180/mo | ~$347/mo |
Two observations before we go further. First, the $2,701 premium collapses to $75/mo once spread over three years — the host-and-power line ($180) outweighs the entire street premium. Anyone modeling "card price = node cost" is missing the bigger fixed line. Second, at 80 tok/s the node serves ~6.9M tokens/day at full tilt, ~207M a month. That capacity denominator is what the next section divides by.
What the Premium Does Per Token
Fixed cost divided by tokens actually served. Three utilization tiers, all three purchase prices:
| Utilization | Tokens/mo | $1,999 MSRP | $3,750 street | $4,700 median |
|---|---|---|---|---|
| 90% | ~186M | $1.46/M | $1.72/M | $1.86/M |
| 60% | ~124M | $2.18/M | $2.57/M | $2.79/M |
| 30% | ~62M | $4.39/M | $5.15/M | $5.57/M |
Read it both ways. Down a column, the premium's damage: MSRP to August-median adds $0.40/M at 90% utilization, $0.61/M at 60%, $1.20/M at 30%. Across a row, utilization's damage: halving utilization from 60% to 30% adds ~$2.20–2.80/M — three to four times what the street premium costs. The sibling H100-breakeven analysis on this list made the same point from the datacenter side (a 10-point utilization gain across a 100-GPU fleet at $2/hr saves ~$175k/year); it holds at single-card scale too.
Whether any row "wins" depends on the hosted endpoint you're substituting. A self-hosted 32B-class model competes with hosted small-model endpoints, not frontier APIs — compare within the weight class. At 60%+ utilization the street-price node lands in the low single dollars per million tokens, competitive with mid-tier hosted small-model pricing while keeping data on your own machines. At 30% it loses to nearly everything metered. The premium didn't change that structure; it narrowed the comfortable margin from "easy win at 40%+" to "needs real traffic past ~55–60%." That is the honest delta, and it belongs in every fleet-sizing spreadsheet this year.
The Other Curve: H100 Time Got ~70% Cheaper
Now the side the buy-the-card discourse keeps pricing stale. At the 2023 peak, on-demand H100s changed hands near $8/hr — $5,760 a month for a single card running 24/7. Early-2026 surveys put the median around $2.99/hr, with RunPod's secure cloud near $2.39–2.89, Lambda in the mid-$2s to low-$3s, Vast.ai's marketplace from under $2, and only hyperscaler on-demand ($7–12) still living in the old world. CloudZero's August 2026 survey summed it up in five words: prices fell hard after Blackwell shipped. Call it $8 to $2.50 — a ~69% fall, squarely in the 64–75% range, with spot and marketplace capacity cheaper still.
Set the two curves against each other as fixed monthly costs, 24/7 both sides:
| Era | Own 5090 node | Rent H100 24/7 | Own-vs-rent ratio |
|---|---|---|---|
| 2023 peak ($8/hr rent, MSRP card) | ~$272/mo | ~$5,760/mo | ~21x cheaper to own |
| Today ($2.50/hr rent, street card) | ~$320–347/mo | ~$1,800/mo | ~5–6x cheaper to own |
Owning still wins the monthly-cost comparison by a wide margin — but the margin compressed roughly fourfold, and the two sides are no longer the same product. The rented H100 brings 80GB of HBM and ~2x the throughput on the same model class, plus headroom for 70B-class models the 32GB 5090 simply cannot fit.
Per unit of capable capacity, the rental side improved even faster than the headline rate suggests. Anyone re-running 2024-era "buy, don't rent" arithmetic with today's numbers on either side is answering last year's question.
Sensitivity and Escape Hatches
Three alternatives change the answer before you pay $4,700:
- Buy used, buy smaller. Used RTX 3090s ($700–999, 24GB) remain the value floor for small-model experimentation, and used 4090s ($2,500–2,800) undercut a new street-price 5090 while giving up little on memory-bound inference per dollar. Marketplace listings even show used 5090s trading far below new-in-stock ask — the scarcity rent is a new-card phenomenon, and the secondhand market partially routes around it.
- Rent fixed-price RTX tin. Hetzner's GPU dedicated line starts around €184/mo for RTX-class cards (EU-only, one GPU per server, no H100 on the menu — their own FAQ confirms RTX-only). An €889/mo RTX PRO 6000 Blackwell box is the top of that ladder. No capex, no $4,700 gamble, flat bandwidth-friendly pricing — at the cost of last-gen VRAM ceilings and whatever availability queue exists this quarter.
- Wait — but price the wait. If scarcity rent unwinds toward MSRP over 12 months, buying today burns ~$75/mo of premium versus waiting. If your workload runs at 70%+ from day one, the tokens served in that year are worth far more than the premium. Idle hardware waiting for cheaper hardware is the most expensive configuration in the table above.
And the one case where buying at street still wins cleanly: sustained, high-utilization traffic on models that fit 32GB, where you value data staying on machines you own and you treat the card's residual value as the upside the 36-month-zero-residual model deliberately ignores. Short, bursty, experimental, or larger-than-32B workloads belong on rented H100 time at today's $2–3/hr — that market finally priced itself into sanity.
What This Means for a Self-Hosted Fleet
For a platform sizing a GPU node pool rather than a single rig, the lesson compounds. Fixed-cost nodes punish variance: ten tenants with uncorrelated bursty traffic multiplex beautifully onto shared 5090 capacity (this is exactly the bin-packing story Kubernetes DRA now makes schedulable), while one tenant's spiky workload on a dedicated card pays the 30%-utilization row. Buy street-price cards only against pooled, measured demand with utilization telemetry (DCGM utilization plus SM-activity counters, not vibes) proving the >55–60% line holds — and hedge the pool's marginal capacity with hourly H100 rentals that cost a third of what they did two years ago. The premium is survivable; unmeasured utilization at street prices is not.
The 5090 premium is a shortage market, not a permanent repricing — size the pool for the utilization you can prove, rent the margin, and let the scarcity rent burn someone else's capex. Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



