Skip to main content

Hetzner's €214 GEX45: What 24 GB of Blackwell VRAM Buys Self-Hosted Inference (and When It Beats Serverless GPUs)

10 min readDora NodaDora Noda
Share
On this page

On September 3, 2026, Hetzner added the GEX45 to its dedicated GPU lineup: an RTX PRO 4000 Blackwell card with 24 GB of GDDR7 ECC VRAM on a full server (i5-13500, 64 GB DDR4, 2× 512 GB NVMe) for €214 a month plus a €209 setup fee. The headline number is the price. The number that actually matters for inference is 24 — four gigabytes more than the GEX44 it sits next to, and a generation jump from Ada to Blackwell with FP4 support. This post works out exactly what that buys: which models newly fit, and at what utilization a flat-rate box beats per-second serverless GPU rentals.

TL;DR: the verdict up front

Hetzner lists the GEX45 at €214/month, or US $249 — we will do the crossover math in dollars and state the exchange assumption once: €214 ≈ $249 at Hetzner's own list prices. Against on-demand serverless GPUs in the 24 GB class ($1.10–1.29 per GPU-hour on Modal and Lambda Labs), the GEX45 wins once sustained utilization passes roughly a quarter to a third of the month — about 190–230 GPU-hours. Against $0.20/hr community spot (RunPod's RTX 4000 Ada tier), it never wins on raw price — spot is cheaper even at 100% utilization — but spot carries no availability promise and no persistent local state. And the 20→24 GB step is the difference between "14B at best" and "27–32B-class models in NVFP4 or 4-bit quants fit with headroom." If your inference load is always-on, order the box. If it is bursty experiments, stay serverless.

The crossover math, first

A 730-hour month turns every serverless hourly rate into a monthly bill at a given utilization. The GEX45 is a flat line — with one wrinkle, the $249 setup fee, which amortizes to about $21/month over a year, $42 over six months, or $83 over a quarter. Here is every row that matters, in one currency:

Monthly cost at utilization →10% (73 hrs)30% (219 hrs)60% (438 hrs)100% (730 hrs)
Spot, RTX 4000 Ada class ($0.20/hr)$15$44$88$146
RunPod-class on-demand (~$0.55/hr)$40$120$241$402
Modal A10, per-second (~$1.10/hr)$80$241$482$803
Lambda A10 on-demand ($1.29/hr)$94$283$565$942
GEX45, setup sunk (yr 2+)$249$249$249$249
GEX45, 12-month amortized$270$270$270$270
GEX45, 6-month amortized$291$291$291$291
GEX45, 3-month amortized$332$332$332$332

Read the breakevens off the table. Against Lambda's $1.29/hr, the flat $249 line crosses at 193 hours — 26% of the month. Against Modal's $1.10/hr, it crosses at 226 hours — 31%. Even on the harshest row, a three-month commitment ($332 effective), the box beats on-demand serverless at 36% utilization. Against the $0.55/hr mid-tier, breakeven needs 62% — a genuinely always-on workload. And against $0.20 spot, breakeven needs 171% of the month, which is to say it never happens.

Three honest qualifications before you budget off this table. First, hourly rates move weekly and the "$0.55" mid-tier row is a representative on-demand quote, not a contract — re-check the current board before committing. Second, the serverless rows buy elasticity the GEX45 cannot: scale to zero, burst to forty GPUs, shut it all off on Friday. The box buys the opposite: a fixed integer count of GPUs that cost the same whether they serve a million tokens or idle. Third, spot's $146 full-month price is real money saved right up until a preemption kills your eval run at hour eleven — fine for batch experiments, disqualifying for anything user-facing. The decision rule that falls out: bursty or experimental (under ~10% sustained) stays serverless, steady production inference (over ~30%) moves to the box, and the 10–30% band is decided by how much you value elasticity over a fixed bill.

The card in numbers

The GEX45's GPU is the RTX PRO 4000 Blackwell SFF Edition. Against the GEX44's RTX 4000 SFF Ada:

GEX44 (Ada)GEX45 (Blackwell)
CUDA cores6,1448,960 (+46%)
Tensor cores192 (4th gen)280 (5th gen, FP4)
VRAM20 GB GDDR6 ECC24 GB GDDR7 ECC (+20%)
Memory bandwidth320 GB/s432 GB/s (SFF)
AI throughputFP8-era Tensor cores~1,290 FP4 TOPS (sparse)

NVIDIA rates the 5th-gen Tensor cores at up to 3× the prior generation's AI throughput with FP4 precision — the format that halves weight memory versus FP8 and is the reason a 24 GB card can now contemplate model sizes that used to need 32 GB or more. Both cards carry ECC VRAM, so the comparison is clean: no reliability asterisk changes sides here. What changes is capacity, bandwidth, and the quant format the silicon accelerates natively.

What fits in 24 GB that didn't fit in 20

VRAM is a cliff, not a slope: a model either fits with its KV cache or it does not run at all. The 20→24 GB step crosses exactly one important cliff — the 27–32B weight class in modern quants:

Workload20 GB (GEX44)24 GB (GEX45)
8B model, BF16 + long contextFits, tight at 128kFits with headroom
14B model, BF16Fits, short context onlyFits comfortably
27B model, NVFP4 (~21.8 GiB weights)Does not fit — weights alone exceed capacityFits (see budget below)
30–32B model, 4-bit (Q4)Does not fitFits
LoRA fine-tune, 8B + adaptersFits, small batchFits, larger batch / longer sequence
Embedding + rerank pool for agent sandboxesOne model at a timeTwo small models resident

The row worth working out is the 27B NVFP4 case, because "fits" deserves arithmetic, not vibes. Published NVFP4 checkpoints of 27B-class models (e.g. Qwen3-27B-NVFP4) ship weights of about 21.8 GiB. The card exposes roughly 23.9 GiB usable. That leaves on the order of 2 GiB for the KV cache, CUDA graphs, and runtime overhead at short context — enough to serve, with batch size and context length as the dials you trade against each other as prompts grow. On the 20 GB card this conversation cannot even start: 21.8 GiB of weights does not fit in ~18.6 GiB usable, full stop. The GEX45 is therefore not "20% more of the same" — it is the cheapest Hetzner tier that serves the current open-weight sweet spot (27–32B) at all.

One more worked number, this time on speed. Token decoding is memory-bandwidth-bound: each generated token streams the full weights once, so single-stream throughput tops out near bandwidth divided by weight bytes. At 432 GB/s against ~23.4 GB of weights, the ceiling is roughly 18 tokens/second for one stream on a 27B NVFP4 model — a roofline sketch with stated assumptions, not a benchmark, and batching changes the picture. The point is directional: this is an interactive-chat-capable box for one house model, not a multi-tenant inference farm. Size the pool as one model per card and the math stays honest.

GEX44 vs GEX45 head-to-head

GEX44 (Ada)GEX45 (Blackwell)
GPURTX 4000 SFF AdaRTX PRO 4000 Blackwell SFF
VRAM20 GB GDDR6 ECC24 GB GDDR7 ECC
Monthly~€184€214 ($249)
Setup fee~€88€209 ($249)
Price per GB VRAM~€9.20~€8.92

You pay about 16% more per month for 20% more memory plus the architecture jump, and the price per gigabyte of VRAM is essentially flat at €9 either way. Two footnotes. First, the GEX44 figures carry "" deliberately: Hetzner repriced several lines in June 2026, and at least one September tracker puts the GEX44 at €234 — verify the live number in Robot before you budget; if the GEX44 really sits at €234, the GEX45 is cheaper per gigabyte by a wide margin and the comparison is over. Second, the setup fee more than doubles (€88 → €209), which punishes short experiments — amortized over a year it adds ~€17/month, over a quarter ~€70. The GEX45 is a box you keep for months, not one you spin up for a weekend eval.

Honest caveats

The box has sharp edges, and three of them are structural. It is one GPU with no autoscaling. Dedicated servers are bought manually in Robot and consumed as static inventory — the cluster autoscaler cannot conjure more of them, so GPU capacity is the integer count of boxes you own. The wiring pattern (bare-metal inventory behind an autoscaled CPU fleet) was covered in depth in our September 19 GEX45/CAPH post and is not re-derived here. The serving stack must actually speak Blackwell. FP4's TOPS number is real silicon, but your framework needs sm_120 kernels and NVFP4 quant support to convert it into tokens — verify vLLM, TensorRT-LLM, or Ollama's current Blackwell support for your exact model format before assuming the datasheet. It is still a workstation card, not a datacenter card. ECC GDDR7 covers memory errors, but you get no MIG partitioning, no NVLink, and a single-GPU memory ceiling — multi-tenant slicing stays a software problem (time-sharing, queuing), not a hardware feature.

Verdict: who orders one this week

Order a GEX45 if you have an always-on inference load in the 8–32B range — a house chat model, an embedding/rerank pair for agent sandboxes, a LoRA fine-tuning loop — running at over ~30% sustained utilization on on-demand serverless today. At that profile the box undercuts the serverless bill every month once past breakeven, and the 24 GB ceiling covers the model class you actually want to serve. Stay serverless if your GPU hours are bursty experiments, if you need to burst past one card, or if preemptible spot meets your reliability bar — $0.20/hr is a price no owned box touches. And if you already hold a GEX44 serving 14B-and-under models happily, there is no forced move: the upgrade case is specifically the 27–32B class that only fits on the new card.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex