Render's August changelog reads like a dare: service builds now run on faster CPU and disk nodes, and the all-runtime median fell from 38 seconds to 21 — a reduction Render itself calls a conservative 40%. If you run a self-hosted PaaS on owned machines, the question writes itself: can a build pool you own match 21-second deploys, and what does it actually cost?
The short answer: on warm-cache execution, yes — published warm-build numbers on comparable persistent builders run 12–15 seconds, comfortably under Render's median. On cold builds, no — dependency installs alone run 30 seconds to several minutes. And Render's 21 seconds is a blended median across runtimes, queue time, and cache states, so a single warm-cache measurement never was the comparison. The honest repro splits all four — cold vs warm, queue vs execution — and this post gives you the exact protocol plus every sourced number.
| Build profile | Render published median | Self-hosted reference (published) | Source |
|---|---|---|---|
| All runtimes, blended | 38s → 21s (Jul 2026) | — | Render changelog, Aug 7 2026 |
| Docker runtime, blended | 87s → 32s (Mar–Jun 2026) | — | Render changelog, Jun 11 2026 |
| Warm cache, execution only | included above | 12–15s | Depot GCB benchmark, Flywheel BuildKit notes |
| Cold cache, execution only | included above | 32s–11min (dep-tree dependent) | Flywheel, registry-cache case study |
| Queue + scheduling | included above | yours to measure — protocol below | — |
Every cell in the right column carries its provenance. Nothing in this post is a number I measured on rented hardware and ask you to trust — the Hetzner column is an expected range from published reference points, and section 5 gives you the commands to replace it with your own silicon.
What Render actually changed
The August 7, 2026 changelog entry, titled "Reduced median service build time by 40% (all runtimes)," says service builds now run on nodes with faster CPU and disk. Before the rollout, in the week of July 13, median build time was 38 seconds, consistent with prior weeks. Since July 27 — the first full week post-rollout — the highest observed weekly median is 21 seconds. Note the arithmetic: 38 to 21 is closer to 45%, which is why Render calls 40% "conservative." Build times remain in that range.
This was the capstone, not the whole project. Three single-runtime wins landed first: Node.js builds down 25% in June, Python builds down 27% in late April, and most substantially, Docker image builds down 60% — 87 seconds to 32 — in the June 11 entry. The Docker list is the one to study, because Render named all four optimizations: tuning chunk size and parallelism for build image uploads, speeding up build scheduling, parallelizing generated image upload with export to build cache, and storing build cache in Render's own image registry with automatic pruning.
Read that list twice. Only the August win is "faster hardware." The Docker win is three parts pipeline engineering and one part cache operations — scheduling latency, upload parallelism, and cache pruning. Anyone reproducing this on owned hardware has to reproduce the ops, not just the clock speed.
The reference workload (so the numbers mean something)
A benchmark without a defined workload is a vibe. Here is the concrete service every number in this post is calibrated against: a multi-stage Dockerfile building a typical web service — a Node frontend-plus-API with roughly 400 MB of node_modules, or a Python equivalent with a pinned requirements.txt around 150 MB — producing a slim runtime image under 300 MB. That shape is deliberately typical: it sits inside the native-runtime mix (Node, Python) and the Docker runtime that Render's own changelog splits out, and its warm path is "rebuild the app layer, reuse the dependency layer," the most common real deploy.
Two cache states, defined precisely. Warm means the dependency layer and all base images are already on the builder: only app code changed. Cold means nothing is cached — fresh builder, empty cache, full dependency install and base-image pull. Every claim below says which state it is in, because the gap between them is the entire story: published cases span from 12-second warm rebuilds to 11-minute cold builds on a 4 GB dependency tree.
Why the headline median is hard to compare directly
Render's 21 seconds blends three things a single benchmark cannot reproduce in one run. First, runtimes: Docker alone medians at 32 seconds post-optimization, so the 21-second all-runtime figure is pulled down by faster native runtimes. Beating "21 seconds" with a warm Node rebuild while your Docker path takes 45 is not matching Render — it is matching one slice of Render.
Second, queue and scheduling time. Render's published medians measure the build as the user experiences it, which includes waiting for a builder and the scheduling work Render explicitly optimized in June. Your local docker buildx build timer measures execution only. Any comparison that puts a stopwatch on execution and calls it a deploy time is flattering the self-hosted side by exactly the queue depth it chose not to measure.
Third, cache state. Render's fleet serves a mix of first builds, dependency-change rebuilds, and code-only rebuilds across thousands of services. Your one warm rebuild is the fastest point in that distribution, not its center. The repro protocol below exists to make the comparison honest: it separates all four cells — cold/warm crossed with queue/execution — so each can be compared against the right reference.
The repro protocol: four cells, no hiding
One AX42-class dedicated box is the reference builder: 8 cores / 16 threads, 64 GB RAM, and NVMe storage — Hetzner's AX42 line ships a Ryzen 7 PRO 8700GE with 64 GB DDR5 and 2×512 GB NVMe. Run a persistent BuildKit daemon on it, not an ephemeral per-build container, and pin one builder per architecture — eliminating QEMU emulation is consistently the single biggest build-time win in published operator notes.
Definitions first, because this is where benchmarks cheat. Cold means you flushed everything: buildx prune --all plus removing the local registry cache, so the next build pulls base images and reinstalls dependencies. Warm means you changed one app file and rebuilt with the daemon cache intact. Queue time is the wall-clock gap between "build requested" and "first build step starts executing" — on a single-box repro with one build at a time this is near zero, which is precisely why you must also run the concurrency test below and report both.
# Cold cell: flush everything, then time the full build
docker buildx prune --all --force
time docker buildx build --push -t registry.example.com/app:cold-test .# Warm cell: touch one app file, rebuild, time it
touch src/app.ts
time docker buildx build --push -t registry.example.com/app:warm-test .Record four numbers per run: queue wait, execution wall clock, bytes pushed, and cache-hit ratio from the BuildKit trace (--progress=plain output marks every CACHED step). Run each cell three times and report the median — Render reports weekly medians, so single lucky runs are not the comparison. Then the concurrency test that turns a box into a pool: fire eight builds at once and record how queue time grows while execution stays flat. That curve, not the single-build stopwatch, is what you compare against a hosted platform's median.
A pool, not a box: where self-hosted usually loses
Single-build latency is the easy part. The published warm numbers — 15 seconds of image build inside a 28-second total in Depot's Cloud Build benchmark, 12 seconds warm (about 7 of pure go build on a daemon-cache hit) in Flywheel's BuildKit notes — all come from persistent builders with warm local disk. Owned hardware matches that trivially, because that is what owned hardware is: a warm disk that nobody else reaps.
The failure mode is cache architecture. A stateless worker that pulls its cache from a registry on every build still has to fetch and unpack every cached layer before it can build on top of it — for a multi-gigabyte dependency layer that is minutes of network transfer for steps buildx cheerfully reports as CACHED. The docker-buildkit-fleet design notes make this the centerpiece: a warm BuildKit worker with the bytes already on local disk skips the transfer entirely, and keeping that warm worker without making it a single point of failure is the actual engineering. Render's June optimization list says the same thing from the vendor side — its own registry cache with automatic pruning is load-bearing infrastructure, not a flag.
Queue behavior is the second half. One box running one build has no queue; a team of twenty pushing at 5pm has nothing but queue. Render's "speeding up build scheduling" line item exists because at fleet scale, scheduling is user-visible latency. Your repro's eight-concurrent-builds curve tells you where your pool's knee is — the point where adding builders beats tuning execution. Size the pool from that curve, not from the single-build stopwatch, and re-run it after every cache-architecture change. Cache mounts deserve a mention here too: one operator took amd64 builds from 8.6 minutes to 0.8 once the daemon was warm with cache mounts plus S3 backing — roughly a 10× drop from cache hygiene alone, no hardware change.
The cost line nobody puts in the benchmark
Here is the fixed cost, stated plainly. Hetzner's AX42-1 (Ryzen 7 PRO 8700GE, 8C/16T, 64 GB ECC, 2×512 GB NVMe, unmetered gigabit) lists at €97.30/month plus a €49 setup fee after the June 2026 dedicated price rise, with a limited-stock tier at €77.30. That is the whole hardware line for a reference builder: no per-minute meter, no egress bill for cache traffic that never leaves the box.
Throughput arithmetic turns that into a per-build cost at any volume. At a 20-second warm execution, one builder sustains roughly 150 builds per hour before queueing dominates — call it 100/hour to leave headroom. A team doing 1,000 builds a month pays about €0.10 per build on the fixed box; at 10,000 builds it is about €0.01. Render includes standard builds in plan pricing and meters only the Performance build pipeline (billed in batches of pipeline minutes), so the honest comparison is this fixed curve against your own current metered line item — plug in your bill, not a rate card I picked to flatter either side.
The hidden line is ops, and Render's changelog itemizes it for you: scheduling, upload parallelism, cache export, and automatic pruning. Self-hosting inherits every one — prune jobs, cache-eviction policy, multi-arch builder fleet, the concurrency curve above — as undifferentiated work your team owns forever. That is not an argument against it; it is the actual price tag next to the €97.30.
Verdict: when owned builders win
Match Render's 21 seconds on warm execution? The published reference points say yes with margin — 12–15 seconds is the warm range on persistent builders, and a warm local NVMe disk is the cheapest way to buy it. Match it on cold builds? No: dependency installs and base pulls run 30 seconds to minutes, and no CPU upgrade changes network-bound layer fetches. Match the blended median end-to-end, queue included, at team scale? That is a pool-sizing and cache-architecture question, and the protocol above is how you answer it for your workload instead of trusting anyone's headline number.
Self-host the build pool when warm rebuilds dominate your deploy mix, your monthly build count pushes per-build cost toward a cent, and someone on the team will own pruning and scheduling as real work. Stay on the hosted pipeline when cold builds dominate (monorepos with shifting lockfiles, heavy native deps), when deploy-time queueing already hurts, or when nobody wants the cache-hygiene pager. Either way, split your own numbers into the four cells first — the median you beat should be the one you actually measured.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



