Skip to main content

Zero-Downtime Docker Compose Deploys Hit HN: What Compose-Native Blue-Green Means for a Self-Hosted PaaS

12 min readDora NodaDora Noda
Share
On this page

On June 24, 2026, the monitoring startup StatusDude published a short, angry, extremely practical post — "Zero-Downtime Deployments with Docker Compose — No Kubernetes Required" — and it did the rounds on Hacker News. The headline numbers are what made it travel: thousands of monitoring checks per minute, workers in three regions, multiple deploys a day, zero dropped requests. The stack behind those numbers is what made it stick: plain Docker Compose behind HAProxy, about sixty lines of proxy config, and a ten-line deploy script. No etcd to babysit at 3 AM.

Strip the post to its load-bearing claim and you get three requirements, stated up front because everything else in this piece hangs off them. First, multiple backend instances, so one can be replaced while another serves traffic. Second, a load balancer that retries a failed request on a different backend, so a dying container never drops anything. Third, a deploy loop that replaces instances one at a time, waiting for health at each step.

That is the whole recipe — and the honest version of this story has two halves: the recipe genuinely works on one box, and it genuinely stops being enough the moment deploys span machines. This post covers both, then answers the PaaS design question the comparison forces: should a self-hosted platform borrow Compose's simplicity for the single-node case, or keep one fleet-grade deploy primitive for everything?

Four hours of Traefik: the proxy is the deploy strategy​

The most instructive part of the StatusDude post is the failure log, because it compresses three deploy-strategy lessons into one afternoon. The team started with Traefik — the default choice for Docker setups, with label-based auto-discovery and a polished dashboard — and ripped it out about four hours later.

Failure one was a label conflict. Their first plan ran a backend_new service alongside backend during cutover, both carrying the same Traefik routing labels. Traefik's Docker provider treats each Compose service as a separate configuration source, so two services with identical labels produced a flat refusal — "Service defined multiple times" — and 404s on every request. The blue-green-shaped plan died to a data-model mismatch: the proxy could not express "these two services are one routable thing during transition."

Failure two was a scale-down race. Reworked to docker compose --scale backend=4 (two old, two new) followed by a scale back to two, the deploy hit routing-table lag: Traefik kept sending traffic to containers already shutting down for several seconds after the scale-down, producing 502s on roughly every other request. Delays, disconnect-before-stop, and passive health checks were all tried; the last was rolled back for being too aggressive. The proxy knew the truth late, and "late" is indistinguishable from "down" at deploy time.

Failure three was the killer, and it is a missing feature, not a misconfiguration. When a request hits a container mid-shutdown — SIGTERM received, Uvicorn draining, connection dropping mid-stream — Traefik's retry middleware retries on the same backend. The dying one. Redispatching to a healthy backend is the long-open upstream request traefik/traefik#2723, and without it every combination of health checks and stop ordering still loses the requests already aimed at the dying container.

Lesson, stated bluntly: your proxy's retry semantics are your deploy strategy's failure semantics. Pick the proxy first and the deploy script writes itself; pick it wrong and no script saves you.

The working recipe: HAProxy plus a ten-line loop​

HAProxy won on one directive: option redispatch. The relevant core of the config reads like this:

haproxy
defaults
  mode http
  timeout connect 3s
  timeout client  30s
  timeout server  30s
  # THE key feature: retry failed requests on a DIFFERENT backend
  retries 3
  option redispatch 1
  retry-on conn-failure empty-response response-timeout 502 503 504

When a request meets a refused connection, an empty response, a timeout, or a 502/503/504, HAProxy retries it — and option redispatch forces every retry onto a different server. During a rolling deploy, the request that lands on a container inside its shutdown window is silently re-issued to the surviving replica. The client never sees the error. StatusDude's own post notes Nginx would do too (proxy_next_upstream has the same redispatch shape); the point is the capability, not the brand.

Around that core sit three independent health layers, each covering the others' blind spots. Per-request retry works in milliseconds and catches the single unlucky request. Passive observation (observe layer7, pulling a backend after three consecutive 5xx responses) catches backends that start erroring under real traffic with no probe cycle to wait for. Active checks (option httpchk against /health, every second, one failure to mark down and one success to restore) catch backends that die silently while receiving no traffic. No single layer would suffice: retry alone masks a broken backend forever, passive checks need traffic to observe, and active checks run on a probe interval rather than instantly.

Discovery is the quiet elegance. A single server-template backend 1-10 backend:8000 line resolves the Compose service name through Docker's embedded DNS (127.0.0.11), re-resolving every two seconds. Containers appear and disappear from DNS as the deploy proceeds; HAProxy follows with no Docker socket mount, no label parsing, and no config regeneration. That is a genuine operational simplification over Traefik's model — a static file instead of a control loop — bought with a two-second discovery lag the redispatch behavior absorbs.

The deploy loop itself is a Makefile target that replaces replicas one at a time:

makefile
prod-deploy:
	@echo "=== Zero-downtime rolling deploy ==="
	@for cid in $$(docker compose -f docker-compose.prod.yml ps -q backend); do \
	echo "Replacing $$cid..."; \
	docker stop $$cid && docker rm -f $$cid; \
	docker compose -f docker-compose.prod.yml up -d --no-deps --no-recreate --wait backend; \
	done
	@echo "=== Deploy complete ==="

Each iteration stops one replica, lets traffic consolidate onto the survivor, then starts its replacement with --wait blocking until Docker's own healthcheck passes — --no-recreate keeps the untouched replica exactly as it is. At every instant, at least one healthy backend serves traffic. Reported result: sub-two-second failover when a backend dies, and the team's first HAProxy deploy dropped nothing after an afternoon of Traefik 502s.

The true blue-green sibling: full second set, one reload, teardown​

One precision is owed here, because the TODO framing says "blue-green" and the loop above is technically rolling replacement — sequential per-replica swaps, not an atomic cutover. Both are compose-native and both are proxy-decided; they are siblings, and a PaaS designer should know exactly where they differ.

True compose-native blue-green runs the entire new set alongside the old one: app_blue and app_green as separate services (or two Compose projects), each independently health-gated as a complete set, with the proxy pointing at exactly one. Cutover is a single proxy reload — nginx -s reload, an HAProxy config swap — moving 100 percent of traffic at once. The old set drains, then tears down. Rollback is the same reload in reverse, instant and total, with the previous version still warm.

Community implementations follow this shape almost uniformly: parallel services, a switch script, reload, teardown.

The tradeoff table between the siblings is short and honest:

Rolling (StatusDude loop)Full blue-green
Spare capacity neededOne replica's worthThe entire set (2x during cutover)
Mixed versions servingYes, during the rolloutNo — cutover is atomic
RollbackRoll forward again, per replicaOne proxy reload back
Deploy durationLinear in replica countOne set-boot plus one reload
Failure blast radiusOne replica at a timeWhole fleet flips at once

Rolling wins when capacity is tight and changes are backward-compatible — the common single-box case, where doubling the set means doubling the box. Blue-green wins when version purity matters (a migration the old code must never see, a protocol break) or when rollback speed dominates every other concern. Note the capacity row carefully: on one machine, "2x during cutover" is not an abstraction — it is RAM and CPU you must actually have free, which is why the single-box world skews rolling and the fleet world, where spare capacity is statistical, can afford blue-green as a default.

Where the single-box recipe stops: the fleet gap​

All of the above fits on one machine. The TODO's real question is what breaks when deploys span machines — and the answer is a capability-by-capability gap, not a single cliff edge:

CapabilityCompose-native (one box)Fleet scheduler (K8s + progressive delivery)
Cross-machine schedulingNone — everything shares one daemon's bin-packingScheduler places replicas across nodes with affinity, topology spread, and resource requests
Per-replica health gating at N replicas--wait per container, sequential, operator-pacedReadiness probes gate each pod; maxSurge/maxUnavailable bound the blast radius declaratively
Automatic rollback on metricsNone — a human notices and re-runs the loop (or reloads back)Argo Rollouts / Flagger run Prometheus-querying analysis steps mid-rollout and abort automatically
Canary percentage trafficNot expressible — the proxy flip is all-or-nothingWeighted traffic shifting in measured steps (1%, 10%, 50%) with promotion gates
Failure domainsThe box is the domain — host death is total outageMulti-node, multi-zone spreading; node drain evicts and reschedules instead of killing
Concurrent tenant deploysOne loop at a time; parallel loops race over shared ports and DNSControl plane serializes per-workload rollouts; dozens proceed concurrently
Declarative desired stateThe Compose file plus operator disciplineGitOps-reconciled objects; drift is detected and corrected, not discovered at 3 AM

The compose-adjacent PaaS tools each close part of this gap, and their ceilings are instructive. docker-rollout packages the rolling trick as a Docker CLI plugin — one POSIX shell script that scales a service to 2x, waits for healthy, removes the old, and rolls back if the new never becomes healthy — but it still talks to one daemon. Dokploy delegates to Docker Swarm's rolling updates, which buys declarative replica counts and multi-node placement at the cost of adopting Swarm's (quietly frozen-feeling) orchestration.

Coolify does health-gated rolling replacement per app, and Dokku ships zero-downtime as a default — both single-box primitives with no metric-driven rollback and no canary percentages. Every one of them is the right answer for one machine and a partial answer for two.

The row that bites first in practice is automatic rollback on metrics. Compose-native rollback is a human decision executed by re-running a script; fleet progressive delivery (Argo Rollouts' Rollout CRD with AnalysisTemplate, Flagger's Canary driven off the same Prometheus signals) aborts a bad release while the canary is still small, before a human has finished reading the alert. That is not a convenience gap — it is the difference between "the deploy failed safely" and "the deploy failed at full traffic and paged someone." A PaaS selling deploys to tenants who sleep through them cannot hand-wave that row.

One primitive or two? The PaaS design decision​

So the design question the TODO poses — borrow Compose's simplicity for the single-node case, or keep one fleet-grade path for every deploy — deserves a direct answer: keep one fleet-grade primitive, and borrow Compose's ideas rather than its topology.

The ideas worth stealing are all in this post's first half. Retry-on-a-different-backend should be a documented property of whatever edge proxy the platform runs, not an accident of defaults — StatusDude's four hours prove the default is often wrong. Layered health detection (request-level retry, passive error observation, active probing) ports to any scheduler unchanged. And the discipline of "at least one healthy backend serves traffic at every instant" is exactly what maxUnavailable: 0 says declaratively; the Makefile loop is the imperative fossil of that declaration.

The topology is what must not fork. Two deploy paths — a compose-simple path for single-node tenants and a fleet path for everyone else — means two failure modes, two rollback stories, two sets of docs, and a migration cliff the day a tenant outgrows one box. That cliff is the exact seam single-box tools already strand teams on: the deploy primitive that assumed one daemon cannot stretch to two, so the team re-platforms under pressure.

A Cluster-API-managed fleet that provisions machine two as routinely as machine one should no more special-case "but what if there is only one machine" in its deploy path than in its networking. The fleet primitive degrades gracefully to one node — replicas share it, the scheduler still gates on readiness, rollback still triggers on metrics — while the single-box primitive cannot grow in the other direction.

There is exactly one honest exception: the platform's own bootstrap story. The machine that runs the control plane before the fleet exists cannot itself be deployed by the fleet primitive — that way lies the "who provisions the provisioner" recursion. A compose-file bootstrap for the seed infrastructure, documented as scaffolding rather than product, is legitimate. Everything tenant-facing should ride the one fleet path from deploy one.

The boring tool wins, then the fleet wins​

StatusDude's post earned its HN run because it is that rare genre: a war story with a parts list. Traefik's retry model, a two-second DNS lag, three health layers, a ten-line loop — each number small enough to verify, each decision explained by the failure that forced it. For a team on one box, that parts list is complete; run it and stop reading Kubernetes release notes guiltily.

For a self-hosted PaaS, the post is chapter one of a longer argument. The single-box recipe nails the semantics every deploy needs — redundant backends, redispatching proxy, health-gated replacement — and the fleet exists to deliver those same semantics across machines, tenants, and 3 AM rollbacks nobody is awake to run.

Borrow the semantics. Build the fleet. And when someone proposes a special simple path for the easy case, ask what happens on machine two.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with fleet-grade rolling deploys instead of a single-box script. Star the repo on GitHub or deploy your first app today.

Sources​

  • StatusDude blog, "Zero-Downtime Deployments with Docker Compose — No Kubernetes Required," June 24, 2026 (Traefik failures, HAProxy redispatch config, three health layers, Makefile rolling loop, thousands of checks/min across 3 regions)
  • Last Week in Cloud Native newsletter, 2026 week 27 articles (HN surfacing of the StatusDude guide)
  • wowu/docker-rollout on GitHub (Docker CLI plugin: 2x scale, health-gated replacement, rollback on unhealthy)
  • traefik/traefik issue 2723 (retry middleware retries the same backend; no redispatch to a healthy one)
  • Dokploy / Coolify / Dokku community comparisons (Swarm rolling updates under Dokploy, health-gated replace in Coolify, zero-downtime default in Dokku)
  • Argo Rollouts and Flagger docs and guides (Rollout/Canary CRDs, Prometheus analysis templates, automated canary rollback)

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex