Skip to main content

Should You Run Plain Docker Compose in Production in 2026? Where the Single-Host Ceiling Actually Sits

13 min readDora NodaDora Noda
Share
On this page

In early May 2026, a Hacker News thread titled "Should I run plain Docker Compose in production in 2026?" collected over 400 points and nearly 300 comments — and the top-voted answer, from Docker educator Nick Janetakis, was a shrug in the direction of yes: "Docker Compose was production ready in 2015 and it still is today." But scroll past the victory lap and the most felt complaint in the thread was quieter: "Could I survive with 10 seconds of downtime, probably, but I'd really like if I could avoid it."

Both comments are right, and the space between them is exactly where the single-host ceiling sits. The thread — a response to Distr's April 2026 production-Compose guide — converged on the same verdict from dozens of angles: Compose is fine until you need more than one host, zero-downtime deploys, or automatic recovery from a dead node. This post turns that consensus into the concrete accounting the thread never quite wrote down: which of those three gaps have cheap fixes, which one is the real ceiling, and what a managed fleet actually closes for a team still running Compose today.

The answer up front​

Here is the whole post in one table. Two of the three gaps have fixes that cost nearly nothing. The third one is the ceiling.

GapPlain Compose todayCheapest fix that actually worksWhat a managed fleet closes
Zero-downtime deploysup --force-recreate stops, then starts: seconds of downtime per deployBlue-green swap via scale-to-2, kamal-proxy, or the docker-rollout plugin — all free, all single-hostRolling updates with readiness gates and automatic rollback, as a default instead of a script you maintain
Self-healingrestart: policies fire only when a container exits; a HEALTHCHECK flipping to unhealthy changes a status light and nothing elseAn autoheal sidecar restarts unhealthy containers; a second €4.50/mo box plus failover covers a dead host — if you build the failoverThe scheduler reschedules onto surviving nodes automatically, for dead containers and dead machines
Multi-hostOne daemon, one machine; no primitive for placing work across hostsSingle-node Swarm, hand-rolled SSH fan-out, or a second Compose host you operate by hand — each one a partial orchestrator you now ownDeclarative multi-host scheduling, service discovery, and sealed secrets across the fleet

If you only need the first two rows closed, stay on Compose — this post will show you how. If you need the third row, keep reading past the "not yet" section, because that is where the honest migration conversation starts.

Gap 1: zero-downtime deploys (the thread's number-one complaint)​

Nothing in plain Compose does a rolling deploy. Recreate a service and the old container stops before the new one is ready; on a typical web stack that is a handful of seconds of refused connections per deploy. The thread's most relatable moment was a Compose user admitting exactly this — ten seconds of downtime they could survive but would rather avoid — and the replies split three ways.

The first camp said: take the ten seconds. "I'm happily taking that 10sec whenever thinking about the lifting I have to do for kube and extra cost," one commenter wrote. For infrequent deploys of internal tools, that math is genuinely correct, and the rest of this post respects it.

The second camp handed over the hand-rolled recipe, and it is worth writing down because it is the actual fix, not folklore: record the running container's ID, scale the service to two instances so a second container comes up, wait for it to report healthy, shift traffic, stop the old container, scale back to one. It works on a single host with stock Compose. Its price is that you are now the deploy controller — the health gate, the traffic shift, and the rollback path are shell script you maintain, and the gap between "script that worked in staging" and "script you trust at 2 AM" is real engineering time.

The third camp pointed at tooling that makes the recipe somebody else's problem. Kamal 2 ships kamal-proxy, which holds requests while containers restart and gives you zero-downtime deploys plus one-command rollback over plain SSH — the same tool 37signals uses to run HEY. A docker-rollout CLI plugin released in early 2026 offers a drop-in docker rollout <service> replacement for docker compose up -d <service>. And several commenters asked the obvious question: why not single-node Swarm, which is barely more complicated than Compose on one machine and brings rolling update_config semantics with it?

Practical bottom line: zero-downtime deploys on one host are a solved problem with free tools. Pick Kamal, the rollout plugin, or Swarm's update config — but stop pretending the ten seconds are free. Count your deploys per week, multiply by the downtime window, and compare against the afternoon it takes to set up the proxy. Almost every team doing daily deploys should have fixed this already.

Gap 2: self-healing (two levels the thread kept blurring)​

Distr's guide names the surprise that bites almost everyone exactly once: restart: unless-stopped fires when a container exits, not when it goes unhealthy. You can watch a container sit in unhealthy for hours while Docker does nothing, because the engine reports health status and acts on exit status, and those are different signals. The thread largely agreed on the fix — run an autoheal sidecar that watches for unhealthy events and restarts the container — and several people noted Swarm restarts unhealthy tasks by default, which is one more reason the single-node-Swarm crowd keeps winning arguments cheaply.

But container-level healing is only half the gap, and the thread kept sliding between the two halves without noticing. A dead container and a dead host are different failures. Autoheal, restart policies, even Swarm on one node — none of them help when the machine itself goes away: the kernel panic, the failed NVMe drive, the datacenter power event, the cloud provider retiring the host underneath you. On plain Compose, a dead host is total downtime until a human intervenes, full stop.

This is where the honest accounting gets interesting, because the hardware half of host-level healing is nearly free and the operations half is the entire job. A second Hetzner box with 2 vCPUs and 4 GB of RAM costs about €4.50 a month — less than the coffee you will drink while setting up failover.

What costs real money is everything around the box: deciding what "failed over" means for stateful services, keeping two Postgres instances in agreement, moving the floating IP, testing the failover path often enough to trust it, and getting paged when the primary dies. One commenter put the stay-simple threshold at roughly "fewer than ten active users, or you can tolerate 30–60 minutes of downtime now and then" — and that tolerance number is the host-level healing budget. If an hour of downtime costs you nothing, a single host with good backups is a complete strategy.

Practical bottom line: close container-level healing this week with an autoheal sidecar (or Swarm). For host-level healing, write down your downtime tolerance in minutes before you buy anything — that number, not the €4.50, decides whether you need a second machine and who builds the failover.

Gap 3: multi-host (the actual ceiling)​

Everything above is repairable on one machine. This gap is not, and it is the one the thread's title question is really asking about. The Docker daemon is single-host by architecture: one daemon, one machine, no scheduler placing work across a fleet. Distr's guide states the consequence plainly — docker compose up runs once and exits; there is no control plane behind it, no reconciler reapplying state, nothing watching the host after the command returns.

The thread offered three exits, and each one concedes the ceiling in its own way. The Swarm exit ("on a single node it isn't really more complicated than Compose") is true but capped: Swarm reuses Compose YAML and adds rolling updates and restart-on-unhealthy, yet even its advocates admit adoption has plateaued and every managed offering has moved to Kubernetes. One commenter wished Docker would reposition Swarm as incremental Compose enhancements — a eulogy disguised as a feature request.

The SSH-fan-out exit (Ansible, docker context over SSH, hand-rolled scripts across N hosts) works until the fleet's state lives in five places and the deploy script needs its own runbook. And the "just add one feature" exit is where the thread linked the essay every Compose veteran has bookmarked: you wanted external-DNS for Compose, then service discovery, then sealed secrets — congratulations, you have built a Kubernetes, except yours has one maintainer and no conformance tests.

There are genuine attempts to push the ceiling up without jumping to Kubernetes. Uncloud, a mostly Compose-compatible tool mentioned twice in the thread, adds multi-host networking, TLS, and rolling deployments to a Compose-like workflow. Single-node k3s got a sincere endorsement ("it's been almost 2 years, thing just works").

Podman plus systemd quadruplets cover the appliance-shaped corner where the Linux host is the product. All of these are reasonable — and all of them prove the point by existing: past one host, you are shopping for an orchestrator; the only question is which one's complexity bill you would rather pay.

Practical bottom line: the day you need workloads placed across machines — for capacity, for availability, or for blast-radius isolation — you have outgrown Compose itself, not your Compose skills. Budget the migration then, not before, and not two years after.

The legitimate "not yet"​

One of the thread's best qualities was its refusal to turn the ceiling into a scare tactic. Nobody serious argued that a side project needs Kubernetes, and the "turkey sandwich" joke that earned 30-plus replies — should you have a turkey sandwich for lunch in 2026? I don't know buddy just do whatever — was the thread's immune system rejecting framework hype. "Not yet, but soon" is a legitimate engineering position, so here it is as a checklist. Stay on Compose while all of these hold:

  • One host holds the load with headroom to spare — CPU, RAM, and disk each under roughly two-thirds utilized at peak, so a traffic spike is boring.
  • Your downtime budget covers a dead host. Single-digit concurrent users, internal tooling, or a business that genuinely tolerates tens of minutes of downtime a few times a year: the thread's rule of thumb, and a good one.
  • Deploys are infrequent or off-peak, or you have already closed Gap 1 with a proxy — seconds of downtime per deploy times deploys per week is a number near zero.
  • State fits the backup story. One Postgres in a named volume with tested restores beats two nodes with untested replication. The day you need live failover for state, you have a real architecture project regardless of orchestrator.
  • You are not building orchestrator features. The moment the roadmap includes service discovery, secret rotation, or multi-host scheduling implemented by you, stop and price the migration — that work is the migration, just without the label.

Between "Compose on one box" and "a fleet" there is also a middle floor the thread kept gesturing at: Kamal for zero-downtime SSH deploys without a control plane, Dokku or Coolify for Heroku-style git-push UX on your own VPS, single-node k3s when you want the Kubernetes API without the cluster. These tools move the ceiling from "one host, downtime per deploy" to "a few hosts, boring deploys" for roughly the cost of learning one new tool. They do not give you automatic rescheduling onto surviving nodes — nothing without a control plane can — but for many teams that row of the table stays empty for years, and that is fine.

What a Cluster-API-managed fleet closes, gap by gap​

For the team that has hit the ceiling — the second host is no longer optional — here is the other half of the accounting the thread implied but never itemized. A fleet managed declaratively (Cluster API provisioning the machines, Kubernetes scheduling onto them) closes each gap as a platform default rather than a script:

  • Zero-downtime deploys become rolling updates with readiness gates: new pods must pass health checks before old ones terminate, failed rollouts pause instead of completing, and rollout undo restores the previous revision. Gap 1's shell script becomes an API object with a status field.
  • Self-healing extends past the container to the machine. A dead pod is rescheduled in seconds; a dead node is detected, cordoned, and its workloads land on surviving nodes while the failed machine is replaced by the provisioning loop. Gap 2's "human intervenes" step becomes the platform's normal Tuesday.
  • Multi-host stops being a project and becomes the premise: the scheduler bins workloads across nodes, service discovery and load balancing are built in, and secrets are sealed objects rather than env files on a box. Gap 3's ceiling becomes the floor.

The honest price, stated the way the thread would want it: control-plane machines that cost money while idle, Kubernetes upgrades that need their own runbook (a bad minor-version bump can do what no Compose quirk ever did — refuse to let nodes rejoin the cluster), and somebody on call for the platform itself, not just the app. Managed control planes and opinionated PaaS layers exist precisely to shrink that bill, but it never reaches zero. That is why the checklist above matters: migrate at the ceiling, where the fleet's monthly ops cost crosses below the cost of the downtime and hand-rolled orchestration you would otherwise eat — not a year early because the tooling looks exciting.

Verdict​

The thread's consensus, compressed to a decision rule: run Compose while one host holds you, fix zero-downtime deploys the week they start costing real users, add host-level healing when your downtime budget says so — and start the migration the day you need a second machine to share load rather than to sit idle as a spare. "Not yet, but soon" is not denial; it is the correct answer for every team below the ceiling, and the thread's 300-comment argument for it is the best peer review a deployment boring enough to forget about will ever get.

When the ceiling does arrive and the second machine needs to share load, not sit idle — that is the gap Bex.co is built for: push a git repo, get a running HTTPS service on machines you own, with the fleet behavior above as the default. Star the repo on GitHub or deploy your first app today.

Related articles

Check your move before you migrate

Free browser tools: check a render.yaml or your Render scripts against bex, or turn a Heroku app or docker-compose.yml into a draft render.yaml. Nothing you paste leaves your browser.

Open the migration tools