Two numbers changed on Hetzner Cloud this year, and both of them can quietly break a fleet that runs on it. The first is 50: for years, every HTTP/HTTPS service on a Hetzner Cloud Load Balancer silently dropped connections idle for more than 50 seconds — no knob, no warning, just a dead WebSocket. Since April, that window is tunable from 30 to 300 seconds per service. The second is an error code: since May 1, deleting a Primary or Floating IP that is still assigned to anything fails with must_be_unassigned. If your teardown or failover automation ever deleted an IP in one step, it now 422s.
Neither change made headlines. Both deserve a runbook update. Here is what changed, when it actually changed, and the two concrete edits to make: a per-service idle-timeout tuning table for your ingress layer, and an unassign-before-delete sequence for IP lifecycle automation — plus a September 23 follow-up that changes how you poll IP assignment state.
What actually changed (and when)
The timeline matters because the two changes shipped months apart and through different channels. Here it is from the Hetzner Cloud changelog:
| Date | Change | What it means |
|---|---|---|
| Jan 29, 2026 | Deleting assigned Primary/Floating IPs deprecated | Announced: after May 1, deletion of an assigned IP stops working |
| Apr 30, 2026 | http.timeout_idle added to the LB HTTP Service schema | Idle timeout for HTTP/HTTPS services goes from a fixed 50s to a per-service 30–300s setting, via API |
| May 1, 2026 | Assigned-IP deletion enforced | Delete on an assigned IP now returns must_be_unassigned; unassign first |
| Jun 29, 2026 | Idle timeout editable in the Console | Same knob, now clickable per service |
| Sep 23, 2026 | Primary IP assignee_type returns unassigned when detached | Polling code that checked assignee_id == null still works; code that matched assignee_type == "server" needs a look |
Tooling followed the API quickly. The hcloud CLI gained --http-timeout-idle on load-balancer add-service and update-service; hcloud-go, hcloud-python, the Terraform provider, and the Ansible collection all carry timeout_idle now. Most importantly for Kubernetes fleets, the Hetzner Cloud Controller Manager (HCCM) added a load-balancer.hetzner.cloud/http-timeout-idle Service annotation (a Go duration string like 30s, closing upstream issue #1233), so the value is declarative per Service instead of a Console click someone forgets to replicate.
On the IP side, terraform-provider-hcloud v1.60.0+ and the Ansible collection v6.7.0+ unassign before deleting, so managed-IaC users were covered if they upgraded. Hand-rolled scripts and operators were not — which is where the runbook edits come in.
Runbook update 1: tune timeout_idle per service
The old behavior was a fixed 50-second guillotine: if neither client nor server sent a byte for 50 seconds, the connection died. For request/response APIs that never mattered. For long-lived connections it was the classic mystery failure — an SSE stream that stalls every minute, a WebSocket that drops during quiet periods, a dashboard that reconnects in a loop nobody can explain from application logs.
Operators have hit exactly this in production. One open-source project running SSE behind a Hetzner LB watched streams drop whenever server keepalives drifted past the idle window, and fixed it the only way available at the time: tightening the application heartbeat well under the timeout. That workaround still works, but the tunable timeout flips the default — set the LB window to fit the workload instead of warping every app's heartbeat around a fixed 50 seconds.
The fix is a small decision table, applied per LB service. The rule underneath it: your heartbeat interval must be comfortably below the idle timeout, and the timeout should be the smallest value that still covers your quietest legitimate gap. A longer timeout is not free — idle connections hold LB and backend slots — so don't blanket everything to 300s.
| Workload on the service | Suggested timeout_idle | Why |
|---|---|---|
| REST/JSON APIs, webhooks | 30–60s (leave default-ish) | Short gaps; fail fast and let clients retry |
| WebSocket apps with app-level ping/pong every ~25s | 60–120s | 2–4x the ping interval absorbs one missed beat |
| SSE streams with 15–30s heartbeats | 120–300s | Heartbeats are cheap; reconnect storms are not |
| gRPC streaming / AI-agent event tails | 180–300s | Long quiet stretches are normal; drops are expensive |
Two caveats before you copy-paste. First, the knob covers HTTP/HTTPS services only — TCP-mode services have their own separate behavior, so check which mode your service uses before assuming the setting applies. Second, raising the timeout does not replace heartbeats: it buys headroom for missed beats and GC pauses, not permission to go silent forever.
For a CAPH-managed fleet, the value lives where your Services live. If HCCM provisions your tenant-facing load balancers, annotate the Service:
apiVersion: v1
kind: Service
metadata:
name: events-sse
annotations:
load-balancer.hetzner.cloud/http-timeout-idle: "180s"
spec:
type: LoadBalancer
ports:
- port: 443
targetPort: 8080For the control-plane or platform LBs you manage by hand, the CLI updates a service in place:
hcloud load-balancer update-service web-lb \
--protocol https --listen-port 443 \
--http-timeout-idle 180sVerify after applying: open an SSE or WebSocket connection through the LB, stay quiet past the old 50-second mark, and confirm it survives. If your app already sends heartbeats faster than the timeout, you should have no behavior change at all — which is exactly the point. The audit to run this week is the reverse: list every LB service still on the implicit default, identify the ones fronting streaming endpoints, and raise those deliberately instead of discovering them via user complaints.
Runbook update 2: unassign before you delete (or reassign)
The May 1 enforcement ended the era of DELETE /floating_ips/{id} doubling as "detach and destroy." The new sequence for tearing down or moving an IP is explicit:
# New teardown order: unassign, then delete
hcloud floating-ip unassign 12345678
hcloud floating-ip delete 12345678The same ordering applies to Primary IPs (hcloud primary-ip unassign first) and to both resource types over the raw API: POST /floating_ips/{id}/actions/unassign (or the Primary IP equivalent) must complete before DELETE. Skip the step and the API answers with must_be_unassigned — a hard failure, not a deprecation warning.
Three places in a typical fleet need the edit:
- Failover scripts. The old pattern of delete-and-recreate to move an IP between nodes must become unassign-then-reassign. Note that reassignment never required deletion at all —
assignafterunassignis faster and keeps the address stable — so the enforcement is also a nudge toward the better primitive. - Cluster teardown / node deprovisioning. Any job that deletes servers and then sweeps "their" IPs will now fail on the sweep if the IPs are still attached. Order the unassign before the delete, or detach IPs as part of server deletion.
- Assignment-state polling. The September 23 change sharpens this: a detached Primary IP now reports
assignee_type: "unassigned"instead of"server"with a nullassignee_id. If your automation gates the delete on "is it detached yet," match onassignee_id == null(stable across the change) or handle the new"unassigned"value — don't assert the old"server"-with-null shape.
If you manage IPs with Terraform or Ansible, the fix is a version pin: terraform-provider-hcloud v1.60.0+ and hetzner.hcloud v6.7.0+ already unassign before deleting. Pin at or above those, run a plan on a scratch IP to watch the ordering, and then grep your own scripts and operators for direct DELETE calls on IP resources — those are the ones the provider upgrade doesn't cover.
One more spot worth checking on CAPI fleets: if your control-plane endpoint rides a Hetzner LB or a floating IP managed outside the provider — bring-your-own-LB setups, or kube-vip fronted by a cloud floating IP — the same ordering applies to whatever re-points that address during control-plane replacement. Provider-managed resources inherit the fix from the provider; hand-managed endpoint IPs get it only from your runbook.
The fleet audit checklist
Both changes reward one focused pass over a Hetzner-backed fleet. Concretely, this week:
- List every LB HTTP/HTTPS service and flag streaming endpoints (WebSocket, SSE, gRPC tails) still on the implicit 50s behavior — raise those per the table above.
- Confirm app-level heartbeats run at most at half the configured idle timeout, so one missed beat can't kill a connection.
- Set
timeout_idledeclaratively (HCCM annotation or Terraform) rather than Console clicks, so the next LB rebuild keeps it. - Grep failover/teardown automation for IP
DELETEcalls and insert the unassign step; pin Terraform/Ansible past the versions that do it for you. - Update IP-state polling to accept
assignee_type: "unassigned"(or key offassignee_id == null).
None of this is a migration. It's the unglamorous middle layer of platform work: two upstream knobs turned slightly, and a fleet whose runbooks now say the true thing about both. The teams that get bitten are the ones running last year's assumptions — a hardcoded 50-second expectation in a reconnect backoff, a teardown script written when delete implied detach. A half-hour audit beats a 3 a.m. page from either.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



