Fly.io's status history logs a single stretch of trouble in its Chicago (ORD) region on July 3, 2026: the first update lands at 00:11 UTC, the last resolution fires at 18:11 UTC. Read as one line, that looks like an 18-hour outage. It wasn't. It was two separate incidents, on the same day, in the same region, both blamed on the same upstream provider — a ~6-hour power failure that resolved at 05:59 UTC, then eleven hours of apparent normal service, then a second, unrelated ~1-hour networking hardware failure that started at 17:00 UTC.
That distinction matters more than the headline number. A single 18-hour outage is a bad day. Two independent hardware failures at the same upstream facility, hours apart, is a pattern — and Fly.io's own status page is the postmortem that shows exactly what a tenant on rented hardware can and can't do about it.
The timeline, in Fly.io's own words
Here's what actually happened, pulled verbatim from Fly.io's status history, incident by incident.
Incident 1 — power failure (00:11–05:59 UTC, ~5h48m):
| Time (UTC) | Status | Update |
|---|---|---|
| 00:11 | Investigating | "We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly." |
| 00:29 | Identified | "We've identified and reported power issues with one of our upstream providers in ORD. We're waiting for an update from our upstream for a resolution." |
| 00:54 | Update | "Our provider has advised us their facilities team is working on restoring the power, we'll provide another update as soon as we learn more." |
| 02:59 | Update | "Power restoration work in a subset of ORD is still in progress and impact remains ongoing for a subset of hosts and some Managed Postgres clusters." |
| 04:21 | Update | "Power restoration is ongoing and we're making sure the hosts are healthy before starting customer workloads to avoid issues." |
| 04:41 | Monitoring | "Customer workloads are now starting and we're monitoring the affected hosts. Affected Managed Postgres instances will be investigated." |
| 05:59 | Resolved | — |
Incident 2 — networking hardware failure (17:00–18:11 UTC, ~1h11m):
| Time (UTC) | Status | Update |
|---|---|---|
| 17:00 | Investigating | "We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly." |
| 17:02 | Identified | "We've identified the issue as a networking hardware failure impacting a subset of hosts at one of our Upstream providers in ORD." |
| 17:49 | Monitoring | "A fix has been implemented and we are monitoring the results." |
| 18:11 | Resolved | "This incident has been resolved." |
The impact splits cleanly along the same line. Incident 1's power failure took a subset of hosts offline long enough that Fly.io held customer workloads back until it could confirm the hosts were healthy — that's the multi-hour "Managed Postgres clusters unavailable" window. Incident 2's networking hardware failure produced the same symptom class (machines unreachable, deploys on affected hosts failing) but resolved in about a tenth of the time, once a fix could be applied on the networking side rather than a facilities crew physically restoring power.
Two different failure modes, two different upstream sub-teams involved, two different recovery curves — bundled by Fly.io's own status page into what reads, at a glance, like one long incident because they shared a region and a root cause category ("one of our Upstream providers in ORD").
Why "upstream provider" is doing so much unexplained work
Every status update above uses the same phrase — "one of our upstream providers in ORD" — and never gets more specific than that. Not which provider. Not which facility. Not whether incident 1 and incident 2 trace back to the same physical building or two different ones sharing a regional label. Fly.io runs on bare-metal hardware it leases from third-party providers rather than owning outright in every region, and "upstream provider" is the boundary past which its own status page stops being able to tell you anything, because Fly.io itself is waiting on someone else's facilities team.
That's not a one-off blind spot specific to ORD in July. It's the recurring shape of Fly.io's 2026 incident history:
- Singapore (SIN), June 18 — an upstream core router failed unexpectedly, causing intermittent but complete loss of connectivity for a couple of hours.
- Singapore (SIN), June 22 — the same fault reappeared on the same route; the affected device was finally replaced afterward.
- Toronto (YYZ), June 26 — scheduled upstream network maintenance, 08:00–09:00 UTC, with up to 15 minutes of possible connectivity loss.
- Frankfurt (FRA), July 18 — scheduled upstream network maintenance, 01:00–03:00 UTC, with up to 20 minutes of possible connectivity loss.
- Chicago (ORD), July 3 — the two incidents above.
A router flaky enough to fail identically four days apart, before the hardware itself gets swapped, is the kind of thing a tenant only learns about after the second outage — because the first one didn't come with "and this will probably happen again until we replace the part." That's not a knock on Fly.io's incident communications, which are unusually candid as status pages go. It's the structural ceiling on what candor from the tenant can tell you: Fly.io can only disclose what its upstream tells it, on that provider's schedule, in that provider's words.
The fault domain you don't get to choose
Strip away the specific outage and what's left is a clean statement about who holds which decision. On a rented-rack region, the facility and hardware vendor underneath any given region is a choice Fly.io made once, for reasons a customer never sees — pricing, capacity, a long-term contract. A customer picking "ORD" as a region isn't picking a vendor or a facility; they're picking a label that currently maps to one, and that mapping can carry an outage history the label alone won't show them.
When that vendor's hardware breaks, the tenant's own operational options collapse to exactly one: wait. Fly.io's status updates say it outright — "We're waiting for an update from our upstream for a resolution" — because there's nothing else to do. No internal team to page for a faster fix, no second facility to fail over to within that same region, no lever to pull that isn't "ask the vendor for a status update." The fix ships on the vendor's timeline, communicated through Fly.io, to a customer three layers removed from anyone actually holding a wrench.
A Cluster API-managed fleet inverts which of those decisions the tenant gets to make. The operator chooses the bare-metal vendor per region directly — Hetzner, or any other provider whose contracts and hardware they can inspect before committing capacity to it — and nothing stops them from provisioning a second vendor in a second facility and splitting workloads across both, the same day they decide reliance on one is a risk they'd rather not carry. That's not a hypothetical menu item; it's the same reconciliation loop Cluster API already runs for machine lifecycle, pointed at a second provider account instead of a second replica count. The difference from Fly's position isn't that self-hosted infrastructure fails less. It's that the entity discovering the failure and the entity who could add redundancy against it are the same team, not separated by a support ticket and a status page.
Owning the hardware doesn't buy immunity — it buys the pager
None of this means self-hosting makes power failures and flaky routers disappear. Hetzner has its own outages; any bare-metal vendor does. A team running its own Cluster API fleet on rented hardware from a single provider in a single facility has reproduced Fly's exact exposure — one upstream, one fault domain, one "waiting for an update" — just with fewer people writing the status page updates.
What changes is not the failure rate of physical infrastructure; power grids and network switches fail at roughly the same rate no matter who's leasing the rack. What changes is who decides whether that exposure gets fixed and when. A platform tenant on Fly, Render, or Railway inherits whichever upstream footprint that platform negotiated, sight unseen, and finds out its shape only when an incident report finally says the word "upstream." A team running its own fleet inherits the same category of risk, but gets to see it before an outage forces disclosure, and can act on it — adding a second vendor, moving a facility, negotiating a different contract — without waiting for anyone else's postmortem to greenlight the move. That's the trade: not "no more hardware failures," but "hardware failures you get to plan around instead of read about after the fact."
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, provisioned through Cluster API against whichever bare-metal providers you choose. Star the repo on GitHub or deploy your first app today.



