On July 3, 2026, at 00:07 UTC, one power feed in a Chicago datacenter went down — and Fly.io's servers responded by throttling their CPUs to the point where, in Fly's own words, they were "unable to do any useful work." A misconfigured safety setting had told every dual-power-supply server to stand down to idle the moment a single feed dropped. That incident traveled: power failures are dramatic, legible, and easy to retell.
But the ORD power failure was the seventh notable entry on Fly.io's status page in nine days, not the first. Read incident-by-incident, the window looks like a run of unrelated bad luck. Read as a tally — the way you'd read it if you were deciding whether to keep your production workloads on the platform — it reads differently. Here is the tally, what each entry actually was, and a reusable method for turning any hosted PaaS's status history into a stay-or-migrate input.
The tally: nine days, seven entries
Every row below comes from Fly.io's own disclosures: its real-time status page (status.flyio.net) and its infra log, the team's written post-hoc explanations. No forum rumors, no third-party uptime aggregators — just what the vendor itself published.
| # | Date | Incident | Layer of your deploy loop it hit |
|---|---|---|---|
| 1 | Jun 25 | Depot remote-builder delays, plus elevated Fly Proxy & Machines API latency in BOM/NRT | Build + control plane |
| 2 | Jun 26 | IPv6 connectivity issues in EWR | Edge network |
| 3 | Jun 28–30 | VictoriaMetrics ingestion delays; dashboards, alerts, and autoscaling signals stale for three days | Observability |
| 4 | Jun 30 | Egress IP connectivity issues in SIN/NRT | Edge network |
| 5 | Jul 1 | Static egress IPv6 issues in NRT, plus elevated GraphQL/dashboard API errors | Edge network + control plane |
| 6 | Jul 2 | SSL certificate issuance outage for several hours | TLS onboarding |
| 7 | Jul 3 | ORD power-supply failure (00:07 UTC), then a separate ORD upstream-network hardware failure that evening | Compute + data adjacency |
Seven entries, nine days, and at least six distinct subsystems: builders, control plane, edge networking (three times, three regions), observability, certificate issuance, and datacenter power. For a team whose entire path to production is "push code, Fly builds it, Fly routes it, Fly measures it" — every layer of that loop failed at least once inside a single fortnight.
What each incident actually was
June 25 — builders and control plane, together. Depot-backed remote builders hit provisioning delays on the same day Fly Proxy and the Machines API degraded in Mumbai (BOM) and Tokyo (NRT). Fly's infra log ties the latter to a Corrosion migration routing change — Corrosion being the gossip-based state store Fly built after outgrowing Consul-adjacent patterns. Two things matter here beyond the outage itself: the control-plane degradation was a migration side effect (the platform was mid-move on its own coordination layer), and builders plus API failing on the same day meant you could neither deploy nor easily check why.
June 26 — EWR IPv6. A regional connectivity incident in Secaucus, disclosed on the status page without an accompanying infra-log entry. On its own, the least interesting row in the table: single region, single protocol family, resolved the same day. It matters only in aggregate, which is exactly the point of tallying instead of headline-reading.
June 28–30 — metrics go stale for three days. The VictoriaMetrics ingestion pipeline fell behind when uneven load distribution into the "aggregator" layer starved a subset of ingestion hosts of CPU, building queue backlogs that then took significant time to drain. Grafana dashboards, alerts, and fly-metrics.net showed missing or delayed data — and, per the community-maintained incident rollup, stale statistics drove bad autoscaling behavior for some apps. An observability outage that also degrades autoscaling is two incidents wearing one coat: you can't see the problem, and the system that's supposed to react to load is reacting to fiction.
June 30 — SIN/NRT egress IPs. Egress IP connectivity issues in Singapore and Tokyo, the second Asia-Pacific edge-network entry in five days. If your app calls allowlisted third-party APIs — the static-egress-IP customer — this is your critical path failing, not someone else's.
July 1 — NRT static egress again, plus API errors. Static egress IPv6 issues returned in Tokyo, joined by elevated GraphQL and dashboard API errors that Fly's infra log attributes to a Redis cleanup job. NRT now appears in three of the five entries so far. Nobody in US-East felt any of this — which is precisely why "was my region affected" belongs in the method below rather than in a platform-wide verdict.
July 2 — certificate issuance stalls. Let's Encrypt had a networking hiccup while failing over datacenters for maintenance; Fly's issuance jobs bailed out and re-queued, and apart from a handful of lucky hostnames, certificate issuance stopped for a few hours (infra log). Fly renews well before expiry, so no existing traffic was affected — but any app mid-onboarding sat waiting for its first certificate. Note the shape: an upstream failure (Let's Encrypt) converted into a platform outage by retry logic that re-queued instead of degrading gracefully.
July 3 — ORD, twice. At 00:07 UTC the power-feed failure and its CPU-throttling surprise (infra log). Then that evening, a separate partial outage: a networking hardware failure at an upstream provider in ORD, with some Managed Postgres clusters in the region unreachable or flapping for about an hour. Same airport code, two unrelated causes, seventeen hours apart. Anyone running Postgres-adjacent workloads in Fly's most popular region got both barrels in one day.
Pattern or bad week?
The honest verdict first: these seven entries share no single root cause. A datacenter power feed, a Let's Encrypt failover, a metrics-pipeline load imbalance, a state-store migration, and a Redis cleanup job are not one failure wearing seven masks. If "pattern" means "common cause," there is no pattern — and anyone selling you that reading is fitting a narrative to noise.
But common cause is the wrong test. The right question for a team deciding where to run production is subsystem spread over time: how many independent layers of the platform you depend on failed, how close together?
By that measure, June 25–July 3 is a genuine signal. Builders, control plane, edge network (three regions), observability-plus-autoscaling, TLS issuance, datacenter power — that is every layer of the push-to-production loop, each failing independently, inside nine days. Independent failures clustering in time don't prove a shared defect; they prove a wide blast surface. When everything fails separately, redundancy at any single layer doesn't save you, because the next failure is somewhere you didn't harden.
And the window wasn't even an isolated spike against a quiet baseline. Zoom out to the full month and June 2026 also held a brief global Machines API outage on June 15 (the widest single incident of the quarter), North American network glitches on June 11 and 24, and SIN/NRT network incidents on June 22. The nine-day cluster sits inside an already-busy month the way a wave sits inside a rising tide. That's not cherry-picking a bad week — it's the opposite: the "bad week" framing understates the cadence, because the surrounding weeks were busy too.
One more check before drawing conclusions: disclosure quality. Every entry above is Fly.io's own telling, and the infra log is unusually candid as vendor postmortems go — it names the misconfigured throttle setting, the CPU-starved aggregators, the upstream failover. A vendor that publishes this level of detail is doing transparency right. But transparency cuts both ways for the reader: the better the disclosure, the more faithfully the tally reflects reality, and this tally says the platform had a rough month across nearly every subsystem.
How to read any status history as a migration signal
This is the same read-the-postmortems method previously applied here to Railway's five outages and to the Vercel/Fly.io/Render July cluster — generalized into five checks you can run against any vendor, in about thirty minutes, before your next renewal or migration decision:
- Cadence vs. the vendor's own baseline. Don't count incidents in a vacuum; compare the window against the surrounding months. Seven entries in nine days means something different on a page that usually shows one a month versus one that always looks like this. Fly.io's June says: elevated, but on a page that was already active — signal, not anomaly-from-nowhere.
- Subsystem spread. List which layers failed. One subsystem failing repeatedly suggests a fixable defect (and a vendor worth watching, not leaving). Five-plus subsystems failing independently suggests blast-surface width that no single fix addresses. This is where the Fly.io window scores worst.
- Blast radius vs. your footprint. NRT appears three times; if you run exclusively in ORD and IAD, your personal tally is two entries, not seven. Filter the vendor's history through your regions, your services (Postgres? static egress? autoscaling?), and your deploy frequency. A builder outage that hits a team deploying twice daily is a different incident than the same outage hitting a team deploying twice a quarter.
- Disclosure depth. Prefer vendors whose postmortems name mechanisms ("uneven load into the aggregator layer starved hosts of CPU") over vendors whose entries say "elevated errors, now resolved." Depth lets you judge whether the fix matches the cause. Shallow disclosure isn't evidence of reliability — it's absence of evidence, and should discount the tally's completeness, not the vendor's incident count.
- Your critical path's recurrence. Ask the narrow question last: did the thing you uniquely depend on fail more than once? Static-egress customers, NRT-region tenants, and heavy-autoscaling users each have a different answer from this window than a single-region hobby project. Migration decisions should be made on your critical path's history, not the platform's average.
Applied back to the tally: a US-only team not using static egress or Postgres-adjacent ORD workloads reads this window as "one rough week, mostly elsewhere" — annoying, not migrating. An APAC team on static egress IPs with autoscaling reads it as "our exact critical path failed three times in nine days" — and that's a migration conversation, or at minimum a multi-region conversation.
The self-hosted counterpoint, honestly
The pitch writes itself: on a self-hosted fleet, the status history is a log you write about your own hardware instead of a vendor's disclosure decision. No waiting for someone else's incident commander to update a page; no wondering whether "elevated errors" means your region. When a power feed drops in your cabinet, you know in seconds, from your own monitoring, with your own runbook open. The ORD throttling surprise is instructive here — that misconfigured BIOS-level default would have bitten an owner-operator exactly once, on hardware they could then reconfigure fleet-wide, with no other tenant's Postgres sharing the blast radius.
But own the failure modes honestly too. Self-hosting doesn't delete a single row from a table like this — it reassigns each row to you. Power feeds still fail. BGP still flaps. Your VictoriaMetrics equivalent still needs its ingestion layer load-balanced, and Let's Encrypt still has maintenance windows (yours just fail differently, through cert-manager retries instead of a vendor's job queue).
What changes is agency and correlation, not physics: your incidents correlate with your change windows instead of a stranger's Redis cleanup, and every fix you ship compounds on hardware you keep. For the APAC static-egress team above, self-hosting trades "three vendor incidents I couldn't influence" for "incidents I cause and can prevent" — a good trade only if the team has the on-call depth to hold up its end.
What to do with this
Status pages are the most underused due-diligence input in infrastructure decisions. Teams benchmark latency, model bills to the dollar, and then migrate onto platforms whose own disclosures — free, public, timestamped — would have told them exactly which subsystems fail and how often. Before your next renewal, spend the thirty minutes: pull the vendor's last ninety days, run the five checks, and filter through your footprint. The answer might be "stay, with eyes open." It might be "leave." Either way it will be your tally, not someone else's headline.
Fly.io's infra log deserves genuine credit here: candid, mechanism-level postmortems are what make a thirty-minute audit like this possible at all. The method only works on vendors that publish — which is itself a selection criterion.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with the status history as a log you write yourself. Star the repo on GitHub or deploy your first app today.



