Skip to main content

Fly.io's Four-Incident Day: What September 23's Networking Pileup Teaches About Depending on Someone Else's Mesh

11 min readDora NodaDora Noda
Share
On this page

Four separate incidents in about twelve hours, across four regions, on three different parts of the platform — compute networking, a marketplace database add-on, the private mesh itself, and the agent-sandbox product. That was September 23, 2026 on Fly.io, the worst networking day in the platform's recent memory, and one of the four was still burning the next morning. If you run production traffic on Fly, you spent the day refreshing a status page. If you are deciding where your next app lives, the day is worth studying incident by incident — not because Fly handled it badly, but because it shows exactly what you can and cannot do when someone else owns your network.

Here is the verdict before the timeline: three of the four failures were networking, none of them had a tenant-side fix, and the pattern stretches back through August. That is not an indictment of Fly's engineering — their status page is among the most transparent in the industry. It is a structural fact about depending on a managed global mesh plus pinned-region marketplace add-ons: when the mesh has a bad day, your multi-region story is a page you watch, not a system you operate. This post walks the full timeline, names what tenants could actually do during each incident, and sketches the owned-infrastructure alternative where the mesh, the regions, and the incident timeline are all yours.

September 23 in one table​

All times UTC, reconstructed from Fly.io's official status page and independent status mirrors (StatusGator, Statussight, Pingoru), which agree on start and resolve times to within a few minutes:

#IncidentWindow (UTC, Sep 23)DurationBlast radius
1IPv6 networking issues in DFW6:23 → 12:15~5h 55mSubset of hosts in Dallas; apps on affected hosts unreachable or degraded
2Upstash unavailability in FRA12:44 → 14:03~1h 30mRedis-backed apps in Frankfurt; cache, sessions, queues down
3Elevated private-networking (6PN) errors14:47 → 15:23~36mCross-machine private traffic in some regions; .internal calls failing
4Partial Sprites outage, starting in SJC18:44 → Sep 24, 14:51~20hSandbox create/connect/manage errors; spread beyond SJC overnight

Two things stand out before we go deeper. First, the incidents chained across the day with barely a breather — DFW resolved at 12:15, FRA caught fire at 12:44, the mesh errors started at 14:47, Sprites failed at 18:44. Second, the day before had been a two-hour scheduled maintenance window, and as we will see, August had already been a rough month for Fly networking. September 23 did not come out of nowhere.

Three of the four were the network​

Fly.io's defining architectural bet is its private WireGuard mesh, called 6PN: every machine in an organization joins an encrypted IPv6 mesh, reachable over .internal DNS as if on one LAN regardless of physical region, with Flycast as the internal load balancer for private services. It is genuinely differentiating — no VPC peering tickets, no cross-region private-link pricing puzzles. But a mesh the tenant cannot see inside is also a mesh the tenant cannot debug, route around, or fail over. September 23 exercised that tradeoff three times.

The DFW incident was the big one: IPv6 connectivity issues on a subset of hosts in Dallas lasting nearly six hours, the longest single-region networking incident on Fly's recent record. "Subset of hosts" is the status page's honest phrasing for the tenant experience that matters — whether your app was down depended on which physical host your machines happened to land on, something you neither choose nor can see. Stateless apps with machines in multiple regions could shift traffic; single-region apps on an affected host simply waited from breakfast to lunch, UTC.

The afternoon's elevated private-networking errors were shorter (~36 minutes) but arguably scarier: error rates spiking on 6PN itself, the fabric every cross-machine call traverses. A region outage takes out machines; a mesh wobble takes out the paths between healthy machines — service-to-service calls, Flycast-routed internals, anything addressed over .internal. Thirty-six minutes is a blip on a quarterly SLA chart and an eternity when your app's internal RPCs are failing and there is no alternate route to configure, because the mesh has exactly one operator and it is not you.

And August says this was a season, not a fluke. Status mirrors record IPv6 networking issues (45 minutes, August 13), network issues in LAX (30 minutes, August 23), and increased packet loss (2 hours 10 minutes, August 28) — three networking incidents plus an end-of-month maintenance, all in the four weeks before September 23. Any one of them is unremarkable; periods of elevated network instability happen to every provider. The point for a tenant is cumulative: roughly monthly contact with incidents whose only available response is monitoring someone else's progress.

The one that wasn't Fly's network: Upstash in FRA​

At 12:44 UTC, half an hour after DFW resolved, Upstash services in Frankfurt went unavailable for about ninety minutes. Upstash is Fly's marketplace Redis partner — fly redis create provisions an Upstash instance wired into your Fly organization — and this incident is the purest illustration of the day's structural lesson, because two vendors' blast radii stacked.

A marketplace add-on inherits both vendors' failure modes while giving you neither vendor's controls. Upstash Redis instances on Fly cannot change their primary region or name after creation; your Frankfurt Redis lives in Frankfurt until you rebuild it elsewhere. During the FRA window, affected tenants could not fail over to another region's replica themselves, could not see whether the fault sat in Upstash's control plane, Fly's hypervisor layer, or the network between them, and could not do anything but wait for two companies' engineers to coordinate. A spring 2026 community thread about a previous Upstash-on-Fly outage is instructive about the genre: the eventual root cause was described as a bad interaction between a kernel upgrade and a hypervisor bug — layers no tenant of either vendor can inspect, let alone patch.

Redis is also rarely "just" a cache in these architectures. Sessions, rate-limit counters, job queues, feature flags, real-time pub/sub — when the Frankfurt instance goes dark for ninety minutes, every one of those becomes an app-level outage with its own shape. Teams that treat marketplace Redis as critical infrastructure without a second-region story or a degraded-mode plan discovered the gap at 12:44 UTC. The honest pre-mortem question is whether a managed add-on whose region is pinned at creation and whose internals are two vendors away counts as a highly available dependency at all.

The one still burning at midnight: Sprites​

The evening incident started at 18:44 UTC as a partial Sprites outage in SJC. Sprites is Fly's agent-sandbox product — persistent, hardware-isolated Linux microVMs with checkpoint and restore, sold at sprites.dev as ready execution environments for AI coding agents, including an MCP server for tool-call access. Sandboxes are the workload Fly is betting its next chapter on, which makes the timeline of this incident the day's most uncomfortable reading.

Per Fly's status page, the outage did not stay in SJC: a 23:59 UTC update notes the issue "now occurring on regions other than SJC." A 02:12 UTC update on September 24 says Sprites API performance "largely recovered" with most users no longer seeing errors creating, connecting to, or managing Sprites. Full resolution did not land until 14:51 UTC on September 24 — roughly twenty hours after it started. That makes the sandbox outage, the day's fourth and last incident, also its longest by a factor of three.

There are two uncomfortable implications. First, anyone running agents against Sprites as a production dependency — long-running coding sessions, scheduled agent jobs — had a twenty-hour window where the substrate could not reliably create or reach sandboxes, with the failure spreading across regions overnight rather than contained in one. Checkpoint-and-restore persistence is cold comfort when the API that manages the checkpoints is erroring. Second, the spread pattern (SJC first, then other regions) smells like shared control-plane or API-tier fate rather than one region's bad luck — the same class of failure that makes single-control-plane managed platforms feel fragile precisely when you most want region independence.

To be fair, Fly's transparency here deserves credit: timestamped updates through the night, an explicit acknowledgment when the blast radius widened, no euphemism about "a small number of users." This is status communication done right. It is also the ceiling of what transparency can buy you as a tenant: perfect visibility into a system you cannot touch.

What tenants could actually do (an honest accounting)​

Strip away the coping strategies and the per-incident tenant playbook for September 23 was thin:

  • DFW IPv6 (~6h): multi-region stateless apps could shift traffic away from Dallas; everyone else waited. Moving machines between regions mid-incident (fly machine clone --region plus DNS/proxy changes) is possible but slow, manual, and itself dependent on the platform's control plane behaving.
  • Upstash FRA (~90m): effectively nothing. Region-pinned instances, no tenant-driven failover, two vendors' internals between you and the fault. The mitigation had to exist before the incident: multi-region Redis, a degraded mode that survives cache loss, or a database you operate.
  • 6PN errors (~36m): nothing. There is one mesh, one operator, no alternate route table. Retries and circuit breakers limited the user-visible damage; nothing a tenant could do shortened the incident.
  • Sprites (~20h): retry, queue, wait — or have a second sandbox substrate ready, which almost nobody does for a workload most teams adopted in the last year.

Notice the shape: the shorter the incident, the less a response was even theoretically possible; the longer the incident, the more the missing capability was architectural (a second region, a second vendor, a second substrate) rather than operational. Nobody's runbook covers "the provider's private backbone errors for half an hour" because there is no runbook — there is only the status page and your retry policy.

The owned-infrastructure alternative does not magically prevent any of these failures. Hosts still lose IPv6, Redis still falls over, control planes still have bad nights. What changes is the response surface: your WireGuard mesh (plain WireGuard, Headscale, NetBird) has logs you can read and routes you can change; your regions are machines you picked, so failing over is a decision you make rather than a migration you request.

Your Redis runs with replicas you configured, so a Frankfurt-shaped event promotes a replica instead of starting a vigil; your incident timeline is written by the people fixing the problem, and inquiries get answers instead of updates. The price is real — you are the on-call, the capacity planner, and the postmortem author — but September 23 is a clean demonstration of what the managed price buys and, crucially, what it does not: it buys freedom from operating the system, not freedom from its failures.

The takeaway for your next deploy target​

Fly.io remains one of the best places to stand up a globally distributed app fast, and a bad day does not erase that. But September 23 belongs in every platform decision document alongside the happy-path benchmarks: four incidents, three of them networking, one marketplace dependency with no tenant-side failover, one twenty-hour sandbox outage that spread regions overnight — on top of an August that was already noisy. Evaluate a platform the way this day forces you to: ask what breaks when the mesh breaks, ask what your database add-on lets you do during its outage besides wait, and ask whether your most load-bearing new dependency (quite possibly an agent sandbox) has a second substrate or just a status page.

If those answers bother you, the alternative is not theoretical. Self-hosted platforms on machines you own — declarative fleets with your own mesh, your own regions, and stateful services you control — trade one set of hard problems for another, but "watch and wait" stops being the entire runbook. That is the whole thesis: own the substrate, own the timeline.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex