At 08:39 UTC on July 2, 2026, an engineer at Railway disconnected a network carrier at one US East availability zone — and accidentally removed the site's last route to the internet. For the next twenty minutes, every tenant behind that route went down together. Then came the stranger part: after routing was restored, storage ran at a third of its capacity for nearly two hours while the cluster reported itself healthy, and roughly 20,000 private-network links sat silently blackholed until someone restarted the mesh agents across the whole zone. The outage, from first packet loss to resolution, ran 07:44 to 12:01 UTC.
That cascade is worth understanding in detail, because it is the clearest public illustration this year of what shared fate costs on a multi-tenant platform. And it was not an isolated bad day. It was the fifth major disclosed Railway outage since November 2025 — each in a different subsystem, each one no customer could see coming or do anything about. The pattern, not the postmortem, is the signal worth tracking.
Twenty minutes with no route to the internet
Railway's official incident report is candid and detailed, and the timeline is the core of the story. All times UTC, July 2:
- 07:44 — Packet loss observed in US East, traced to one upstream Tier 1 carrier whose backbone degradation was spilling traffic onto the path between Railway's US regions. Railway's internal probes had caught the degradation nearly two hours before it touched user traffic. A public incident was declared.
- 07:44–08:32 — Railway disconnected from the degraded carrier at all US border routers. Most US paths recovered. US East did not fully, because a secondary carrier there was handing return traffic back through the degraded network.
- 08:39 — Railway disconnected that secondary carrier at the affected zone too. Unknown to the team, it was the only carrier still supplying the site's default route. The zone is a first-generation site: unlike Railway's newer datacenters, which generate their own default route at the border, it depends on carriers to supply one — and the one remaining carrier's connection does not provide it. The site lost its route to the internet for roughly twenty minutes. Failed connections in and out of US East; private networking heavily degraded.
- 08:59 — The secondary carrier was reconnected. Routing stabilized. And this is where a routine carrier incident became a platform outage.
With the fabric's default route gone, every server's OS had fallen back to the only route left: the slow internal management network. A default Linux behavior let servers answer for their storage addresses on that network too, so storage traffic began flowing over a path with a small fraction of the fabric's capacity.
Connections don't re-check their route once established — so the storage connections created during that twenty-minute window stayed pinned to the management network after the fabric recovered. Routing tables looked correct. The storage cluster reported healthy. Throughput sat at roughly a third of capacity, I/O wait at 58%, for nearly two hours, until engineers found connections coming from management-network addresses and terminated them across every host in the zone. They reconnected over the correct fabric within seconds; I/O wait dropped under 5% in about fifteen minutes.
Meanwhile the private-networking mesh had captured its own bad state. Railway's tunnels learn peer addresses from received packets, and during the disturbance traffic funneled through a device that rewrites source addresses — so thousands of tunnels learned the wrong peer address and kept it. The mesh only re-verifies addresses on membership change, and idle tunnels never send the packet that would fix them.
At peak, about 20,000 host-to-host links were blackholed, including inter-region traffic into US East. Recovery required a fleet-wide restart of the mesh agents at 11:49. Full resolution at 12:01.
Railway's own summary of the mechanism is the sentence to remember: stale connections "captured a bad path during a brief window of instability, and held onto it after the network recovered, because nothing in the system re-asserted the correct state."
What "one ISP having a bad day" costs when the fate is shared
A carrier degradation is routine internet weather. Railway connects every Metal datacenter to at least three Tier 1 carriers precisely so one can fail safely — and the multi-carrier design did its job. Everything after 08:39 was self-inflicted: an older site design that depended on carriers for its default route, a disconnection made without checking what routing would remain, and dependent systems that latched onto bad state silently.
But notice what that means from a tenant's side of the boundary. You could not see any of it. Routing tables were correct, the storage cluster said healthy, and your services were slow or unreachable for reasons no dashboard available to you would explain. If your monitoring showed elevated latency that morning, you spent the first hour debugging your own application — the wrong system entirely.
There was no region to fail over to, no configuration to change, no fix to apply. The repair required terminating stuck connections on hosts you can't reach and restarting mesh agents across a fleet you can't touch. Your incident response was refreshing the status page.
That is the blast radius of a shared control plane and shared data path: one zone's bad twenty minutes became every tenant's degraded storage and broken private networking for hours, with the failure invisible to everyone affected. Restoring the network did not restore what depended on it — and the dependency graph lives entirely on the vendor's side.
Five outages, five subsystems, seven months
July is the one to study technically. But the evaluation signal is the rate. Since November 2025, Railway has published five major incident reports, each in a different layer of the stack:
| Date | Failure | Impact |
|---|---|---|
| Nov 20, 2025 | GitHub webhooks dipped, then surged 10x, flooding the deploy queue; workers locked up under memory pressure | All deploys delayed ~2.5h; Free, Hobby, then Pro deploys disabled in stages. Running workloads unaffected. |
| Feb 11, 2026 | An automated abuse-enforcement rule misclassified legitimate services during a staged rollout | Under 3% of the fleet received SIGTERM, including Postgres and MySQL services — while the dashboard kept showing terminated workloads as active. |
| Mar 30, 2026 | A CDN config change enabled caching on domains that had it explicitly disabled | ~0.05% of domains for 52 minutes; GET responses without an explicit Cache-Control header could be served to a different user than the one they were generated for. |
| May 19–20, 2026 | Google Cloud incorrectly suspended Railway's production account | Platform-wide blackout of roughly eight hours. API, dashboard, control plane, builds, deploys, databases down; once cached routes expired, workloads on Railway Metal and AWS returned 404s despite containers still running. |
| Jul 2, 2026 | Upstream ISP degradation, then a carrier disconnection that removed a zone's last default route | ~20 min with no internet route; storage at a third of capacity for ~2h; ~20,000 private-network links blackholed. |
One outage is a bad week. Five failures in five unrelated parts of the stack — deploy queue, abuse automation, CDN edge, a cloud-provider account, network/storage/mesh — is a pattern, and it changes what you can plan for. You are no longer hardening against a known weakness. You are waiting to discover which layer fails next, knowing that in every case so far the information needed to diagnose it and the controls needed to fix it sat on the other side of a boundary you cannot cross.
May is the one that should change a risk model. Railway runs workloads across its own Metal datacenters, AWS, and GCP — and the edge proxies resolved routes through a network control plane API hosted inside the suspended GCP account. The mesh held for about an hour on cached routes, then caches expired and healthy containers became unreachable 404s. Multi-cloud bought nothing, because route discovery sat on every path. As The Register and The Stack reported, the trigger was an automated suspension of an account Railway says costs around $2M a month — suspended without warning, restored hours later. Your compute being healthy is not the same as your compute being reachable, and reachability was never yours to control.
The departures corroborate the pattern rather than proving it. SoundBoost.ai moved to a Hetzner dedicated server after May at roughly half the cost — and reports lower latency to its mostly-US customers from Germany than it got from Railway's US East, a reminder that a closer region with inconsistent behavior loses to a farther server with a steadier path. Every moved its products to Render after downtime "every few days" on Railway versus one total incident on Render. A B2B team migrated to Azure mid-outage on May 19 itself, back up in hours — possible only because its database had never lived on Railway. That last detail is the actionable one, and we'll come back to it.
The same internet weather on machines you own
Here's the honest comparison the TODO framing demands: a carrier degradation hits your upstream too if you self-host. Internet weather is shared. What differs is everything after the first packet drops.
On an owned fleet — say, a few Hetzner machines behind your own routing — a Tier 1 degradation is your incident to see and your incident to fix. Your probes catch it, your routers reroute it, and the blast radius is exactly your fleet: no other tenant's storage storm shares your fate, no mesh restart across someone else's hosts gates your recovery. There is no first-generation site whose default-route design you didn't choose. If a bad path gets captured somewhere, it's in a config you can read, on hosts you can reach, with logs you own. The July outage's cruelest property — everything downstream of the vendor boundary reporting healthy while running at a third of speed — is a property of opacity, not of networks.
The honest limits cut the other way, and a fair accounting names them. You own the 3am page. A single server is single-homed unless you design otherwise — multi-carrier redundancy at Railway's level is not what a default dedicated server gives you, and diversifying transit is real work.
And five postmortems notwithstanding, Railway's incident reports are genuinely good: detailed, candid, public. Most teams evaluating a host will never get this clear a view into anyone's failure modes. The question is never "does this infrastructure fail" — it's "when it fails, can I see it, can I act on it, and does anyone else's bad day become mine."
The audit to run this week
Whether you stay or leave, the July outage plus the four before it suggest a concrete evaluation, and it takes an afternoon:
- Pull the status-page history and count. Not the latest postmortem — the rate. Five major disclosed outages across five subsystems in under eight months is the number that matters, and it answers a different question than any single root cause: how often will I be a spectator at my own incident?
- Ask what you could have done during each one. For all five Railway incidents, the answer is nothing customer-side: no failover to trigger, no region to shift to, no config that would have helped. In February the console actively misled; in May the console was gone; in July everything looked healthy at a third of speed. If that is an acceptable position for a workload, a hosted PaaS remains a reasonable choice. Decide explicitly per workload, not once for the company.
- Inventory what exists only inside the platform. Environment variables set as literals, internal hostnames in connection strings, volume data you can only dump rather than copy, cron defined in the console rather than the repo, DNS and TLS held at the vendor edge. The mid-outage Azure migration worked because the database already lived outside Railway. Every item on that list is both a migration task and, today, a single point of failure — and exporting secrets plus de-hardcoding internal hostnames are no-op changes worth making while you stay.
- Price your escalation path. Below Railway's $5,000-a-month Business Class tier there is no contractual response time; Pro gets a private thread typically within 72 hours with no SLO. Know which tier your revenue runs on before the next incident, not during it.
Outage count over time is the signal because any single postmortem can be explained and any single fix can be shipped. Railway is migrating first-generation sites to self-generated default routes, fixing the management-network fallback, and adding the missing alerts — each a real fix for a real link in the July chain.
What no fix changes is the architecture the table reveals: a platform where the deploy queue, the abuse automation, the CDN edge, a cloud account, and the network fabric can each, independently, take every tenant down together, invisibly, with nothing for the tenant to do but wait. That is what the fifth outage since November actually reveals. Plan accordingly.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



