In July 2026, somebody counted open roles on three PaaS careers pages and the numbers told a story that was almost too neat: Railway was hiring for 18 roles, Render for 22, and Fly.io for zero. The takes wrote themselves. Railway and Render were healthy and growing; Fly.io, which was also sunsetting GPU support that same quarter, was obviously winding down. Then September arrived and every number moved — Railway dropped to 8, Render climbed to 37, Fly.io reappeared with 3 — and a funny thing happened to the neat story: the vendor doing the most hiring had just come off five major outages in six months, while the vendor hiring nobody had just raised $25 million. Here is the verdict first: a careers page is a snapshot of recruiting pipeline, not a reliability instrument, and reading vendor health off headcount gets the answer backwards in both directions.
| Vendor | July 2026 snapshot | September 2026 re-check | Headline reliability event, same window |
|---|---|---|---|
| Railway | 18 open roles | 8 open roles, 5 of them infra/platform | Five major incidents Nov 2025–May 2026, including an 8-hour platform-wide blackout |
| Render | 22 open roles | 37 open roles across product, infra, and agents | Steady major/critical status history, including a critical Oregon disruption on Jul 24 |
| Fly.io | 0 open roles | 3 open roles (infra ops, networking, strategic accounts) | GPU sunset on Aug 1 — plus a $25M raise and a new CEO in July |
Treat the July column as dated on purpose: the numbers decayed within about eight weeks, and that decay is the thesis. A signal that inverts this fast was never a signal. The rest of this post reads each vendor's hiring against its own incident and deprecation record, then gives you a six-signal checklist that actually answers "should I stay or should I migrate."
Railway: hiring fast, breaking anyway
Railway entered 2026 as the growth story of the three — and as the outage story. Between November 2025 and May 2026 the platform suffered five major incidents: a GitHub webhook surge in November, a cryptominer exploit causing CPU starvation in December, a multi-day DDoS compounded by a Cloudflare BGP outage in February (Railway's own incident report covers Feb 18–21), a CDN misconfiguration that exposed user data in March, and the May 19–20 GCP account suspension that took the whole platform down for roughly eight hours.
The May outage deserves one paragraph because its anatomy matters more than its duration. Google Cloud's automated systems suspended Railway's production account — a customer reportedly spending about $2 million a month on GCP — with no prior outreach. But Railway runs workloads on its own Railway Metal hardware and on AWS as well as GCP, while the network control plane lived on GCP alone.
So as cached routes expired, unreachable-everywhere spread to machines Google never touched. Railway's incident report owns the architectural decision plainly: one upstream provider action should never have been able to cascade platform-wide. (The full exit fallout is covered in our post-Railway destinations piece; this post is about what the hiring numbers meant while it was happening.)
Now look at the hiring column with that record in mind. Eighteen open roles in July did not buy reliability — the outages happened while the company was staffed to grow. And the September re-check reframes the number further: Railway's board is down to 8 roles, but 5 of the 8 are infrastructure and platform positions — datacenters, storage, observability, baremetal orchestration. That is not a company coasting; it is a company hiring directly into its outage record, consistent with a summer reliability program that cut over to a distributed network router and spread database quorum across Metal and AWS. Headcount direction told you nothing; headcount composition told you where the pain was. A careers page read as "18, growing, healthy" missed both halves of that.
Fly.io: hiring nobody, not dying
Fly.io's half of the July snapshot looked like an obituary: zero open roles in the same quarter the company fully deprecated GPU Machines, unavailable after August 1, 2026. If you squint, "hiring freeze plus product sunset" reads as "winding down." It was, in fact, a refocus — and the deprecation details matter for judging how the company treats customers on the way out of a product.
Scope the sunset fairly and it looks narrow, not existential. GPUs were an experimental surface with a small user base; affected customers were contacted directly, docs carried the banner for months, and community staff said so outright. The deprecation was arguably too quiet — one forum regular admitted to being unnerved that the announcement thread was unlisted — but "quiet sunset of a side product with direct outreach" is a deprecation-behavior data point, not a shutdown signal. Note the methodology choice here: Fly's risk surface shows up in deprecations and focus pivots, not in a count of status-page entries. Fly files incidents per region, so raw totals run into the hundreds per quarter and mean nothing next to another vendor's platform-wide filings; platform-wide events are the only apples-to-apples comparator, and Fly's headline story this quarter was the GPU exit, not an outage record.
Meanwhile the "zero" turned out to be a trough between strategies, not a flatline. In July 2026 Fly.io raised $25 million to build "computers for agents", brought in Scott Johnston as CEO with founder Kurt Mackey moving to the board in an advisory role, and centered the company on long-running stateful compute for AI agents under its Sprites product. By September the jobs page listed 3 roles — infrastructure operations, proxy/networking engineering, and a strategic-accounts hire — which is exactly the shape of a funded refocus: keep the fleet healthy, fix the routing layer agent workloads churn hardest, land bigger accounts. Anyone who read "zero open roles" as "migrate now" in July would have migrated away from a vendor weeks before it got funded to double down on its core product.
Render: the control group that breaks naive counting
Render is the vendor the naive theory should predict best — the most hiring, hence presumably the most reliable — and it is the one that breaks naive counting entirely. Open roles roughly doubled from 22 in July to 37 by September, spanning product, infrastructure, data, and a visible agent cluster (agent auth, agent observability, agent experience). Hiring acceleration is real here. Reliability data is also real, and it is voluminous: Render's public status history runs to hundreds of filed incidents across 2026, with recent entries including degraded builds and deploys on September 2, service instability in Singapore on August 31, and a critical service disruption in Oregon on July 24.
Does that mean Render is less reliable than Railway? You cannot conclude that from the counts, and understanding why is the most transferable skill in this post. Status pages are not calibrated instruments: vendors differ in filing granularity (one entry per affected region versus one per underlying event), in severity thresholds, and in how aggressively minor degradations get posted. Render files per region, so a single bad day can mint a dozen records; Railway's record is dominated by a handful of platform-wide events with narrative incident reports. Comparing 500 filed entries against 5 tells you about filing culture, not uptime. The honest read is narrower: Render had a critical region-wide disruption in July and repeated build/deploy degradation through August and September — real marks against it — while hiring fastest of the three. Growth and incidents coexist, again. Anyone using headcount as a reliability proxy would have picked Render as the safe harbor without ever opening its status page, which is precisely the failure mode.
The 6-signal vendor-health checklist
One signal misleads; six triangulate. Run this in about thirty minutes before you bet production on a vendor — or before you panic-migrate off one.
- Platform-wide events, not incident counts. Pull the status page (Render and Fly.io expose Statuspage APIs; Railway publishes narrative incident reports) and count only full-platform or region-wide events over the trailing two quarters. Ignore per-region minor filings when comparing vendors — granularity differs too much. Ask: did the control plane ever go down? Did workloads that were healthy become unreachable?
- Deprecation behavior, not deprecation existence. Every vendor sunsets products; judge how. Check the notice period, whether affected customers were contacted directly, whether docs carried the warning, and whether a migration path existed. Fly's GPU exit scores well on direct outreach and poorly on visibility — that split verdict is more useful than "they deprecated something, run."
- Funding and leadership events. A raise, a CEO change, or a strategy pivot within the last two quarters reframes every other signal. Fly's $25M plus new CEO in July inverted the "zero roles" reading within weeks. Check press releases and the company blog before concluding anything from headcount.
- Where the hires go, not how many. Open the actual listings. Five of Railway's eight open roles being infra/platform (datacenters, storage, observability, baremetal) says "reliability investment"; a wall of growth-marketers-and-AEs says "revenue push." Also check posting dates: evergreen listings from 2024 sitting next to fresh 2026 infra roles tell different stories about urgency.
- SLA and outage-posture fine print. Is there a contractual SLA on your plan, or only on enterprise? Are backups and data exports reachable during a control-plane outage, or do they go down with the dashboard? Third-party roundups flag Railway's lack of a standard-plan SLA and outage-inaccessible backups — verify the current docs yourself, because this is the term that decides whether an incident is a bad day or a hostage situation.
- Exit cost and single-dependency architecture. Could you leave mid-outage? Teams whose databases lived outside Railway migrated during the May blackout itself; Railway-native data had no such option. And ask what the vendor itself depends on: Railway's GCP-hosted control plane taking down Metal and AWS workloads is the canonical warning — your vendor's single point of failure is your single point of failure.
Score the vendor across all six. Any single red column — including a careers page at zero — is a prompt to investigate, never a verdict to migrate on.
Conclusion: own the signal
Step back and the pattern is almost symmetrical. Railway hired the most and broke the most; Fly.io hired nobody and got funded; Render hired fastest while filing incidents by the hundred. Headcount predicted none of it, because headcount measures recruiting pipeline and strategy phase — growth push, post-pivot trough, reliability rebuild — not the probability your deploy succeeds tonight. The signals that actually predicted customer pain were architectural (a control plane with one upstream), contractual (no SLA, backups behind the dashboard), and behavioral (how sunsets get communicated). None of those live on a careers page.
There is a reason this genre of tea-leaf reading exists at all: on a managed PaaS, the vendor's internals are opaque, so customers divine health from whatever is public — job boards, funding news, forum vibes. Self-hosting inverts that. When the machines, the control plane, and the incident history are yours, there is no signal to divine; there is a dashboard to read and a runbook to run. That is the deeper lesson of the 18-vs-zero summer: the teams with the least anxiety about vendor health were the ones who had stopped renting it.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



