Railway spent three years calling four regional proxies an "edge network." Then, in June 2026, it announced Hikari: a from-scratch CDN with 60 points of presence, 180 nodes, and a headline number — 30 million requests per second — absorbed during a real DDoS attack. Built in 30 days, in Rust, with a WebAssembly dataplane that upgrades itself without dropping a packet.
The headline is fun. The architecture is the actual gift. Railway published the whole design — the control plane, the BGP automation, the per-request WASM guest, even the incident where it went wrong — and almost every pattern in it transfers to a self-hosted fleet on machines you own. This post reads the engineering writeup as a parts catalog: what to steal, what to skip, and what the bandwidth math says about building your own edge instead of renting one.
The numbers, honestly stated
Railway's author, Phin Walton, is unusually careful about what "30M RPS" means, so let's receipt it before anyone copies the headline into a pitch deck:
| Number | What it actually is |
|---|---|
| 30M RPS | Sustained during real DDoS attacks, zero disruption |
| ~1M RPS | Ordinary peak-hour traffic |
| ~150M RPS | End-to-end benchmark ceiling under ideal conditions |
| 60 POPs | At launch, with ~40% still being delivered |
| 180+ nodes | 16-core EPYC, 256 GB RAM, 8 TB NVMe, 100G networking each |
| Tens of terabits | Total network capacity |
| 30 days | Build time for the software, not the procurement |
Two honesties matter here. First, the 30-to-1 gap between absorbed-attack traffic and daily peaks is the whole point of owning capacity: the fleet is sized for the worst day, and the other 364 days are headroom you already paid for. Second, "30 days" covers the software build — the datacenter leases and hardware procurement started months earlier. Anyone promising you a global CDN in a month without mentioning the purchase orders is selling the demo, not the deployment.
Why they built it — and how 30 days was possible
Railway tried buying first. Off-the-shelf CDN products failed on two counts: Railway's internal velocity outpaced the vendors' ability to ship features, and Railway's users have non-standard connection behaviors the vendors were never engineered to support. But the deciding number was a support metric: network-related issues made up around 20 percent of all ticket volume — CDN configuration, nameservers, the usual. Owning the stack meant owning the fix for a fifth of support load.
The deeper reason is structural. A platform that operates both the edge and the origins has a routing superpower no standalone CDN can match: it knows exactly where your workload runs, so it never has to guess where the origin is. Traditional CDNs guess with muddy signals — an origin 10 ms away on layer 3 might proxy an asset living 200 ms away. Railway also ships cache defaults other CDNs won't: it respects your HTML Cache-Control and automatically purges HTML cache on every deploy, where Cloudflare won't cache HTML or JSON by default at all. That default is only safe because the CDN and the deploy pipeline are the same company.
So how do you build that in 30 days? The writeup names the enablers, and they're all about sequencing:
- Hardware first, software second. POP procurement ran months ahead of the build. Day one of the 30 days started with racks, not RFPs.
- Deployment designed before the dataplane. The team explicitly engineered in reverse: with 180 nodes serving a million RPS, the upgrade story had to exist before the thing being upgraded. They'd been burned by Ansible's lossy, hard-to-audit behavior on the old proxies and refused to repeat it.
- One-command node bootstrap. Onboarding a node means running an authenticated
curlthat installs thehikari-keepersidecar, which phones home to theedge-cpcontrol plane and reconciles itself into service. - Reconcile loops everywhere. Node config, BGP announcements, dataplane versions — every layer converges toward declared state on a timer, so a lost message is a delay, never a stuck fleet.
The self-hosted takeaway from this section isn't "be fast." It's that Railway's 30 days were possible because the boring work — capacity, bootstrap, convergence — was either pre-done or designed first. A three-node cache tier on Hetzner gets the same benefit at 1/60th the scale: provision the machines before you need them, and write the bootstrap script before you need it.
Four patterns worth stealing
Hikari splits into four services: edge-cp (control plane), hikari-keeper (per-node sidecar), hikari (the dataplane that terminates TLS/HTTP2 and manages cache), and hikari-guest (a versioned WASM binary invoked per request for cache/WAF/limit decisions). Each split exists to solve a problem self-hosters also have. Here's the transfer table:
| Hikari pattern | Problem it solves | Self-hosted equivalent | Verdict |
|---|---|---|---|
hikari-keeper reconcile steps in Rust | Ansible is lossy, slow, and scary at fleet scale | Cluster API + GitOps, or a small reconcile agent of your own | Steal the shape. You don't need their Rust; you need "converge on a timer, audit every step." |
| WASM guest per request (wasmtime) | Dataplane rule changes without dropping connections | wasmtime/Wasmer sidecar, or Caddy/Nginx with reloadable config for smaller scale | Steal selectively. The per-request WASM + thread-local pool design is gold above ~100k RPS/node; below that, graceful reloads are fine. |
| Aggressive cache defaults + purge-on-deploy | Stale HTML vs. no caching at all | Surrogate keys + deploy-hook purge in Varnish, Caddy, or your app gateway | Steal immediately. This is the highest value-per-line item in the whole project. |
| Edge + origin under one operator | Guessing where origins live | Run cache and workloads on the same fleet/network | You already have this. Any self-hosted PaaS co-locates cache and compute — it's the default, not a feature. |
Two details deserve a closer look. First, the WASM design: wasmtime cold-instantiates a guest in about 10 microseconds, which Railway still considered too slow for the hot path. Their answer is a thread-per-core model — each worker is a single-threaded Tokio runtime pinned to one core, no work-stealing — with a thread-local pool of warm guest instances per live version. Acquire and release are a pop_front and push_back on a deque no other thread ever touches; the kernel pins connections to workers with SO_REUSEPORT plus NIC RSS. Multiple guest versions stay loaded simultaneously, so a rollout swaps the per-request brain underneath a dataplane that holds connections open forever. That is the correct shape for zero-downtime rule changes, and it runs just as well on one owned box as on 180.
Second, the purge-on-deploy default is the pattern with the best effort-to-impact ratio for small teams. Most self-hosted setups pick one of two bad options: cache HTML aggressively and serve stale deploys, or don't cache HTML at all. Railway's third option — cache it, and invalidate it from the deploy pipeline — needs exactly one webhook between systems you already operate. If you steal one thing from Hikari this week, steal this.
What not to copy: global anycast (and the incident ledger)
The section of Railway's writeup that should make a small team nervous is BGP anycast. BGP was never designed for anycast; anycast is a side effect of the protocol. BGP doesn't understand geography or latency — it sees autonomous systems, and a user one ASN away from your prefix might be routed on a world tour through a transit network you only peer with on another continent. Railway fights this with dynamic BGP community tagging ("don't export this route outside Japan"), reconciled every 5 seconds — and still names outlier ISPs like Liberty Global that expose no communities at all, likely to force content networks into paid transit.
The verdict for a self-hosted fleet: do not build your own global anycast network. The failure modes are measured in continents, the debugging tool is a looking glass, and the fix is often a commercial relationship with a transit provider. Cheaper alternatives cover 95% of the value: put a managed anycast front door (Cloudflare's free tier, Bunny) in front of origins you own, or serve two to three regions with geo-DNS and let the cache tier do the real work. Anycast is the one Hikari subsystem where Railway's scale is load-bearing, not incidental.
The second caution comes from Railway's own incident ledger. In March 2026 — three months before the Hikari announcement — a CDN configuration update enabling surrogate keys accidentally enabled caching on ~0.05% of domains that had it opted out. For 52 minutes, GET responses without explicit cache headers, including potentially authenticated content, were served from edge cache to the wrong users. Origin Cache-Control was respected and Set-Cookie responses were never cached, which bounded the blast radius — but the window still happened.
Railway's mitigations are your checklist, free of charge:
- Caching is opt-in per domain, with
Authorization-bearing requests bypassing cache andSet-Cookieresponses never stored. These two rules are what kept a 52-minute incident from being a catastrophe. - Behavior tests gate every cache-config change — tests for what must and must not be cached.
- Roll out cache changes over hours, sharded, not minutes fleet-wide. The incident's root cause was a config push; the fix was a slower push.
Notice the shape of this failure: the team that built a 150M-RPS benchmark machine still got bitten by a config default. Cache safety is a policy problem wearing an infrastructure costume, and the policies above cost nothing to adopt on day one.
The money math
Here's the question the whole project poses for self-hosters: what does "don't rent a third-party edge" cost? The answer starts with bandwidth, where the price spread in 2026 is enormous:
| Egress path | Effective $/GB | 10 TB/month costs |
|---|---|---|
| AWS CloudFront | ~$0.085 | ~$850 |
| Bunny CDN | ~$0.005 | ~$50 |
| Hetzner dedicated (20 TB included/server) | ~$0.0013 effective | $0 marginal within allowance |
At 10 TB a month, serving bytes off your own Hetzner machines instead of CloudFront saves roughly $800 every month — before you count a single cache hit. And cache hits are the multiplier: every request your Varnish/Caddy tier absorbs is a request your app servers never see, which means the cache tier pays for its own hardware by shrinking the compute tier. Railway's 8-TB-NVMe nodes are sized so the working set lives on local flash; a self-hosted cache node with even 2 TB of NVMe holds the entire static corpus of most small platforms in memory-mapped comfort.
The honest accounting has to include what you don't get: no 60-POP map, no DDoS absorption at 30M RPS, no one else's NOC. But weigh that against what a typical self-hosted PaaS actually serves. At tens of millions of requests a month rather than per second, two cache nodes in different Hetzner regions with NVMe, a deploy-hook purge, and a managed DNS front door deliver Hikari's behavioral wins — fast cache hits, fresh HTML after deploys, no per-GB meter running — for the price of hardware you were already renting. Railway proved the ceiling is 150M RPS per well-built fleet. You don't need the ceiling. You need the shape.
Copy the control plane, rent the map
Hikari's real lesson isn't that every platform should build a CDN. It's that the valuable half of a CDN — reconcile-loop operations, versioned dataplane logic, purge-on-deploy caching, edge-and-origin co-design — is separable from the expensive half, which is a global anycast map with 60 leases and a BGP team. Railway needed both halves. A self-hosted fleet on owned hardware needs the first half and can rent the second for the price of a free-tier account.
So steal the control plane. Write the bootstrap script before you need it. Put your cache rules in a versioned artifact with a slow rollout. Purge from your deploy pipeline. And leave the looking glass to the people with an ASN.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



