Five days and eighteen hours. That is how long OVHcloud's web hosting incident took to go from onset to full mitigation in the first half of 2026 — the longest single mitigation window IncidentHub flagged for any web hosting provider in six months, inside a half-year that produced 30,246 outages across 1,082 providers.
Not five minutes. Not five hours. Nearly six days during which every tenant on that hosting tier shared the same status page, the same incident commander, and the same queue for a fix — whether their own site was fully offline, intermittently degraded, or just slow enough to lose a checkout.
This is not a postmortem of OVHcloud's root cause. The status page said what status pages say. It is an accounting of what a 138-hour recovery window actually costs a tenant who picked one multi-tenant vendor as their only substrate, and what the same class of failure looks like when you own the fleet and set your own remediation priority.
The number that doesn't fit on a status page
IncidentHub's H1 2026 Cloud and SaaS Reliability Report, published July 20, watches public status pages and counts every provider-acknowledged degraded-performance or outage event until the provider marks it resolved. From January 1 to June 30 — 181 days — it recorded 30,246 incidents across 1,082 providers spanning cloud, developer tools, edge and CDN, IAM, AI and LLM, payments, observability, collaboration, and EdTech.
| H1 2026 census | Value |
|---|---|
| Collection window | Jan 1 – Jun 30 (181 days) |
| Providers with continuous coverage | 1,082 |
| Total incidents | 30,246 |
| Daily average | ~167 incidents per day |
| Busiest month | May: 6,070 incidents (20% of the half-year in one month) |
| Top category | Cloud providers: 4,723 |
| Runner-up | Developer tools: 4,589 |
That average — 167 chances per day for some provider to have a bad morning — is the background radiation of rented infrastructure. Most of those incidents are not yours. Until one is.
Inside that volume, the OVHcloud web hosting incident stands out on duration, not frequency. IncidentHub's tracking records its mitigation window at 5 days and 18 hours, or about 138 hours. The July 2026 Pagerly summary of the report and OpsMatters syndication both repeat the same structure: a small number of long-mitigation incidents sit on top of a very broad base of brief ones. OVHcloud's is the web hosting outlier that makes the base look short by comparison.
Two points of context matter before the cost math:
First, how IncidentHub counts. An incident opens when the provider acknowledges degraded performance or an outage on its status page and closes when the provider marks it mitigated. That is not a third-party probe declaring your site down. It is the vendor's own clock. Some tenants recovered sooner, some later; some were degraded rather than dark. The 5d18h window is what the provider needed to call its own platform fully healthy again. Your queue behind that work is exactly as long as their queue.
Second, OVHcloud has done long recoveries before at a different scale. On March 10, 2021, a fire that began at 00:47 CET in OVHcloud's Strasbourg SBG2 data center destroyed one hall, badly damaged part of SBG1, and required the other two halls (SBG3 and SBG4) to be switched off while firefighters worked. Millions of sites went offline; the company later described restoration as a "real nightmare," and a French class action for the incident eventually sought more than €10 million on behalf of 140 clients. The 2026 web hosting incident is not a fire and not a data-center loss — it is a control-plane and hosting-tier degradation — but the 2021 history is useful calibration: when a shared facility or shared control plane fails, tenant-level recovery is bounded by provider-level recovery, and that bound can be days, not hours. The mechanism changes; the queuing problem does not.
That queue is the cost that never appears on the timeline. The rest of this post puts numbers on it.
What 138 hours actually costs when you have no other substrate
A tenant with no failover plan experiences a multi-day incident as a single exposure: the hosting tier is degraded, and there is nowhere else to serve traffic from. The bill for that exposure has four layers — SLA breach, revenue at risk, support load, and SLA credit illusion — and only one of them shows up on the vendor invoice.
1. SLA math: how badly a six-day window blows any monthly budget
Monthly availability budgets are tiny compared to 138 hours:
| SLA target | Monthly error budget | 138 hours as a multiple |
|---|---|---|
| 99.99% ("four nines") | ~4.38 minutes | ~1,890x |
| 99.9% ("three nines") | ~43.8 minutes | ~189x |
| 99% | ~7 hours 18 minutes | ~19x |
Even a 99% monthly target — the kind of target a small team might informally accept for a low-tier hosting plan — is blown by nearly 19 times over. A 99.9% target is blown by almost two orders of magnitude. There is no interpretation of "five days of degraded hosting" that fits inside a single month's budget; the incident is the month.
Uptime Institute's Annual Outage Analysis Report 2026 adds a sobering baseline: across nine years of publicly reported outages, roughly two-thirds originated with third-party IT and data-center providers — cloud and internet giants, telecommunications firms, and colocation companies. The 138-hour window is an outlier on duration, but the kind of dependency that produced it — a third-party shared platform — is the statistically normal place for outages to come from.
2. Revenue at risk: a worked table
The most intuitive cost — lost revenue — depends on whether the site was fully offline or degraded (slower, intermittently erroring, losing a fraction of conversions). A simple model covers both:
- Offline — every hour at zero conversions.
- Degraded — a fraction of traffic still converts, but at a depressed rate.
Assume a SaaS or content-driven business whose revenue is roughly uniform across the month (an approximation; real businesses are spikier, which makes the variance worse).
| Monthly recurring revenue | Revenue per hour | At risk if fully offline for 138h | At risk if degraded (50% loss) for 138h |
|---|---|---|---|
| $1,000 | ~$1.39 | ~$192 | ~$96 |
| $10,000 | ~$13.89 | ~$1,917 | ~$958 |
| $50,000 | ~$69.44 | ~$9,583 | ~$4,792 |
For e-commerce, replace "MRR" with daily order volume times expected span — a shop doing $2,000 per day in orders that is offline for 5.75 days has ~$11,500 of orders that never happened, plus the downstream cost of restocking, refunds, and abandoned carts that convert elsewhere. These numbers are direct and obvious, which is why they are also the floor. The more expensive layers are indirect.
3. Support load and churn: the cost that never appears on the timeline
For every hour a hosting tier is degraded, a tenant's own support queue fills with a class of ticket the tenant cannot resolve:
- "Is the site down?" — answered with a link to someone else's status page.
- "Where is my order / post / upload?" — answered with "we're waiting on our provider."
- "Can you fix it?" — answered with "we're monitoring the situation."
A solo founder handling support themselves pays this in sleep. A two-person team pays it in the sprint that did not ship. A larger team pays it in overtime or in the churn of customers who interpret "we're waiting on our vendor" — accurately — as "we have no second substrate."
Churn is the quiet multiplier. If a degraded week pushes even 2% of a $10,000 MRR base to cancel, that is $200 per month forever until replaced, not just $1,917 once. The incident lasts six days; the cohort damage lasts quarters.
4. SLA credits: the number that compensates almost nothing
Every managed hosting SLA defines service credits for missed availability. They sound like compensation; arithmetically they are not. The Uptime Institute's own commentary on cloud SLAs uses a deliberately small example to make the point transferable: if a single virtual machine on a $3-per-month instance misses a 99% monthly target (down for more than about 7 hours), the credit at the 10% tier is $0.30. Scale that to a $50 or $200 hosting tier and the same pattern holds — a credit is a fraction of a fraction of one month's hosting fee, not a fraction of the tenant's business loss.
For the OVHcloud window, any credit tied to the hosting tier's monthly fee is bounded above by that fee. Even a 100% credit — the maximum any SLA offers and one almost never granted for a single incident — returns one month of hosting spend. It does not return 138 hours of revenue, support time, or churned customers. The SLA makes the provider whole on its own invoice; it does not make the tenant whole on the tenant's P&L.
That gap — between what the vendor remits and what the tenant lost — is why "we have an SLA" is not a recovery plan. A recovery plan is a second substrate.
Where the same failure mode goes on a fleet you own
A tenant on a single multi-tenant hosting tier experiences the 138-hour window as one event with one owner: the vendor. A fleet operator on owned bare metal — for example, a Cluster API fleet on Hetzner or OVHcloud bare-metal machines — experiences the same underlying failure classes as many small, local events, each with its own remediation path that the operator set ahead of time.
The mapping is not one-to-one, but it is instructive:
| Shared-hosting failure mode | Fleet-native analogue | Who remediates and how fast |
|---|---|---|
| Control-plane or host-OS degradation taking the tier degraded | Kubelet death, disk pressure, or network partition on a single node | MachineHealthCheck watches node conditions; the controller creates a remediation object and the provider reconciles it — minutes, not days |
| Capacity or scheduling stall (no room for new workloads) | Cluster Autoscaler / Cluster API MachineDeployment scale-out | Declarative: add a machine; the fleet provisions it from the owned pool |
| Storage or data-plane slowness | Local NVMe vs network volume latency; CSI behavior | Operator chose the substrate: etcd on local NVMe where wal_fsync_duration p99 stays under 10ms, not on a network volume that blows the threshold |
The key mechanism is remediation priority. On a shared tier, your incident is one of every other tenant's incidents. The vendor's incident commander triages across the whole population. You cannot fork the fix, reprioritize your tenant, or route around the degraded tier without infrastructure you built before the incident.
On a fleet you own, the incident commander is your own controller. Cluster API's machine-health machinery — the pattern maturing across recent cluster-api-provider-hetzner work on HCloudRemediation (#2054), retireConditions (#2227), and the Retire vs Reuse onExhaustion choice (#2144) — lets an operator declare:
- Which node conditions trigger remediation (not just "node is NotReady" but
DiskPressure,MemoryPressure, custom conditions). - How many reboots to attempt before giving up (
retryLimit). - What happens when reboots are exhausted — reuse the host or retire it and provision a fresh one.
Concretely, a kubelet that stops reporting, a disk that fills, or a network partition that isolates a node each follows the same loop without human intervention: detect the condition, cordon the node, reschedule its workloads, attempt remediation, and — if the host is genuinely failed — retire it and bring up a replacement via the MachineDeployment. The window is bounded by nodeStartupTimeout and the provider's provisioning time, typically single-digit minutes to low tens of minutes for a cloud VM and somewhat longer for dedicated bare metal. That is not zero downtime — workloads on the failed node still restart elsewhere — but it is not 138 hours of shared degradation either.
There is an honest limit to this argument, and the 2021 Strasbourg fire is the right place to name it. Bare metal can burn. A data center can be lost. Hetzner's own local NVMe is excellent for etcd's fsync tax (small sequential writes with fdatasync where p99 over 10ms is already a warning), but even excellent disks do not survive a hall fire. A fleet spread across a single building is still a single building. The structural win of a self-owned fleet is not "it never fails" — it is that its failures are uncorrelated with 1,081 other providers' incidents that same week and governed by a remediation priority the operator set, not a vendor's queue.
That distinction shows up operationally in how the fleet is tested. The teams that get value from MachineHealthCheck are the ones that have watched it fire — killing a worker's kubelet, partitioning a control-plane node, filling a disk — and asserted that remediation converges without a human, the pattern LitmusChaos's 2026 update and Cluster API v1.14's machine-remediation work now make straightforward to automate as a quarterly drill. A "self-healing" fleet that has never been chaos-tested is a promise; a fleet that has is a measured recovery budget.
Two smaller operational choices amplify the same point:
- etcd placement. etcd commits every write through
fsyncand documents hard latency expectations. A fleet operator who places etcd on Hetzner local NVMe and monitorswal_fsync_durationandbackend_commit_durationlearns the disk is wrong before the API server starts timing out. A tenant on a managed web hosting tier never gets to choose where the control-plane datastore lives. - Control-plane recovery. The management cluster that owns the fleet's
Machines,MachineDeployments, and provider secrets is itself a single point of concentration unless the operator has practiced restoring it — from an etcd snapshot, viaclusterctl moveto a standby management cluster, or by re-bootstrapping from GitOps. The same way tenant workloads need a second substrate, the fleet's brain needs one.
None of this makes a fleet free. A fleet has its own incident load — but it is your incident load, on a schedule you control.
The decision the six-day window forces
The value of a 5-day-18-hour incident to a prospective fleet operator is not the drama of the outage. It is the clarity of the question it forces: whose clock do you wait on when the platform you rent is the one that is degraded?
For a tenant on a single provider's web hosting tier, the answer is always the provider's clock. For a fleet operator on owned machines, the answer is a set of smaller clocks the operator defined: the MachineHealthCheck interval, the nodeStartupTimeout, the autoscaler loop, the backup and restore drill cadence.
A useful way to decide which side of that line you should be on is a four-question checklist. If you answer "no" to more than one, a six-day vendor window is not an abstract risk.
1. Do you have a second substrate you can serve traffic from without the degraded provider? That substrate does not have to be exotic. A second Hetzner region, a standby machine in a different availability zone, or even a cold standby that Cluster API can promote in minutes counts. "We would rebuild somewhere" is not a substrate; "we have MachineDeployments that can place workloads on a different pool" is. If the honest answer is "we have one place where the app can run," a multi-day provider degradation is a single point of failure with a six-day recovery objective you did not choose.
2. Can you reprioritize remediation for your own workloads, or do you wait in a shared queue? On a shared tier, every tenant's urgency is averaged into the vendor's incident priority. On a fleet, your controller's priority is your priority. If your checkout flow and your marketing site share a hosting tier, a fleet lets you reprioritize one over the other; a shared tier does not. The question is not "which is faster on average" but "who decides what is urgent when everyone is urgent at once."
3. Is your blast radius correlated with 1,081 other providers' incidents that week? IncidentHub's census exists to quantify dependency risk: 30,246 incidents in six months means any sufficiently connected application sits downstream of many status pages. A self-owned fleet does not opt out of failures — it opts out of correlated failures. Its nodes fail for local reasons (disk, kernel, network) on a schedule unrelated to Fastly's 161 H1 2026 incidents, GitHub Actions' 37, or OVHcloud's web hosting window. When the internet has a bad May (6,070 incidents in one month), uncorrelated is the cheapest kind of reliability to own.
4. Have you measured your own recovery, or are you assuming it?
A fleet that has never killed a kubelet on purpose, filled a disk to DiskPressure, or restored its management cluster from an etcd snapshot does not know its recovery time. It knows its intended recovery time. The teams that benefit most from owning the fleet are the ones that treat remediation like any other feature — tested quarterly with a chaos suite, timed, and documented — before the 3 a.m. page.
If those answers point toward owning the substrate, the shape of the platform matters as much as the decision to own one. That is where Bex.co sits.
Bex is an open-source, AI-native alternative to Render that runs as a Cluster API fleet on machines you own — typically Hetzner bare metal or cloud VMs, including OVHcloud bare metal where that is the right substrate. The fleet owns machine lifecycle declaratively: you describe the desired machines, and controllers reconcile toward that state. Adding a second machine when the first is full, replacing a failed host when remediation retires it, and spreading control-plane etcd across local NVMe are not tickets to a vendor's queue. They are reconciliations your own control plane runs. The hosting tier that OVHcloud's incident degraded for 138 hours has a fleet analogue that degrades one node at a time and heals in minutes — not because bare metal is magic, but because the remediation loop is local.
That is a deliberately narrow pitch. Bex does not try to bundle your database onto the same bill the way some managed PaaS products do — bring your own managed Postgres (Neon, Supabase, RDS) or self-host it on the same fleet — and it does not promise that owned hardware never fails. What it does promise is that when hardware does fail, the wait that follows is yours to bound, measure, and improve.
Thirty thousand incidents in 181 days is not a forecast. It is a census. Rented infrastructure fails constantly, visibly, and on other people's clocks. The teams that sleep best in that environment are not the ones whose providers fail least. They are the ones who picked a substrate where the longest wait they can face is the one they set themselves.
Sources: IncidentHub H1 2026 Cloud and SaaS Reliability Report (July 20, 2026) and its OpsMatters syndication; Pagerly summary of that report (August 3, 2026); Uptime Institute Annual Outage Analysis Report 2026 (May 13, 2026) and cloud SLA commentary; OVHcloud Strasbourg fire coverage via Reuters, The Register, and Data Center Dynamics; CAPH remediation issues #2054, #2227, #2144. Census figures and the 5d18h window are as reported by IncidentHub; per-tenant impact varies by how degraded the hosting tier was for each site.