The most consequential Kubernetes announcement at KubeCon EU 2026 was not a faster scheduler or a new GPU API. It was a backup tool changing owners. In front of more than 12,000 attendees in Amsterdam, Broadcom handed Velero — the Kubernetes-native backup, restore, and migration tool running in production clusters everywhere — to the CNCF as a Sandbox project, moving it out from under single-vendor control and into neutral governance.
Backup-tool governance does not usually make headlines. This one did because of who the donor was. Broadcom's stewardship of the VMware portfolio has spent three years teaching platform teams to flinch, and Velero sat squarely inside the blast radius. The donation is Broadcom's answer to the trust deficit — and for teams running their own fleets on machines they own, it reopens a question many had quietly shelved: is Velero safe to bet a disaster-recovery story on?
TL;DR — what this buys you:
- The governance win is real: Velero now lives under CNCF oversight with maintainers beyond Broadcom, and its roadmap can no longer be redirected by one vendor's earnings call.
- What Sandbox does not guarantee: maturity. Sandbox is the CNCF's entry tier, not a stamp of production-readiness — Velero earns that from a decade of production use, not the donation.
- The minimal setup on owned hardware: Velero plus an S3-compatible backup target, a nightly schedule, and a restore drill you actually run. The whole blueprint is one section below.
- Who should wait: nobody needs to wait on governance anymore. Wait only if your storage layer cannot give Velero what it needs — details in the verdict.
From Heptio to the CNCF in 90 seconds
Velero's ownership chain reads like a decade of Kubernetes history compressed into one project. It started at Heptio, the startup founded by Kubernetes co-creators Joe Beda and Craig McLuckie, where it was built as the missing piece every production cluster eventually needs: a way to back up not just persistent volumes but the cluster state itself — deployments, configmaps, secrets, CRDs — and restore all of it onto a different cluster. VMware acquired Heptio in 2019 and folded Velero into the Tanzu portfolio. Then Broadcom acquired VMware for 61 billion dollars in 2023, and Velero became, technically, a Broadcom product.
That last hop is where teams started to worry, for two concrete reasons. First, Broadcom overhauled VMware licensing from perpetual licenses to subscription bundles, repricing infrastructure many teams had budgeted as a fixed cost and pushing a visible share of customers to evaluate alternatives. Second, Broadcom later revoked public access to the VDDK-based VMware migration tooling, a reminder that anything living under a single vendor's roof can have its terms rewritten. Velero itself stayed Apache-licensed and open throughout — but "open today" is not the same as "governed in the open," and platform teams choosing a backup layer they might depend on for a decade noticed the difference.
Broadcom's own framing at KubeCon was unusually candid about this: the company said it did not want people to mistrust the open-source project or believe it was somehow a VMware-only thing when it had not been one for a long time. The CNCF application was filed in February 2026, the move was announced alongside the Amsterdam conference in late March, and the repository has since moved to the neutral velero-io GitHub organization. In July, Broadcom followed up by joining the CNCF as a Platinum member — dues, roadmap seat, and all — which reads as a costly signal that the donation was strategy rather than spring cleaning.
What "CNCF Sandbox" actually guarantees — and what it doesn't
For a team evaluating Velero as a backup layer, the donation changes the risk math in specific ways. It helps to be precise about which guarantees are real and which are wishful thinking.
What the move concretely buys:
- Vendor-neutral governance. The project answers to the CNCF Technical Oversight Committee, not to a Broadcom product manager. A future licensing pivot or portfolio cull cannot quietly redirect the roadmap, because no single company owns the roadmap anymore.
- A wider maintainer base. Velero now counts maintainers from Broadcom, Red Hat, and Microsoft. That matters more than logos: it means the project's bus factor and review bandwidth no longer depend on one employer's staffing decisions.
- A graduation path to watch. CNCF projects move Sandbox to Incubating to Graduated against public criteria — adoption, committer diversity, security posture, governance maturity. That ladder gives adopters an independent signal to track instead of reading vendor tea leaves.
And what it does not buy:
- A maturity stamp. Sandbox is the CNCF's entry tier — the starting gate, explicitly not a recommendation. Velero's production credibility comes from roughly a decade of clusters depending on it, not from the donation ceremony.
- Support or an SLA. Neutral governance does not page anyone at 3 AM. If you run Velero on your own fleet, you still own upgrades, plugin compatibility, and the restore drill.
- Immunity from ecosystem churn. The CSI, storage, and Kubernetes APIs underneath Velero keep evolving. Governance fixes who decides; it does not freeze what they must decide about.
| CNCF stage | What it signals | What to do as an adopter |
|---|---|---|
| Sandbox | Neutral home, early governance | Adopt on technical merit; track committer diversity |
| Incubating | Growing adoption, maturing process | Treat as default choice for the category |
| Graduated | Proven at scale, hardened security | Bet multi-year platform decisions on it |
CNCF CTO Chris Aniszczyk's comment on the donation pointed at the underlying shift: as cloud-native workloads scale, the industry's focus is moving from simple orchestration to long-term resilience and data management. Backup is becoming load-bearing infrastructure, and load-bearing infrastructure wants neutral ownership. That framing also explains why the donation landed alongside a KubeCon narrative about stateful AI workloads — model weights, checkpoints, fine-tuning data — where "we can reschedule the pod" was never a backup strategy.
The DR blueprint: Velero on a fleet you own
Governance answers whether Velero will survive. This section answers whether it can carry your disaster recovery — concretely, on machines you own rather than a managed cloud.
The single-cluster setup that actually matters
Velero's architecture has one hard requirement that shapes everything on owned hardware: backups must land in object storage. Velero persists cluster-state snapshots and volume data to a BackupStorageLocation, which speaks S3. On a hyperscaler that means a bucket. On your own fleet it means an S3-compatible endpoint — Hetzner Object Storage, Garage, RustFS, or Ceph RGW all work through the AWS plugin with path-style addressing.
One honest caveat on that list: MinIO, the longtime default answer here, has spent 2025–2026 retreating from its community edition — source-only releases, then an archived public repository. If your runbooks still say MinIO, treat the storage-backend choice as a live decision again, not inherited wisdom.
With the target settled, the setup that covers most tenants is small:
velero schedule create nightly-tenant-dr \
--schedule="0 2 * * *" \
--include-namespaces 'tenant-*' \
--snapshot-move-data \
--ttl 168h0m0sThree decisions hide in that one command. The namespace pattern scopes blast radius — system namespaces get their own schedule with different retention, because restoring kube-system alongside tenant data is how you turn a drill into an incident. The --snapshot-move-data flag says volume data leaves the cluster for object storage instead of lingering as local CSI snapshots, which is the difference between surviving a node failure and surviving the loss of the storage array. And the TTL encodes retention policy as code rather than as a wiki page nobody enforces.
The remaining choice is CSI snapshots versus file-system backup, and the rule is simple: use CSI snapshots wherever your CSI driver supports them, and file-system backup (Velero's Kopia-based data mover running through a node agent) for everything else — local-path volumes, hostPath-backed dev tiers, drivers without snapshot support. Most owned-hardware fleets end up running both, split by storage class rather than by tenant.
Then the part everyone skips and nobody should: the restore drill. A backup you have never restored is a hypothesis, not a recovery plan. Schedule a quarterly restore into an empty namespace on a scratch cluster, time it, and record the result. Velero's restore is namespace-remappable precisely so drills are cheap — use that.
The fleet layer: backing up a Cluster API topology
A single-cluster schedule is necessary but not sufficient for a Cluster API fleet, because a CAPI topology has two kinds of clusters with different things worth saving.
The management cluster holds the fleet's declared state: Cluster, MachineDeployment, MachineHealthCheck, and provider-specific objects that describe every workload cluster. Losing the management cluster without a backup does not delete running workloads — but it deletes your ability to reconcile, upgrade, or reprovision them, which is arguably worse. Back up the management cluster's CAPI namespaces on their own schedule, with longer retention, and include the CRDs themselves — restoring custom resources without their definitions fails in confusing ways.
The workload clusters hold tenant state and get the per-tenant schedules above. Here Velero's second talent matters as much as backup: migration. A Velero backup taken on cluster A restores onto cluster B, which makes it the fleet's cluster-replacement tool. The migration-restore workflow is the thing to practice:
velero backup create pre-migration \
--include-namespaces 'tenant-*' --snapshot-move-data --wait
velero restore create --from-backup pre-migration \
--namespace-mappings 'tenant-acme:tenant-acme' \
--restore-volumesRun that against a freshly CAPI-provisioned cluster the way you would run a fire drill — new machines, empty control plane, tenant namespaces materializing from object storage. If it works, you have proven both your backup and your reprovisioning path in one exercise. If it fails, you have found out on a Tuesday afternoon instead of during an outage.
The 2026 features that change the blueprint
Two recent Velero releases alter the blueprint above rather than just polishing it. Version 1.16 parallelized item-block backup — correlated resources back up across a thread pool instead of sequentially — which is the difference between a backup window that fits the night and one that bleeds into business hours once tenant counts grow. The same release added wildcard namespace patterns, so the 'tenant-*' scoping above works natively instead of requiring an enumerated list someone must maintain. Version 1.17 extended node-selection to CSI data-movement restores, letting you keep restore traffic off tenant-facing nodes.
None of these invents a new concept; all of them remove a reason the blueprint quietly stops working at scale.
Owning the backup target vs renting the snapshot button
Is backup for a fleet you own meaningfully different from the managed-snapshot story a hosted PaaS bundles by default? It is — not in the Velero layer, which behaves the same either way, but in everything around it that a hosted platform absorbs invisibly.
| Responsibility | Hosted PaaS (Render, Heroku, Railway) | Self-hosted fleet with Velero |
|---|---|---|
| Backup target durability | Their S3 replication, their problem | Your object storage, your replication policy |
| Scheduling and retention | A dashboard toggle | Your Schedule CRs in git, your TTL math |
| Encryption at rest | Provider-managed keys | Your bucket encryption and key management |
| Restore testing | Provider's internal drills (trust us) | Your quarterly drill, your runbook, your timing data |
| Control-plane state (etcd) | Included invisibly | Separate concern — Velero backs up objects via the API, not etcd itself |
| Cross-cluster migration | Redeploy from git (stateless bias) | Velero restore onto a fresh CAPI cluster, volumes included |
Two rows deserve emphasis. First, the restore-testing row cuts both ways: owning the drill is work, but it is also evidence. A hosted platform's "trust us" has failed loudly enough times — multi-tenant PaaS outages that took tenants' data stories down with the control plane — that a drill log with timestamps is a genuine competitive feature of running your own stack.
Second, the etcd row is the sharp edge. Velero backs up Kubernetes objects through the API server, which covers disaster recovery for everything declaratively described — but a true control-plane rebuild still needs etcd snapshots or a CAPI-driven reprovision plus Velero restore. Know which of those two your runbook assumes before you need it.
Should you adopt Velero now?
The verdict first: for a self-hosted fleet, the CNCF move removes the last structural objection to Velero. Adopt it if the checklist below fits; the thing to keep watching is graduation progress, not governance risk.
Adopt Velero now if:
- You run clusters you own (bare metal, Hetzner, colo) and need volume-level DR, not just "reapply the manifests and hope the PVCs survived."
- You operate more than one cluster, or plan to — migration restore is worth the setup cost alone the first time you replace a cluster instead of nursing one.
- You can give it S3-compatible storage with real durability (replicated, off the same failure domain as the clusters it protects).
- Your CSI story is settled enough to pick the snapshot-versus-filesystem split per storage class.
Watch, but proceed:
- Graduation progress. Sandbox to Incubating is the signal that committer diversity and security process have matured under neutral governance. Check the CNCF landscape page yearly, not weekly.
- Contributor concentration. The multi-vendor maintainer list is the headline; the commit graph over the next year is the verification. If Broadcom authorship stays dominant, discount the neutrality story accordingly.
- Plugin maintenance. Velero's provider plugins (AWS, Azure, GCP, CSI) evolve separately from core. Pin versions in git and read the upgrade notes — the 1.16 to 1.17 path is the current one to rehearse.
Wait if:
- Your storage layer cannot offer snapshots or stable file-level reads — Velero cannot conjure durability the substrate lacks.
- You have exactly one small cluster and no state worth more than its manifests —
etcdsnapshots plus git may genuinely be enough, and a backup operator is operational surface area you do not need yet.
Velero spent a decade becoming the default answer to Kubernetes backup, then spent three years under an owner that made teams nervous about defaults. The CNCF donation does not change a line of its code — but it changes who gets to decide what the next lines are, and that was the actual risk all along. For a fleet you own, on machines you own, with a backup target you own: that is now somebody else's governance problem, solved in your favor.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



