Skip to main content

SNCF Ships Its Cluster API Providers as OCI Artifacts: What an ORAS-Based Provider Supply Chain Buys a Self-Hosted Fleet

9 min readDora NodaDora Noda
Share
On this page

The most interesting sentence in CNCF's SNCF case study is the one nobody wrote a post about. Buried in the GitOps section, between the ArgoCD paragraph and the metrics table, it reads: "ORAS was integrated into the supply chain to manage Cluster API's providers as OCI artifacts, ensuring seamless lifecycle management of every high-level primitive through GitOps." One sentence — no diagram, no command, no follow-up.

That sentence deserves the follow-up. It describes a genuinely unusual choice: taking the components that provision your entire fleet — the Cluster API providers themselves — off the default clusterctl init path of pulling YAML from GitHub releases, and distributing them as versioned, signed artifacts from an OCI registry instead. This post unpacks what that pattern is mechanically, what it buys over each of the three officially documented alternatives, and whether a Hetzner-scale fleet running four providers should copy it.

The short version, up front:

Distribution pathWhat clusterctl pulls fromPinned?Signed?Air-gap ready?GitOps-consumable?
Default: GitHub releasesgithub.com release assets over HTTPSTag only — a re-pushed tag moves under youNo check in the fetch pathNo — needs github.com + proxy.golang.orgOnly via a wrapper that runs clusterctl
Overrides + image mirrorsYour $XDG_CONFIG_HOME files + mirrored imagesYes, by file pathNoYes, but workstation-local and manualNo — files live outside git
GitLab generic packagesYour GitLab package registryYes, per package versionNo (relies on GitLab auth)Yes, with a mirror jobPartially — YAML still fetched by clusterctl
OCI artifacts via ORASYour OCI registry, by tag or digestYes — immutable digest pinningYes — cosign referrers on the same digestYes — oras cp replicates registriesYes — registries are API-driven stores

If you run one management cluster and four providers, the bottom row is probably overkill today — but the reasons it exists (digest pinning, signature verification, registry replication) are the exact failure modes that bite you the week you add a second management cluster or lose GitHub access mid-incident. Read on for the mechanics and the honest adoption line.


What clusterctl init pulls by default​

Every clusterctl init --infrastructure hetzner starts from a provider repository: a well-known location holding, at minimum, a metadata.yaml (which release series maps to which Cluster API contract) and a *-components.yaml (every CRD, controller Deployment, and RBAC object the provider needs), plus workload-cluster templates. That contract is defined in the clusterctl provider contract, and the default repository for nearly every provider is a GitHub release: a semver tag with the YAML files attached as downloadable assets.

Version resolution has its own hidden dependency. clusterctl discovers available provider versions through a Go module proxy — GOPROXY, defaulting to proxy.golang.org — falling back to the GitHub API only when the proxy is disabled or the provider doesn't follow Go versioning. So a default install phones home to at least two public services before it applies a single object, and what it downloads is addressed by a mutable tag: if a maintainer ever re-pushes a release (it happens — yanked images, rebuilt manifests), the bytes behind v1.0.7 change without the version string changing.

For a single management cluster, this is fine. It is the documented happy path, and the CAPH quickstart rightly teaches it first. The failure modes only show up at fleet scale or under pressure: no cryptographic check that the YAML you fetched is the YAML the maintainer published, no story for installing when github.com is unreachable, and no central record of which provider version each management cluster resolved — because "latest" is evaluated at init time on whoever's laptop ran it.

What "providers as OCI artifacts" changes mechanically​

ORAS (OCI Registry As Storage, a CNCF Sandbox project since July 2021, CLI at v1.3) treats a container registry as a generic artifact store: oras push uploads arbitrary files addressed by tag or digest, with an artifactType declaring what the blob is rather than masquerading as a container image. A provider bundle maps onto this almost embarrassingly well — the components YAML, the metadata YAML, and the cluster templates are just files:

bash
oras push registry.internal/capi/providers/caph:v1.0.7 \
  --artifact-type application/vnd.capi.provider.bundle.v1+yaml \
  infrastructure-components.yaml:application/yaml \
  metadata.yaml:application/yaml \
  cluster-template.yaml:application/yaml

Three properties fall out of that one command that none of the clusterctl-native paths give you together. First, digest pinning: every push produces an immutable sha256: digest, so a fleet can standardize on caph@sha256:9f3a… and make "which bytes did cluster #12 install" a question with a permanent answer. Second, attached provenance: OCI 1.1 referrers let you hang cosign signatures and SBOMs off the same digest (oras discover lists them, cosign verify checks them), so the signature verification the GitHub-release path lacks becomes a registry-native step. Third, registry mechanics: replication (oras cp between registries), RBAC, retention policies, and pull-through caching are problems your registry already solved for container images — the provider bundles inherit all of it.

That is what the case study means by "lifecycle management of every high-level primitive through GitOps": the providers stop being special snowflakes fetched by a CLI from a forge, and become versioned artifacts flowing through the same store-and-verify pipeline as everything else the fleet runs.

The seam the case study doesn't show​

Here is the honest gap: CNCF's one sentence says SNCF stores providers as OCI artifacts, not how those artifacts get installed. And there is a real seam here, because the two halves of the toolchain don't natively meet. clusterctl init understands GitHub releases, GitLab generic packages, and local directories — it has no "install from OCI artifact" source. ArgoCD, which SNCF uses as its GitOps layer across hundreds of public-cloud clusters, has no first-class "sync from OCI artifact" source either (Flux does, via its OCIRepository kind — one of the few places Flux leads Argo on this axis).

So the pull-and-apply step has to live somewhere. The realistic options:

WiringHow it worksCost
CI pushes, job appliesPipeline oras pulls the pinned digest and kubectl applys it (or feeds it to clusterctl via the overrides layer)A small amount of glue YAML; the digest pin lives in git, which is exactly where GitOps wants it
Flux OCIRepositoryFlux polls the registry tag/digest and applies the fetched artifact as a kustomizationRequires Flux alongside or instead of ArgoCD on the management cluster
Registry webhook → re-applyRegistry notifies on new digest; automation opens a PR bumping the pinMost machinery; buys you Renovate-style provider updates

None of these is exotic — the first is a cronjob and a service account — but notice what they share: the OCI artifact is the store, and something boring does the apply. Anyone selling this pattern as "zero glue" is skipping the row that matters. SNCF's scale (a national railway that cut provisioning from a month to 30 minutes and now updates every cluster monthly with zero drift) justifies glue that a smaller fleet should think twice about.

Should a four-provider Hetzner fleet copy this?​

Count the actual surface. A CAPH-based management cluster installs four providers: core cluster-api, bootstrap kubeadm, control-plane kubeadm, and infrastructure hetzner. SNCF's fleet spans the same four provider roles — core, bootstrap, control-plane, plus its OpenStack infrastructure provider — multiplied by every datacenter cluster that needs identical versions. Your blast radius from provider-version skew is one management cluster; theirs is a railway.

Adopt the ORAS pattern when one of these is true:

  • You run more than one management cluster (multi-region, or separate prod/staging fleets) and have already been bitten by version skew between them. Digest-pinned artifacts are the cheapest way to make "same providers everywhere" mechanically true.
  • You need an install path that survives losing github.com. Incident-time reinstalls during a forge outage are the scenario the overrides layer technically covers but nobody keeps current. A self-hosted registry (Zot and Harbor both speak everything ORAS needs) with replicated provider bundles is the version of that runbook that actually works at 3am.
  • Compliance wants signatures on the control plane's own supply chain. If you already cosign-verify workload images with Kyverno (SNCF runs Kyverno fleet-wide), extending the same verify-before-apply policy to provider bundles is coherent — and only possible once the bundles live in a signature-aware store.

Skip it — for now — if none of those apply, and take the two cheaper wins instead. First, vendor the exact provider YAML you installed into git next to your management-cluster definition; clusterctl generate provider emits it, and a vendored copy plus a comment with the source tag answers "what is this fleet running" without any new infrastructure. Second, mirror the provider container images to your existing registry with clusterctl init list-images plus the documented images: repository override — that covers the air-gap case for the bytes that actually get pulled at runtime, which is most of the outage risk for a fraction of the machinery.

The pattern to watch is the day these two stop being enough: the second management cluster, the first compliance questionnaire asking who signed your CAPI providers, or the first incident where clusterctl upgrade can't reach a release page. At that point the migration is small precisely because ORAS treats providers as dumb files — push the bundles you already vendored, pin the digests, and point your existing glue at the registry. SNCF built the cathedral; you can start with the parish church and upgrade along the same road.


Running your own fleet on machines you own is exactly the problem Bex.co exists for — push a git repo, get a running HTTPS service on your own hardware, reconciled the same declarative way SNCF reconciles its clusters. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex