Skip to main content

One ClusterProfile vs One Flux per Cluster: What Sveltos v1.13 Changes About Fleet Add-On Delivery

11 min readDora NodaDora Noda
Share
On this page

Every Cluster API fleet eventually hits the same bookkeeping problem: each new workload cluster needs the same boring substrate — CNI, CSI, metrics-server, policy engine, ingress — before it can run a single tenant pod. The default GitOps answer is one Flux installation per cluster, each bootstrapped to its own clusters/<name> path. It works. It also means the number of Flux installations, bootstrap pipelines, and per-cluster upgrade chores grows linearly with the fleet.

Sveltos inverts the topology: one addon brain in the management cluster pushes add-ons to workload clusters it auto-discovers from Cluster API itself. Its v1.13.0 release (July 27, 2026) added fleet-wide Helm-update visibility and remote Kustomize sourcing that sharpen the comparison. This post is the concrete before/after: what classifier-matched, template-driven rollout buys a small CAPI fleet over hand-rolled Flux-per-cluster — and where Sveltos stops, so Cluster API stays the source of truth.

The short version, up front:

ConcernFlux per clusterSveltos ClusterProfile
Onboarding a new clusterBootstrap Flux into it (flux bootstrap --path=clusters/<name>), add its path + Kustomization to gitZero touch: Sveltos watches clusters.cluster.x-k8s.io and discovers CAPI clusters automatically
TargetingOne directory / Kustomization per cluster, or per-cluster overlays over a baseOne ClusterProfile with a label selector (env: prod); matching clusters get the add-ons, no per-cluster config
Bootstrapping a bare clusterChicken-and-egg: pull model needs networking + controller running firstPush model from the mgmt cluster deploys CNI/CSI onto empty clusters
Per-cluster variationKustomize overlays per clusterTemplates instantiated per cluster from cluster metadata + mgmt-cluster ConfigMaps/Secrets
Fleet-wide rollout pacingEach cluster reconciles on its own schedulemaxUpdate bounds concurrent cluster updates; staged/progressive rollouts across environments
DriftFlux re-syncs its own cluster's git stateContinuousWithDriftDetection deploys a watcher per managed cluster and re-syncs drifted resources
Dry runflux diff per cluster, one at a timesyncMode: DryRun simulates across all matching clusters; sveltosctl show dryrun lists every pending change
Upgrade visibilityCheck each cluster's HelmReleasesv1.13 records latestVersion/lastCheckedTime per release per cluster; sveltosctl show helm-updates shows the whole fleet
Controllers to upgradeN Flux installations (plus their CRDs)Sveltos in the mgmt cluster (+ a lightweight agent per managed cluster)

If that table settled it for you, the rest is the evidence. If it looks too good to be true, skip to the last two sections first — Sveltos has real costs, and a single-cluster fleet should probably ignore all of this.

The Flux-per-cluster tax​

Flux's multi-cluster story is well documented and honestly standalone: each cluster runs its own controllers and reconciles its own path. The official shape is one directory per cluster under clusters/, each with its own bootstrap or FluxInstance config. For a fleet of CAPI workload clusters, that means every clusterctl-provisioned cluster needs a bootstrap step — CI/CD installs Flux, points it at clusters/<name>, and the controller starts pulling.

Three costs follow from that shape. First, onboarding is a pipeline step, not an event: nothing in Flux notices that a new CAPI Cluster object exists. Your automation must bridge "Cluster API created a cluster" to "Flux is now running in it." Second, the pull model has a genuine chicken-and-egg problem at the bottom of the stack: a cluster pulls its CNI from git only once it has networking and a working controller, which is exactly what a fresh kubeadm-bootstrapped node does not have yet. Teams work around it with pre-baked images or bootstrap jobs, but the workaround is theirs to maintain. Third, fleet-wide change is N independent reconciliations: N Flux versions to upgrade, N flux-system namespaces to keep healthy, and per-cluster overlays (<dir>/overlays/<cluster>) whenever one cluster differs from the base. ArgoCD at least ships ApplicationSet generators for fan-out; Flux's own ecosystem answers "how do I target a fleet" with per-cluster resources plus a bootstrap repo.

None of this is broken. For one cluster, or a handful of long-lived clusters, it is arguably the right call — each cluster is self-sufficient, and a compromised hub cannot reach into it. The tax only bites as clusters churn: dev/test clusters spinning up weekly, per-tenant clusters, multi-region expansion. That is the fleet where the per-cluster bootstrap starts to feel like toil.

How Sveltos inverts it​

Sveltos runs in the management cluster — the same cluster that already runs Cluster API — and treats workload clusters as delivery targets, not GitOps peers. The single most load-bearing sentence in its docs is that when Sveltos is deployed in a management cluster with Cluster API, no further action is required: it watches clusters.cluster.x-k8s.io instances and programs the matching ones. A new CAPI cluster with the right labels gets its full add-on set with zero onboarding. "Adding a new cluster with the right labels automatically brings everything to the desired state" is the project's own summary, and the mechanism is exactly that boring: label selectors over discovered Cluster objects.

Delivery is push-based. The addon-controller in the mgmt cluster renders each ClusterProfile's payloads and applies them to matching workload clusters through their kubeconfigs, which it already holds via CAPI. Because nothing needs to be running inside the target first, Sveltos can bootstrap a bare cluster bottom-up: CNI and CSI go in via the same mechanism as everything above them, and the chicken-and-egg simply does not exist. A small agent (sveltos-agent, plus a drift-detection manager only when a profile asks for drift detection) lands in each managed cluster to report health, classifier data, and drift events back — but the brain stays in one place.

The targeting primitive is the ClusterProfile, a cluster-wide CRD that says "these add-ons, on all clusters matching this selector." One profile carries Helm charts (helmCharts), raw manifests via ConfigMap/Secret references (policyRefs, deployable locally to the mgmt cluster or remotely to the workload cluster), and Kustomize sources (kustomizationRefs, which can consume Flux GitRepository/OCIRepository/Bucket objects directly). Updating the profile updates every matching cluster — the fan-out Flux lacks is the entire point of the tool.

Worked example: one profile for the whole substrate​

Concretely, a CAPH-provisioned Hetzner fleet's baseline — Cilium, metrics-server, cert-manager, Kyverno — collapses to a handful of profiles like this:

yaml
apiVersion: config.projectsveltos.io/v1beta1
kind: ClusterProfile
metadata:
  name: baseline-networking
spec:
  clusterSelector:
    matchLabels:
      env: prod
  syncMode: ContinuousWithDriftDetection
  helmCharts:
  - repositoryURL: https://helm.cilium.io/
    repositoryName: cilium
    chartName: cilium/cilium
    chartVersion: 1.17.3
    releaseName: cilium
    releaseNamespace: kube-system
    helmChartAction: Install

Every current and future env: prod cluster gets Cilium, drift-watched. The interesting half is the Classifier: it labels clusters dynamically from runtime state — Kubernetes version constraints, deployed resources — so profiles can track properties instead of names. The canonical example pins an add-on version to the cluster's Kubernetes version: a Classifier matching >= 1.24.0, < 1.27.1 stamps k8s-version: v1.24 onto the Cluster object, a ClusterProfile selects that label, and upgrading the cluster automatically re-labels it into the newer add-on profile. Add-on upgrades follow cluster upgrades with no human editing selectors.

Per-cluster variation comes from templating rather than per-cluster overlays: add-on manifests and Helm values are templates instantiated per cluster from cluster metadata and mgmt-cluster ConfigMaps/Secrets. One profile, N instantiations — the overlay directory per cluster disappears. And when a cluster stops matching (re-labeled out of env: prod, decommissioned), stopMatchingBehavior decides the ending: the default WithdrawPolicies removes everything the profile deployed; LeavePolicies leaves the workload in place. Decommissioning a cluster cleans up its add-ons as a side effect of how targeting works, not as a runbook step someone must remember.

What v1.13 itself adds to this comparison​

The features above are the long-standing Sveltos shape. v1.13.0 sharpens the Flux-per-cluster comparison in four specific ways, all verifiable in the release notes:

Fleet-wide Helm-update visibility. Sveltos now periodically checks whether each deployed chart has a newer version or same-minor patch upstream (HTTP repo or OCI registry) and records latestVersion, latestPatchVersion, and lastCheckedTime on the owning ClusterSummary — per cluster, not blended across the fleet. sveltosctl show helm-updates lists every behind release everywhere; the dashboard adds an "Update Available" column. Detection only — nothing auto-upgrades — but it answers "which clusters run stale charts" in one command, a question Flux-per-cluster answers by visiting N clusters.

Remote Kustomize sourcing. KustomizationRef.RemoteURL fetches content directly from an HTTP/HTTPS endpoint or OCI registry while preserving directory structure (which Kustomize needs, unlike the flattened PolicyRef OCI path). Fewer manifests must be mirrored into mgmt-cluster ConfigMaps first.

Dynamic paths from mgmt-cluster data. PolicyRef.Path and KustomizationRef.Path/Components can now be Go templates resolved against TemplateResourceRefs — resources read from the management cluster — so one profile can pick its overlay path dynamically instead of fixing it per profile.

Sharper failure surfaces. EventTrigger instantiation errors land on EventReport.Status instead of hiding in event-manager logs; Helm post-rendering gains an explicit PostRenderStrategy (combined, separate, nohooks) for the hooks-vs-templates edge; the MCP server grows to 22 tools with fleet-wide classifier-label visibility. (v1.14.0 followed in August with per-chart scoping of drift redeploys — one drifted chart no longer forces every chart in its profile to re-upgrade.)

Rollout safety, concretely​

Fan-out without pacing is a footgun, so the rollout controls matter as much as the targeting. By default a ClusterProfile change updates all matching clusters concurrently; maxUpdate caps how many move at once, and the rolling-update strategy pairs it with ValidateHealths checks so the rollout only advances while clusters stay healthy. For coarser staging, progressive rollouts walk environments in order — deploy to staging-labeled clusters, verify, wait out a delay, then promote to prod — with ClusterPromotion objects tracking the waves. And syncMode: DryRun simulates a profile change across every matching cluster first, with sveltosctl show dryrun rendering the per-cluster action table (install/upgrade/delete per resource) before anything touches a workload cluster.

Drift handling is the ContinuousWithDriftDetection mode from the example above: Sveltos deploys a lightweight watcher into each managed cluster configured with that profile's resource list, the watcher reports modifications (spec, labels, annotations, rules) back to the mgmt cluster, and the profile re-syncs. Resources that legitimately mutate in place (autoscaler-touched replicas, operator-managed fields) opt out with the projectsveltos.io/driftDetectionIgnore annotation. Reconcile outcomes and drift counts are Prometheus metrics (projectsveltos_reconcile_outcome_total per profile and status, projectsveltos_total_drifts) with a Grafana dashboard — the per-cluster consistency view the TODO item promises is a real endpoint, not a slogan.

Where Sveltos stops​

The scope limit is as important as the feature list: Sveltos delivers add-ons; it does not provision clusters, version machine images, or own the upgrade of Kubernetes itself. Cluster API remains the source of truth for "what clusters exist and what version they run" — Sveltos reacts to that truth via discovery and Classifiers, and even its version-following trick is downstream of CAPI having upgraded the cluster first. If you want one throat to choke for cluster lifecycle, it is still CAPI.

The healthy end state the ecosystem has converged on is the "super-cluster" split, visible in production homelab and platform setups: CAPI provisions, Sveltos delivers infra add-ons, Flux or ArgoCD keeps app-level GitOps. Sveltos explicitly does not compete with GitOps controllers — its own docs recommend the pairing both directions. Mature shops run Flux in the mgmt cluster to version the Sveltos CRDs themselves (GitOps for the brain: ClusterProfiles, Classifiers, and referenced ConfigMaps all in git, so the push-based delivery still has a reviewed, revertable source), while Sveltos installs and configures Flux or ArgoCD inside workload clusters as just another ClusterProfile payload for teams that want pull-based app deploys. Each tool owns the layer it is good at instead of one tool stretching across both.

Verdict: adopt or skip?​

Adopt Sveltos when the fleet looks like the tax description: CAPI-managed clusters, churn (dev/test/ephemeral/tenant clusters), and a substrate larger than two charts. The payoff scales with cluster count and turnover — each cluster you never bootstrap by hand and each fleet-wide chart bump you make in one profile edit is the return. The v1.13 Helm-update detection is a quiet force multiplier here: the bigger the fleet, the more "which clusters are stale" dominates addon-ops time, and that is now one command.

Skip it — keep Flux-per-cluster — when the fleet is one long-lived cluster, or a handful of static ones. Sveltos is another control plane to run, upgrade, back up, and understand; its mgmt cluster becomes a blast-radius concentrator (push access to every workload cluster's kubeconfig in one place), which is precisely the failure mode Flux's standalone per-cluster model avoids. If your clusters rarely churn and your substrate fits in one clusters/<name> directory you already understand, the inversion buys ceremony, not leverage.

For the Hetzner-scale CAPI fleet in between — a few prod workload clusters, staging that churns, a roadmap pointing at more tenants — the migration path is low-drama: install Sveltos in the mgmt cluster, write one ClusterProfile for the substrate you already deploy via Flux, let auto-discovery adopt the existing clusters by label, and keep Flux around for app deploys (managed by Sveltos itself, if you like the symmetry). The comparison table at the top stops being a pitch at that point and becomes a checklist you can verify cluster by cluster.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Fleet addon delivery like this is the layer a self-hosted PaaS stands on; star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex