Skip to main content

Kairos + k0rdent + bindy: Is This the 2026 Reference IDP a Small Platform Team Should Copy?

11 min readDora NodaDora Noda
Share
On this page

A regulated bank runs 50-plus Kubernetes clusters across VMware and multiple clouds, every node booting from a versioned OCI image, every cluster a reconciled CRD, every DNS record a Git commit. No ticket queues, no snowflake nodes, no out-of-band state. The stack is Kairos, k0rdent, and bindy — stitched together with FluxCD — and the build log is public on the CNCF blog.

So here is the question a small platform team should be asking: is this the 2026 reference internal developer platform? Not "is it impressive" — it clearly is — but "should we copy this exact stack for machines we own?" I read the anniversary retrospective and the production build log, and my answer is: adopt the fleet layers, diverge at the app layer. The bundle is a superb fleet IDP. It is not a PaaS, and it does not pretend to be one.

The verdict up front: adopt the fleet, diverge at the app layer​

LayerReference componentSmall-team verdict
Node OSKairos (immutable, OCI-booted)Adopt — the highest-leverage layer in the bundle
Cluster lifecyclek0rdent KCM over Cluster API + k0sAdopt — stop hand-rolling CAPI glue
Add-on statek0rdent KSM (SveltOS default, Flux-backed)Adapt — pick one state provider deliberately
Observability + FinOpsk0rdent KOF (OTel, OpenCost, VictoriaMetrics)Adopt — chargeback ships in the box
DNSbindy (DNS-as-CRDs, RFC 2136)Adapt — proven pattern, young project
Self-service APIFoundry (Rust, event-driven)Diverge — roadmap, not software you can install
Build → deploy pathNot in the bundleDiverge — buildpacks and per-tenant routing are yours to build

The rest of this post is the evidence behind that table: what k0rdent shipped in year one, what RBC Capital Markets proved under compliance load, and where the honest gaps are.

What k0rdent's first year actually shipped​

Mirantis open-sourced k0rdent in February 2025, and the project just closed its first year with a retrospective and a CNCF On-Demand anniversary webinar. The headline number is velocity: v0.1 on February 6, 2025 to v1.7 within twelve months, with a v2.0 roadmap already in sight (Mirantis retrospective). That pace cuts both ways — fast shipping plus moving contracts — but the architecture held stable around three components mapped to three platform problems.

ComponentProblem it answersBuilt on
KCM (Cluster Manager)How are clusters provisioned, upgraded, secured, recovered?Cluster API, k0s, infrastructure providers
KSM (State Manager)What runs on every cluster, and how is it governed over time?Service templates, SveltOS default provider, Flux-backed Helm/Kustomize
KOF (Observability + FinOps)How are metrics, traces, and costs collected and charged back?OpenTelemetry, Prometheus/Grafana, OpenCost, VictoriaMetrics storage

KCM went from prototype to a production-ready cluster layer, expanded provider support across AWS, Azure, GCP, OpenStack, and VMware, added anonymized usage telemetry, and introduced a "regional clusters" concept for multi-site fleets. KSM treats platform beachhead services — CNIs, CSIs, runtime dependencies — as template-driven service templates with intermediate ServiceSet objects, roughly analogous to how ReplicaSets manage pods. KOF is the quiet overachiever: 14 releases since February 2025, auto-configuration of the observability stack onto child clusters via a cluster-role label, and real multi-tenancy through tenant-ID credential isolation.

Two external signals corroborate the maturity story. In July 2026, both k0s and k0rdent achieved CNCF Certified Kubernetes AI Conformance at v1.35, which means GPU scheduling and inference-networking behavior is tested rather than asserted. And the ecosystem is filling in around the edges: Tigera ships a managed Calico integration for k0rdent-provisioned clusters, and Mirantis paired with Netris in March 2026 so network automation ships as part of cluster delivery instead of a post-hoc ticket.

What RBC's 50-cluster build log proves in production​

The reference-stack claim would be hand-waving without a production witness. The May 2026 CNCF build log from RBC Capital Markets, authored by Erick Bourgeois, is that witness — and it is unusually honest about operating under SOX, PCI-DSS, and Basel III constraints, where auditability and drift prevention are legal requirements, not nice-to-haves.

RBC's platform runs more than 50 clusters across on-premises VMware and multiple clouds. Their framing is three questions no single off-the-shelf tool answered together: how do you manage cluster lifecycle itself, how do you make every node reproducible and tamper-evident at boot, and how do you integrate Kubernetes service discovery with enterprise DNS without a ticket queue? Each question got one project.

Kairos answered the node question. Every node boots from an OCI image built on a RHEL-derived base with approved security configuration baked in, and node behavior — SSH keys, networking, SSSD authentication against Active Directory, Kubernetes agent registration — is versioned cloud-config YAML flowing through FluxCD like any other component. The discipline detail matters more than the tool choice: Kairos images go through a GitHub Actions pipeline that builds, runs integration tests against a live VM, and publishes a new OCI tag only on a clean pass, with nightly builds catching upstream regressions before production. VM provisioning itself is GitOps-native through VirtRigaud, an operator that manages vSphere, Libvirt/KVM, and Proxmox VMs as CRDs — so creating a node is a pull request.

k0rdent answered the cluster question, paired with k0s and k0smotron. k0s is a single-binary distribution with no host OS dependencies beyond the kernel, which makes it a natural fit for immutable Kairos nodes — no package managers, no runtime systemd surgery, none of kubeadm's host assumptions. k0smotron goes one step further by running control planes as workloads inside the management cluster, so even the control plane is a reconciled CRD with no out-of-band state. RBC is extending the same model toward a spot-computing scheduler that absorbs donated physical servers dynamically.

bindy answered the DNS question, and this is the most instructive layer. RBC's enterprise DNS is Infoblox, and every record request went through a human ticket workflow measured in hours or days — while 50-plus clusters each spun up dozens of endpoints. Bourgeois built bindy as a Rust kube-rs operator that manages DNS zones and records as first-class resources: DNSZone and ARecord CRDs in Git, RFC 2136 dynamic updates to push changes, and a bindcar sidecar exposing an RNDC-style REST interface for zone lifecycle operations. Provisioning time dropped from hours to seconds, and the audit trail became Git history instead of a ticket system.

Note the roadmap signals too: CRD-based compliance scoring for zone health and a future MCP server interface for AI-driven platform tooling. DNS here is being built as an agent-operable API surface, not a console.

The architectural punchline is one sentence from the post: a single Git repository fans out to k0rdent/CAPI manifests for clusters, Kairos cloud-config for nodes, and bindy CRDs for DNS, with FluxCD as the single reconciliation plane. Drift at the node, cluster, and network level is structurally prevented rather than operationally managed.

Layer by layer: adopt, adapt, or diverge on owned hardware​

Now the verdict matrix, expanded for a small team running owned machines — Hetzner-class bare metal or self-managed VMs — rather than a bank's VMware estate.

Kairos: adopt. Immutable, OCI-booted nodes with A/B atomic upgrades, rollback, and recovery are the highest-leverage layer in the bundle, and Kairos is a CNCF Sandbox project with broad base-image support (Hadron by default, plus Alpine, Debian, Fedora, openSUSE, Rocky, and Ubuntu) on x86-64 and ARM64. The Kairos-plus-k0s pairing has been an explicit collaboration since April 2025, so the "standard" images with Kubernetes baked in are a maintained path, not a hack. Copy RBC's image CI discipline too — tested OCI tags and nightly upstream builds — because immutable infrastructure without image validation just moves the snowflakes into the registry.

KCM over Cluster API: adopt. If you run more than a couple of clusters on owned hardware, KCM's templated, declarative cluster lifecycle beats hand-rolled clusterctl glue and shell scripts from day one. The provider surface already covers the clouds plus OpenStack and VMware, and the project's stated direction is monthly or bi-weekly releases. The one caveat is release velocity itself: twelve months from v0.1 to v1.7 with v2.0 ahead means pinning versions and reading upgrade notes is mandatory, not optional.

KSM: adapt. The beachhead-services problem is real — every cluster needs its CNI, CSI, and runtime dependencies governed identically — and template-driven service templates are a sound answer. But KSM's default state provider is SveltOS, with Flux-backed Helm/Kustomize as the alternative path, and that choice deserves a deliberate decision: teams with deep GitOps expertise may prefer the Flux-native route, while teams that want multi-tenancy without GitOps-specific headcount will take SveltOS. Adopt the layer, choose the provider consciously.

KOF: adopt. Full-stack observability plus OpenCost chargeback, auto-applied to child clusters by label, with tenant-isolated credentials, is exactly the layer small teams postpone until the first cost dispute or the first 3 a.m. page with no traces. Fourteen releases in year one and extensive docs covering architecture, upgrades, storage, and scaling suggest this component is past the demo stage. Cost visibility via OpenCost is the feature that pays for the adoption effort.

bindy: adapt. The pattern is proven — DNS-as-CRDs with GitOps reconciliation killed an hours-to-days ticket queue — but bindy itself is a young, single-origin project (a bank engineer's operator, now public at firestoned/bindy) rather than a governed multi-vendor standard. Adopt the architecture: zone and record CRDs, dynamic updates to your backend, Git as the audit trail. But go in with eyes open about bus factor, and watch whether the compliance-scoring and MCP-server roadmap items land before you bet multi-cluster DNS on it.

Foundry: diverge (for now). RBC's planned Rust self-service API for governed, event-driven cluster and DNS provisioning is exactly what a platform team eventually wants developers to touch instead of raw CRDs. It is also explicitly future work described in a blog post, not software you can install. Do not put "Foundry" on your architecture diagram; put "self-service API (to be built)" and let KCM/KSM/KOF's CRDs be the API underneath it until then.

Build → deploy path: diverge. This is the boundary the reference stack does not cross. Nothing in the bundle turns a git push into a running HTTPS service: no buildpacks, no image builds from source, no per-tenant routing, no app-level auth or billing. A git-push PaaS adopts this stack for everything below the app layer and builds the app layer itself — that is the correct division of labor, not a criticism of the bundle.

The honest gaps and what to watch​

Three risks deserve plain language. First, contract churn: a project that ships v0.1 to v1.7 in a year and is heading to v2.0 will move APIs under you; anyone adopting k0rdent in 2026 should budget for at least one breaking upgrade and keep provider manifests versioned and tested. Second, single-project bets: bindy and the Kairos-based spot scheduler with k0smotron and Kata Containers are genuinely innovative and genuinely early — the spot work was announced as a future post, not a shipped feature. Third, the SveltOS-versus-Flux decision inside KSM is unresolved by default; defaults are fine, but this one shapes your team's skill requirements for years, so make it explicitly.

What to watch, in order of signal value: whether k0rdent v2.0 stabilizes the KCM/KSM/KOF contracts or reshuffles them; whether bindy's compliance scoring and MCP server interface ship, which would make DNS the first fully agent-operable layer of the reference stack; whether RBC publishes the promised spot-computing follow-up with real utilization numbers; and whether the KOF multi-tenancy model grows the audit and policy hooks regulated adopters beyond one bank will demand.

The bottom line stands: 2026 finally has a credible, open-source, all-CNC-ecosystem answer to "what should our fleet look like" — Kairos for nodes, k0rdent for clusters and state and cost, bindy-shaped DNS, FluxCD holding the reconciliation plane together. Copy that. Then build the half the bundle deliberately leaves out: the path from a developer's git push to a running service.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex