Skip to main content

One Scheduler Per Cluster Isn't a Fleet Strategy: How KubeStellar, Cluster API, and Hive Split the Multi-Cluster Job

11 min readDora NodaDora Noda
Share
On this page

Your Kubernetes scheduler goes blind the day you add a second cluster. Inside one cluster, kube-scheduler sees every node, every pod, every GPU, and places each workload where it fits. Across two clusters, it sees nothing: there is no scheduler above the cluster line, so "which cluster does this workload land on" gets answered by whoever runs kubectl, whichever GitOps path happens to point there, or nobody at all.

The fix is a three-layer split the ecosystem has finally started naming explicitly. Cluster API provisions the clusters. KubeStellar places workloads across them with declarative BindingPolicy objects. And hive — the KubeStellar project's AI-agent orchestrator — coordinates the fleets of coding agents that increasingly operate the whole system. Provision the fleet, place onto the fleet, orchestrate the agents running on it: three jobs, three tools, one stack.

This matters now because the workload pushing teams past one cluster is no longer just traffic growth. It is AI agents: sandboxes that need GPU affinity, data locality, and isolation from the web services next to them. If your platform's agent story still assumes a single cluster with a single scheduler, here is the seam, the tooling that covers it, and an honest call on whether placement belongs in your control plane or stays someone else's problem.

Provisioning Is Not Placement: The Seam Cluster API Leaves Open​

Cluster API (CAPI) is very good at its job: declarative lifecycle management for Kubernetes clusters themselves. You declare a Cluster object, the infrastructure provider (CAPH for Hetzner, CAPA for AWS, and dozens more) reconciles machines into existence, and upgrades, repairs, and scaling happen through the same desired-state loop you already trust for pods. What CAPI hands you at the end is a registered, healthy, labeled cluster.

What it does not hand you is an answer to the next question: given five healthy clusters, which one runs this workload? That is placement, and it is a genuinely different job from provisioning. Provisioning is about machines converging to a declared shape. Placement is about matching workloads to clusters along dimensions the machine layer never sees:

  • Bin-pack vs. spread. Should this Deployment's replicas stack onto the cheapest cluster with room, or spread across clusters for blast-radius isolation? kube-scheduler answers this for nodes; nobody answers it for clusters by default.
  • Accelerator affinity. An agent sandbox that needs a GPU must land on a cluster whose node pool actually has one — and ideally the right one, once Dynamic Resource Allocation lets clusters advertise device topology instead of opaque resource counts.
  • Data locality. A training job or an agent with a big scratch volume wants the cluster near its data, not the cluster with the most free CPU.
  • Tenancy and trust. Untrusted tenant sandboxes may belong on a hardened cluster that never runs control-plane or platform-system workloads.

CAPI's own ecosystem acknowledges the gap. The KubeStellar Console added a Cluster API federation provider in 2026 that reads cluster.x-k8s.io objects directly, precisely so the provisioning layer and the placement layer can share one view of the fleet. Provisioning produces labeled clusters; placement consumes those labels. Confuse the two and you end up either hand-assigning workloads per cluster (which stops scaling at about three clusters) or pretending one big cluster can absorb everything (which stops being true the day GPU sandboxes and web services start fighting over the same nodes).

How KubeStellar Answers It: BindingPolicy in 20 Lines of YAML​

KubeStellar is the CNCF Sandbox project that owns the middle layer: multi-cluster workload placement as a declarative API. Its architecture has three roles worth learning, because every placement decision flows through them:

  • WDS (Workload Description Space) — the hub cluster where you write what you want ("this Deployment should exist on GPU-bearing clusters").
  • ITS (Inventory & Transport Space) — built on Open Cluster Management (OCM), this is the registry of known clusters plus the delivery machinery that carries manifests outward and status back.
  • WECs (Workload Execution Clusters) — the clusters that actually run things. Your CAPI-provisioned fleet, enrolled as members.

The entire placement contract is one custom resource, BindingPolicy. Here is a realistic one for the workload this post cares about — agent sandboxes that need GPUs:

yaml
apiVersion: control.kubestellar.io/v1alpha1
kind: BindingPolicy
metadata:
  name: agent-sandbox-to-gpu-clusters
spec:
  clusterSelectors:
    - matchLabels:
        accelerator: "nvidia-l4"
        tenancy: sandbox
  downsync:
    - apiGroup: apps/v1
      resources: ["deployments"]
      namespaces: ["agent-sandboxes"]
      objectSelectors:
        - matchLabels:
            workload: agent-sandbox

Read it as intent, not procedure: any Deployment labeled workload: agent-sandbox in the agent-sandboxes namespace gets propagated to every cluster labeled accelerator: nvidia-l4 and tenancy: sandbox. Add a fourth GPU cluster with those labels and the workload fans out to it with no manifest change — the project's own multi-WEC walkthroughs demonstrate exactly this scale-out without touching workload YAML. Status (replica counts, conditions) flows back up through the ITS so the hub knows what actually landed.

Two properties make this more than "kubectl with extra steps." First, placement is label-driven and continuous: drift between intent and reality gets reconciled, the same way a Deployment controller reconciles pods. Second, KubeStellar deliberately builds on top of OCM rather than replacing it — OCM handles cluster registration and manifest transport, KubeStellar adds the policy engine that decides what goes where. If you already run OCM, KubeStellar is an additive brain, not a migration.

Where Hive Fits: Agents Are the Workload and the Operator​

The third layer is the one most likely to be misunderstood, so here is the crisp version: hive orchestrates AI agents; it does not place Kubernetes workloads. Conflating the two produces bad architecture — an agent framework duct-taped into doing scheduling, or a placement engine asked to reason about agent autonomy. They are adjacent layers, not substitutes.

Hive (developed as kubestellar/hive) is a single Go binary that orchestrates fleets of AI coding agents — Claude, Copilot, Gemini, Goose as interchangeable backends — against software projects. It files issues, opens pull requests, reviews code, and at the highest operator-selected autonomy level, merges on green CI. Agents run as CLI subprocesses inside tmux sessions, in a container or a Kubernetes pod, under deterministic guardrails rather than vibes. Its genuinely novel contribution is ACMM, a maturity model for agent autonomy that lets an operator dial how much authority agents hold, from "suggest" up to full self-management — the level the KubeStellar project itself dogfoods, with hive managing its own development repos around the clock.

Why does this belong in a post about placement? Because those agent fleets are the workload that breaks single-cluster thinking fastest. A hive deployment supervising dozens of agents, each wanting an isolated sandbox with bursty CPU and occasional GPU access, is exactly the tenant that needs bin-pack-vs-spread policy, accelerator affinity, and trust-tier separation from the previous section. And symmetrically, the agents hive orchestrates are increasingly the operators of the fleet — triaging failed placements, opening PRs against BindingPolicies, reviewing cluster-upgrade plans. The stack composes: CAPI provisions the clusters, KubeStellar places the agent sandboxes across them, hive orchestrates what the agents do once they land. The project's own shorthand for the full stack says it in one line: CAPI provisions the clusters, KubeStellar places workloads across them, hive orchestrates the agents that manage the whole system.

The roadmap is already converging on this composition. KubeStellar's MCP server plans to drive intelligent placement through BindingPolicy from natural-language ops prompts, and its console ships an MCP agent interface for multi-cluster operations. "Tell the fleet where things should run" is becoming an API agents can call — which only raises the stakes on the placement layer being declarative and auditable rather than a shell history of kubectl commands.

The Honest Alternatives Table​

KubeStellar is one answer, not the only one. Before adopting any placement layer, weigh it against the field:

ApproachWhat it decidesMaturity signalWhen it wins
KubeStellar BindingPolicyWhich clusters get which objects, label-drivenCNCF Sandbox; v0.30 series in 2026; OCM-based transportYou want placement as declarative policy separate from both provisioning and GitOps
KarmadaScheduling + failover + autoscaling across clusters via a federated control planeCNCF graduated 2026 (v1.19); Bloomberg, Alibaba Cloud, Trip.com in productionYou want scheduler-grade semantics (failover, cross-cluster autoscaling, DR), not just propagation
ArgoCD ApplicationSetsWhich clusters a Git revision deploys toMature; ubiquitous GitOpsPlacement == "where does this repo revision go"; your fleet is already GitOps-driven end to end
Raw OCM (Subscriptions/ManifestWork)Transport + basic placement rulesCNCF graduated; the foundation KubeStellar builds onYou need cluster registration and manifest delivery and will write your own policy logic
Per-cluster kubectl / scriptsWhatever the operator typedTimeless; terribleOne cluster, or a demo that must never grow

Two honest notes on this table. First, Karmada's 2026 graduation — with strengthened cross-cluster scheduling aimed explicitly at distributed AI workloads — makes it the default enterprise answer, and a team that needs federated failover should start there, not here. KubeStellar's pitch is narrower and, for a small platform, arguably better shaped: it does placement policy without asking you to adopt a second full control plane. Second, ArgoCD and KubeStellar are complementary, not rivals: GitOps answers "what revision runs," placement answers "where it lands." Teams that try to express GPU-affinity policy in an ApplicationSet generator invariably end up reimplementing half a placement engine in YAML generators; that logic has a cleaner home in a BindingPolicy.

Should Placement Live in Your Platform or Stay Tenant BYO?​

The build-vs-borrow question from the top deserves a direct answer, because "adopt a placement engine" is real operational surface: another hub cluster, another CRD family, another reconciliation loop to monitor. Here is a decision framework with concrete thresholds:

Keep placement out of your platform (tenant BYO-scheduler) when: you run one cluster, or a handful of clusters that rarely change; tenants deploy through your git-push pipeline and never ask where things land; and no workload class needs hardware-specific targeting. A single well-run cluster with node pools, taints, and ResourceQuotas genuinely covers this. Do not adopt KubeStellar for a fleet of one — label-driven placement over a single member cluster is ceremony without payoff.

Promote placement to a platform primitive when two or more of these are true:

  1. Cluster count is growing past three, especially across regions or hardware types, and per-cluster deploy targets are already a source of toil or incidents.
  2. A workload class needs hardware targeting — GPU-backed agent sandboxes are the canonical trigger. The day a tenant asks "run this only where there is VRAM free," node selectors stop composing and cluster-level policy starts paying rent.
  3. Tenants run multi-cluster agents themselves. If your customers operate hive-style fleets (or plain Kubernetes jobs) that span clusters, "which cluster" becomes a question your API should answer declaratively instead of each tenant hand-rolling it.
  4. Blast-radius policy needs teeth. "Payment services never share a cluster with untrusted sandboxes" is a placement invariant; enforcing it by convention survives until the first 2 a.m. deploy.

For a small self-hosted PaaS, the pragmatic path is staged, not big-bang. Stage one costs almost nothing: enroll your CAPI clusters with consistent labels (region, accelerator, tenancy) today, so the option stays open. Stage two is one hub running KubeStellar with a single BindingPolicy for your most placement-sensitive workload — usually the sandbox tier — while everything else keeps its existing deploy path. Full migration of every workload to policy-driven placement is stage three, and most platforms should let stages one and two earn it before committing.

One more consideration that cuts toward adopting sooner: agent-operated fleets need machine-readable intent. An AI agent can read, diff, and propose a change to a BindingPolicy far more reliably than it can reconstruct "where things run" from five kubeconfigs and a runbook. If your roadmap says agents will operate this fleet within a year, the placement layer is not overhead — it is the interface your future operators (human or otherwise) will drive.

Placement Is the Primitive the Agent Era Needs​

Step back and the shape of the next few years is clear. Provisioning is solved and commoditized — CAPI and its provider ecosystem turned "give me a conformant cluster" into a reconciled API. Agent orchestration is arriving fast, with hive's autonomy-maturity model pointing at how fleets of coding agents get supervised rather than merely launched. The layer in between — deciding, declaratively and continuously, which cluster each workload lands on — is the one most self-hosted platforms have not built yet, because until agent sandboxes showed up, one scheduler per cluster was nearly enough.

It is not enough anymore. Whether you adopt KubeStellar's BindingPolicy, Karmada's federated scheduler, or GitOps-driven targeting that you grow into real policy later, make "which cluster" a declared, versioned, reconcilable decision. Your second cluster is coming; your fiftieth agent sandbox is coming faster. Build the layer that answers where they go before you are answering it by hand at 2 a.m.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide