Until March 2026, every combination of runner attributes in Actions Runner Controller needed its own scale set. Linux plus GPU plus PCI-compliant network zone was not one pool with three labels — it was a separate AutoScalingRunnerSet, a separate Helm release, and a separate listener pod, multiplied across every tenant you served. A platform team offering three operating systems, two hardware tiers, and two compliance zones was staring at twelve scale sets before a single tenant job had run.
ARC 0.14.0, released March 19, 2026, ends that sprawl with multilabel support: one scale set carries many labels, and workflows target runners by combining them. The same release swaps ARC's internals onto the standalone actions/scaleset Go client — the module GitHub put in public preview a month earlier so platform teams can build scale-set autoscaling without Kubernetes at all — and adds stale-config protection that stops autoscaling when a runner set's configuration is outdated.
This post is the walkthrough: what it actually takes to run ARC's runner-scale-set operator on a Cluster-API-managed fleet to give tenants their own git-push-triggered CI. The short version of the architecture, up front so the rest is grounding rather than suspense: a cluster-wide controller watches AutoScalingRunnerSet objects; each one owns a long-lived listener pod holding a message session open to GitHub; when a tenant pushes, the listener picks up the queued job, the controller mints a just-in-time runner registration, and an ephemeral runner pod runs exactly that one job and is deleted. Scale to zero between pushes, no shared long-lived runner, no per-minute meter from anyone.
The moving parts, concretely
ARC's modern architecture — the gha-runner-scale-set mode, which is the only mode that matters for new deployments — has four pieces. First, the gha-runner-scale-set-controller Helm chart installs the cluster-wide operator, typically into an arc-systems namespace. It reconciles the custom resources and owns all communication with GitHub's scale-set APIs.
Second, each runner pool is a gha-runner-scale-set Helm release, which creates an AutoScalingRunnerSet custom resource. This is the object you tune: which GitHub scope it serves, how many runners it keeps warm, how high it can burst. One tenant pool, one release — and after 0.14.0, one release can wear all the labels that pool needs.
Third, each scale set runs a listener pod. The listener is the only long-lived runner-adjacent process: it holds a persistent message session to the GitHub Actions service and reacts to queued jobs and scale signals. It does not run builds. Think of it as the pool's ear, not its muscle — which is why its scheduling requirements differ from the bursty runner pods it triggers.
Fourth, the ephemeral runners themselves. Each job gets a fresh runner pod registered with a just-in-time token that is minted for that job and useless afterward. When the job finishes, the pod goes away. There is no runner agent accumulating state, disk cruft, or leaked secrets across jobs — the property that makes per-tenant CI on shared node pools defensible in the first place.
What multilabel changes, before and after
Before 0.14.0, ARC matched workflows to scale sets on exactly one label: the scale set's name. A workflow wanting a large Linux GPU runner in a compliant zone needed a scale set literally named for that combination, so operators multiplied scale sets across the cross product of attributes. Each one carried its own listener pod, its own Helm values, and its own min/max tuning to drift out of sync.
After 0.14.0, one scale set carries many labels:
# One scale set, many attributes (ARC >= 0.14.0)
runnerScaleSetName: tenant-acme-pool
scaleSetLabels:
- self-hosted
- linux
- x64
- largeand a workflow targets the combination directly:
jobs:
build:
runs-on: [self-hosted, linux, x64, large]For tenant CI this collapses the topology from "a scale set per attribute combination per tenant" to "a scale set per tenant pool." The labels describe the capacity; the scale set's GitHub scope describes who may use it. That separation is what makes the per-tenant walkthrough below stay legible instead of exploding into a matrix of Helm releases.
Walkthrough: tenant CI on a Cluster-API-managed Hetzner fleet
Start with the honest grounding: ARC is provider-agnostic and runs on any conformant cluster, so nothing here is Hetzner-specific. What matters for a Cluster-API-managed fleet is the node-pool split you already operate — typically a small system pool for control-plane-adjacent workloads and larger workload pools for tenant compute — because the listener pods and the ephemeral runner pods want different homes, and both should land on worker capacity you own rather than metered runners you rent.
1. Install the controller once
helm install arc-controller \
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set-controller \
--namespace arc-systems --create-namespaceThis is cluster-scoped and tenant-blind. One installation serves every scale set on the fleet.
2. One scale set per tenant scope
Each tenant pool gets its own gha-runner-scale-set release. The binding between a tenant and a pool is githubConfigUrl — the repo or org URL whose jobs this pool serves — plus the labels workflows use in runs-on:
helm install tenant-acme-runners \
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set \
--namespace tenant-acme --create-namespace \
--set githubConfigUrl=https://github.com/acme-corp \
--set runnerScaleSetName=tenant-acme-pool \
--set minRunners=0 --set maxRunners=20 \
--set containerMode.type=kubernetesminRunners=0 is the scale-to-zero posture: between pushes the pool costs you nothing but its listener pod. maxRunners is your blast-radius cap per tenant — twenty concurrent jobs, not two hundred, no matter how enthusiastic their matrix build gets. containerMode.type=kubernetes runs each job step in its own container; if tenants build container images inside CI, you need containerMode.type=dind instead, with all the privileged-container trade-offs that implies (more in the gotchas). Note that Kubernetes mode also needs a dynamically provisioned work volume per runner pod, so your fleet must have a working StorageClass and provisioner before the first job lands — on a self-hosted fleet, that is a real prerequisite, not a default.
3. Follow one push end to end
Here is the concrete flow for a single tenant repo, with nothing hand-waved:
- A developer pushes to
acme-corp/webapp. GitHub queues the workflow job against the labels inruns-on: [self-hosted, linux, x64]. - The
tenant-acme-poollistener pod, holding its long-lived message session to GitHub, receives the job assignment because its scale set carries exactly those labels. - The controller mints a just-in-time runner configuration — a single-use registration credential — for that job.
- An ephemeral runner pod spins up on a workload-pool node, registers with the JIT config, runs the one job, then is deleted. Nothing persists to the next job: no
/tmp, no daemon state, no credentials.
The tenant's mental model stays exactly the GitHub-native one — push, watch the checks run — while the compute is your fleet's spare capacity and the isolation boundary is a fresh pod per job.
4. Schedule the listener against existing pools
The listener is long-lived and light; the runners are bursty and heavy. On a fleet with a system/workload pool split, the listener usually belongs with the system-adjacent workloads (stable, always-on, cheap to place), while runner pods burst onto the workload pools where tenant compute already lives. 0.14.0 helps here in a small but real way: the listener pod now ships with a default nodeSelector of kubernetes.io/os: linux, so in a mixed-OS cluster it can no longer land on a Windows node and fail on platform incompatibility. You can still override or remove the default through listenerTemplate — the escape hatch matters if you pin listeners to dedicated nodes with your own tolerations.
5. Isolate tenants by scope, not by trust
Tenant isolation in this design comes from three compounding boundaries: a separate namespace and Helm release per tenant pool, a githubConfigUrl scoped to that tenant's org or repos so no other tenant's jobs can route to the pool, and per-job ephemeral pods so no job ever shares a filesystem or process space with another tenant's job. Compared with the old pattern of one shared long-lived self-hosted runner per team — a pet server accumulating state and secrets — this is the difference between multitenancy by convention and multitenancy by construction.
Stale-config protection, with its real limitation
The quietest important change in 0.14.0 is that ARC can now fully stop autoscaling for a runner set whose configuration is outdated. When a runner exits with exit code 7, the controller switches off autoscaling for that runner set, so stale runners stop provisioning while a new configuration rolls out — no more jobs landing in a half-updated environment during a runner-image change.
The honest caveat, straight from the changelog: this depends on a runner-side change (the exit-code-7 signal) shipping in a future runner release, and under the runner version support policy the feature won't become fully effective until two releases after that. So treat it as a committed direction, not a switch you can flip today — your rollout hygiene for runner-image changes still needs its own care until the runner fleet catches up. (The release also lets you set custom Kubernetes labels and annotations on ARC-managed internal resources via resource.all.metadata, plus experimental rewrites of both Helm charts — useful, but not the reason to upgrade.)
The escape hatch: scale sets without Kubernetes
The most structurally interesting line in the 0.14.0 changelog isn't a feature at all: ARC now uses the standalone actions/scaleset Go module as its sole client for the GitHub Actions service APIs, with the old internal client removed. That module — MIT-licensed, in public preview since February 2026 — packages scale-set registration, message sessions, and JIT runner configuration as a library anyone can build on, aimed explicitly at "containers, virtual machines, and bare metal."
GitHub's own docs frame it carefully: the client is not a replacement for ARC, which remains the reference implementation and the recommended Kubernetes solution — it is the complementary tool for the same APIs outside Kubernetes. And teams are already using it that way: the shader-slang project runs a Go-based GCP VM scaler on actions/scaleset in production since late February 2026, community Proxmox VM-template scalers build on it, and the Terraform AWS runner project is implementing scale-set orchestration against the same protocol.
When should a self-hosted platform choose the library over ARC? If your runner substrate is already Kubernetes, don't — ARC is the maintained path and 0.14.0 just made it strictly better. Reach for actions/scaleset when your CI capacity lives somewhere Kubernetes doesn't: Proxmox hosts, bare-metal provisioning, a VM fleet managed by Terraform, or GPU workstations that will never join the cluster. The decision is about where your spare capacity lives, not about which GitHub API you prefer — it's the same API either way now.
Where this sits next to the metering churn
The cost context, briefly, because sibling posts own the full math: GitHub cut hosted-runner prices up to 39% on January 1, 2026 (standard Linux 2-core from $0.008 to $0.006 per minute), and the proposed $0.002-per-minute platform charge on self-hosted runners was shelved indefinitely before it ever took effect — we covered the breakeven analysis (owning wins past roughly 3,000 metered minutes a month) and the charge reversal separately.
ARC is the operational answer to the question that math raises. If your CI volume sits above the breakeven and the compute is already yours — spare capacity on a fleet you're running anyway — then the remaining cost of self-hosted CI is purely operational: installing, upgrading, and babysitting the runner substrate. Multilabel support, stale-config protection, and the extracted client each shrink a different slice of that operational cost. None of them changes the per-minute price of anything; they change how many engineer-hours the own-side column of the spreadsheet needs.
Honest gotchas before you ship tenant CI on this
- Docker builds need privileged mode. In default
kubernetescontainer mode, runner pods cannotdocker build. Tenants that build images inside CI needcontainerMode.type=dind(Docker-in-Docker sidecar, privileged) — which reintroduces exactly the node-level trust questions ephemeral pods were supposed to dissolve. Scope dind pools narrowly and keep them off nodes you care about most. - Scale-to-zero has a cold start. A job arriving at an idle pool waits for pod scheduling, image pull, and JIT registration. Pre-warming with
minRunners=1during business hours is the standard compromise; measure your p50 queue-to-start before promising tenants hosted-runner latency. - The runner support policy lags features. As the exit-code-7 caveat shows, controller features that need runner-side changes take multiple runner releases to become effective. Pin your runner image versions deliberately and track the support window — "latest ARC" does not mean "every ARC feature."
- GitHub's API is still in the loop. ARC hammers GitHub's scale-set APIs all day; rate limits, API changes, and GitHub-side incidents are your incidents now. The shelved platform charge was explicitly motivated by this control-plane cost, so budget for the possibility that orchestration gets metered again in some future form.
- Windows and GPU pools are separate lanes. Multilabel combines attributes within a scale set, but Windows runners and GPU runners still need their own pools, node selectors, and images. The sprawl fix compresses the matrix; it doesn't eliminate the genuinely different machine types.
ARC 0.14.0 doesn't change what self-hosted CI costs. It changes what it costs to operate: fewer scale sets to drift, a config-staleness backstop on the way, and — via the extracted client — a supported path for the capacity that will never be a pod. For a platform team already running a Cluster-API fleet, tenant CI as ephemeral per-job pods is now a Helm release per tenant plus scheduling discipline, not a bespoke autoscaler project. That is a meaningfully smaller box to own than it was a year ago.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



