Skip to main content

Kubernetes 1.36's Sharded List and Watch: the N× Watch Tax It Kills, and the Ceiling It Doesn't Touch

9 min readDora NodaDora Noda
Share
On this page

Run four replicas of a sharded controller against a big cluster and you pay for the same watch stream four times. Every replica deserializes every event, throws away the three-quarters it doesn't own, and the API server ships the full firehose down every single connection. Scaling out the controller doesn't divide that cost. It multiplies it.

Kubernetes 1.36 ("Haru," released April 22, 2026 with 70 enhancements: 18 stable, 25 beta, 25 alpha) ships the fix as an alpha feature: server-side sharded list and watch (KEP-5866). Each replica tells the API server which hash range it owns, and the server filters at the source, so every replica receives only its slice.

But read the name carefully, because it oversells. This feature does not shard the API server. It does not partition etcd or split list-serving across backends. It shards controller fan-out: N replicas that used to cost N× the stream now cost 1× combined. That distinction decides exactly which scalability ceiling moves for your control plane, and which one stays exactly where it was. This post gives you the math, the mechanism, and the operator's bill for turning it on.

The N× watch tax, in one table​

Call S the full serialized stream for one watched collection: a LIST snapshot plus every subsequent watch event. Today, a horizontally "sharded" controller like kube-state-metrics still makes every replica swallow all of S. The waste is exact linear math, not a benchmark estimate:

Replicas (N)Total API-server egress, client-side shardingTotal egress, server-side shardedPer-replica deserialize work
11 × S1 × SS (no waste either way)
22 × S1 × SS → S/2
44 × S1 × SS → S/4
88 × S1 × SS → S/8

Two things to notice. First, at N=1 nothing changes: there is no waste to kill, so a single-replica watcher gains exactly zero from this feature. Second, the savings grow with both replica count and collection size, because S itself scales with object count, object size, and churn. A worked example with clearly labeled illustrative assumptions makes that concrete.

Assume a 10,000-pod collection with an average serialized pod object of 1.5 KiB, so one full LIST snapshot S is roughly 15 MB. Assume four kube-state-metrics-style replicas relisting every five minutes: unsharded, that is 4 × 15 = 60 MB of list traffic per cycle versus 15 MB sharded. Now add churn: at 50 pod events per second averaging 1.5 KB each, the steady-state watch stream is about 75 KB/s, and every unsharded replica pays all of it (300 KB/s combined) versus about 19 KB/s each sharded. Double the pods or double the churn and S doubles with it; the sharded column stays at 1 × S no matter how many replicas you add.

The honest caveat: upstream has published no measured benchmarks yet. Benchmarks showing real client-side wins, plus scalability tests proving no API-server throughput regression, are explicit beta graduation criteria for the KEP, which means they don't exist at alpha. Size your expectations from the formula above, not from a graph nobody has published.

How it actually works​

The client passes a new shardSelector field in ListOptions, written in a small CEL-based grammar built around one function:

text
shardRange(object.metadata.uid, '0x0000000000000000', '0x8000000000000000')

The API server computes a deterministic 64-bit FNV-1a hash of the named field and returns only objects whose hash falls in [start, end). The filter applies to both LIST responses and watch event streams, and the hash is stable across API-server instances, so the feature is safe behind multiple apiserver replicas. Exactly two field paths are supported in alpha: object.metadata.uid and object.metadata.namespace.

Server-side, the implementation extends the Cacher's SelectionPredicate with hash-based filtering: each event's key field is extracted, hashed, range-checked, and either dispatched to the watcher or dropped. Nothing about etcd reads changes; the filtering happens after the watch cache, per watcher, at dispatch time. That placement is precisely why this is a fan-out optimization and not storage sharding.

Controllers adopt it per informer through client-go's WithTweakListOptions:

go
shardSelector := "shardRange(object.metadata.uid, '0x0000000000000000', '0x8000000000000000')"
factory := informers.NewSharedInformerFactoryWithOptions(client, resyncPeriod,
    informers.WithTweakListOptions(func(opts *metav1.ListOptions) {
        opts.ShardSelector = shardSelector
    }),
)

A two-replica deployment splits the 64-bit space in half; one replica can also cover non-contiguous ranges with ||. The feature is alpha behind the ShardedListAndWatch feature gate, which must be enabled on the API server. UID is the natural default key: Kubernetes UUIDs are already uniformly distributed, and hashing on top keeps the spread even if IDs ever go sequential (uuidv7 is the KEP's example). Namespace-keyed sharding is the tenant-shaped alternative, with the skew risk you'd expect: one giant namespace makes one giant shard.

Which ceiling moves, which doesn't​

A growing single-cluster control plane has (at least) two distinct watch-path ceilings, and this feature moves exactly one of them.

Moves: the watch-fan-out ceiling. This is the cap on how many watch-heavy controller replicas one control plane can feed before API-server egress, serialization CPU, and client-side deserialize cost become the binding constraint. It bites multi-tenant platforms first, because tenant count multiplies both the watched collections and the controllers watching them. Server-side sharding divides per-replica cost by N and holds total egress flat at 1 × S, so it directly raises the number of replicas — and therefore tenants and controllers — a single control plane serves before you must split clusters to escape fan-out cost.

Doesn't move: the etcd and watch-cache ceiling. Full-collection reads out of etcd on cache initialization and cache-miss fallback, etcd write throughput, and the cost of serving clients that can't shard (kube-controller-manager still scales vertically; sharding it is an explicit KEP non-goal) are all untouched. If your pain is huge LIST latency, cache-init memory spikes, or etcd disk and I/O pressure, sharded watch will not help, because that traffic never reaches the dispatch filter that got cheaper.

The complementary fix for that other ceiling already exists one release later: Kubernetes 1.37 graduates etcd RangeStream to beta (enabled by default with etcd v3.7), streaming large etcd reads in chunks instead of assembling whole responses in memory. The two features attack opposite ends of the same path: RangeStream cheapens the API-server-read side, sharded watch cheapens the controller-fan-out side. When you plan the multi-cluster-split decision, frame it as whichever-ceiling-hits-first: fan-out saturation argues for sharded controllers on one cluster longer, while etcd-side saturation argues for RangeStream now and a split later regardless.

The operator's bill: what turning it on costs​

Alpha features don't just cost a feature gate. Here is the full checklist before this earns a place in a production runbook:

  1. Gate every apiserver. --feature-gates=ShardedListAndWatch=true on each kube-apiserver; alpha means off by default. During a mixed-version rollout, old servers silently ignore the unknown query parameter and send the full, unsharded stream, so a half-upgraded control plane quietly un-shards you.
  2. Opt in every controller, per informer. There is no framework-level support yet: informer/reflector integration is a beta criterion, so each sharded controller must inject shardSelector itself via WithTweakListOptions. Audit your controllers (kube-state-metrics' --shard/--total-shards model is the canonical client-side pattern this replaces) and assume nothing adopts it automatically.
  3. Verify shardInfo, or you may be unsharded. When the server honors your selector, the LIST metadata echoes it back in a shardInfo field. If shardInfo is absent, you received the complete collection and must fall back to client-side filtering. Treat a missing shardInfo as an alert-worthy condition, not a quiet edge case. Server support is also discoverable through the OpenAPI v3 discovery document.
  4. You own the shard math and every rebalance. Coordination and resharding are explicit KEP non-goals: you assign ranges, and your ranges must tile the full 2^64 space with no gaps and no overlaps. A gap silently drops objects from every replica's view; an overlap double-processes them. Scaling from 4 to 5 replicas means recomputing and rolling out five new selectors yourself. Range-prefix partitioning at least keeps each reshard moving only a fraction of the keyspace.
  5. Pick the shard key per workload. UID-hash spreads evenly and is the default. Namespace-keyed sharding maps neatly onto per-tenant replicas but inherits tenant-size skew. Either is better than client-side discard, but neither rebalances itself.
  6. Budget for alpha churn. The grammar, supported fields, and shardInfo contract can all change before beta. The hash and range-evaluation logic lives in the shared k8s.io/apimachinery library so your client-side fallback computes exactly what the server computes, but pin your client-go and read the release notes on every upgrade.
  7. Load-test your apiserver. Per-watcher hashing is cheap FNV-1a, but it runs per event per sharded watcher, and the no-regression proof is itself a pending beta criterion. Confirm serialization savings outweigh filter cost in your own load tests before declaring victory.

Should a small self-hosted fleet enable it now?​

Use a simple decision rule: multiply replicas × watched-collection size × churn for your watch-heaviest controllers. If that product is small — single-digit nodes, one kube-state-metrics, a handful of operators watching mostly quiet resources — the N× tax rounds to zero and this feature buys you nothing. Wait for beta and informer-framework support, then adopt it by upgrading controllers rather than hand-wiring selectors.

It wins today for exactly the workloads the KEP names: kube-state-metrics-class exporters at scale, per-tenant operator replicas watching cluster-scoped high-churn resources like Pods, and any controller you would scale horizontally if the full-stream penalty didn't punish every added replica. For a Cluster-API-managed fleet on owned machines, that makes it a single-cluster-life-extension primitive: it pushes the fan-out-driven multi-cluster split further out without changing your etcd sizing math one bit. Know which ceiling you're actually hitting, enable the gate for the workloads that pay the N× tax, and leave the rest alone until beta.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex