Skip to main content

etcd 3.7 RangeStream: What Streaming List Reads Mean for a Cluster API Fleet

10 min readDora NodaDora Noda
Share
On this page

Every controller on your management cluster starts its day the same way: with a LIST. The Cluster API controllers that reconcile your Machines, the scheduler, your GitOps agent — each one opens by asking the API server for the full state of the world, and the API server answers by asking etcd for the same. On a small fleet that handshake is invisible. On a fleet with thousands of Machines and tens of thousands of Pods, it is a memory spike on the most fragile box you own: the control-plane node.

Here is the verdict up front. etcd v3.7.0, released July 8, 2026, ships a server-streaming RangeStream RPC that returns large result sets in chunks instead of buffering the whole response. Kubernetes v1.37 pairs with it through the EtcdRangeStream feature gate — beta and on by default — so watch-cache initialization and uncached list reads stream instead of spiking.

If you run a Cluster API fleet, this is a real reduction in control-plane memory pressure as machine counts grow. But be precise about where the win lands: etcd and the API server get leaner during large reads, while your controllers' own informer caches still prefetch everything they watch. RangeStream fixes the serving side of every LIST, not the holding side.

The cost of large reads: what unary Range actually holds in memory​

To see why streaming matters, follow one large read through the old path. The API server keeps most list and watch traffic in its in-memory watch cache, but populating that cache — at startup and on every re-initialization — requires reading a resource's full state out of etcd. For a resource like Pods, with many objects, large objects, or both, that read is expensive.

The API server already paginated these reads, asking for a fixed number of keys at a time. But a page bounded by key count has no awareness of object size, so a page of large objects can still be very large.

Worse, the cost is paid twice over. As KEP-5966 documents, etcd's unary Range assembles each page in full before sending — the KV slice, the serialized protobuf, and the gRPC send buffer all coexist in server memory — and the API server then holds the page while decoding it. The same payload sits in memory on both sides at once, and a bad combination of object size and concurrent reads is enough to trigger an OOM. Most of that cost lands on etcd, which is also where streaming helps most.

Pagination has a second, quieter tax. Every paginated page re-walks etcd's full B-tree index to recompute the total count, so a client paging through N pages pays for N full index traversals. The work grows with the number of pages, not with the number of keys returned.

This is not a hypothetical failure mode. A September 2026 production anecdote describes a cluster whose operators never pruned finished Jobs: etcd grew to 6.4 GB, list calls against large namespaces started paging for tens of seconds, and scheduler resyncs stretched into minutes. Every one of those retained objects was a row in etcd that the API server kept in a watch cache and that every controller listed on resync. On a Cluster API management cluster, the equivalent pile-up is Machine, MachineSet, and infrastructure objects accumulating across a growing fleet — each one taxing the same read path.

How RangeStream works: same request, chunked stream​

etcd v3.7 adds a streaming version of the range read. The RangeStream RPC takes the same RangeRequest as Range and returns the same result set, but instead of building the whole response up front, etcd splits it into chunks and streams them. The design, specified in KEP-5966, has four properties worth knowing:

Adaptive, byte-bounded chunks. The first chunk uses a small key-count limit; after each chunk the server doubles or halves the limit based on the previous response size relative to a target derived from MaxRequestBytes. Chunk sizing converges on the observed value sizes without the client guessing, so a collection of large objects is bounded by bytes rather than by key count. Memory is freed as the stream progresses instead of being held until a whole page is assembled.

One revision for the whole stream. The server pins a single MVCC revision after the first chunk and reuses it for every subsequent chunk, so the merged stream is snapshot-consistent. The total count comes from a running tally of streamed keys, eliminating the per-page B-tree walk that pagination paid. The contract is simple: merging chunks sequentially in stream order produces a result identical to one unary Range call.

Short-lived storage transactions. Instead of holding one long-running bbolt read transaction for the entire stream, the server opens a new short-lived read transaction per chunk. That sidesteps the bbolt caveat where a long-lived read transaction can block write transactions when the database needs to remap or allocate pages — a meaningful detail on a busy control plane where reads and writes contend.

A retryable compaction race. If the pinned revision is compacted mid-stream, the server closes the stream with ErrCompacted and the client retries. The API server surfaces this as a watch-cache initialization failure and the cacher retries automatically — behavior identical to a paginated list racing compaction. In practice this only bites at absurd scale: if a stream takes longer than the API server's 5-minute default compaction interval to complete, the equivalent unary response would be approaching protobuf's 2 GB message limit anyway.

RangeStream is available on both the gRPC API and via etcdctl, and real clients are already adopting it: the Milvus BirdWatcher tooling upgraded to etcd v3.7.1 specifically for the streaming RPC, noting that a limit=1 request now stops the server index traversal right after the limit instead of scanning.

What changes on a Cluster API fleet​

This is the section the headline promised. Concretely, here is what a self-hosted, Cluster-API-managed fleet gets, what it must run to get it, and where the boundary lies.

The version matrix is a conjunction, not a choice. Streaming requires etcd v3.7 or later and Kubernetes v1.37 or later, with the EtcdRangeStream gate enabled on kube-apiserver — which it is by default, since the gate is beta in 1.37. The API server resolves etcd's support at startup and also falls back at runtime if a call returns Unimplemented, so pairing a new API server with an older etcd silently keeps the paginated path. Upgrade etcd first, then the API server, or you will be running the new code on the old path without knowing it.

Two read paths get streaming. The primary integration point is watch-cache initialization: the cacher's sync() consumes KV.GetStream and converts each chunk's key-value pairs into synthetic created events queued inline, without ever assembling the full list. The same treatment applies to direct GetList calls — the fallback path where a list cannot be served from cache and reads etcd directly, which is exactly what a controller's uncached LIST hits. The store decodes each chunk inline as it arrives, overlapping network I/O with decode, and storage.Interface is unchanged.

Large read on the fleetBefore (paginated unary Range)With etcd 3.7 + K8s 1.37
Watch-cache initFull page buffered on etcd and API server, per pageChunks decoded into synthetic created events inline
Uncached controller LISTSame double-buffering, plus a recount per pageStreamed chunks, count from a running tally
Controller informer cacheFull prefetch per watched kindUnchanged — still full prefetch (WatchList is the other half)

Verify with one metric. The API server records streamed reads under their own operation label. A non-zero count means RangeStream is in use:

promql
etcd_request_duration_seconds_count{operation="listStream"}

If it stays at zero after your upgrade, the API server is still on the paginated path — most likely because etcd is older than v3.7. Put this query on your fleet dashboard next to etcd DB size and API server memory; it is the cheapest possible confirmation that the upgrade actually changed the read path.

Now the honest boundary: your controllers' caches do not shrink. controller-runtime informers are not lazy loaders — the first request for a kind starts an informer that prefetches everything of that kind into local memory. RangeStream makes the API server cheaper to ask, but the CAPI controller that watches ten thousand Machines still holds ten thousand Machines.

Teams running large management clusters already fight this directly: one Cluster API provider commit, for example, disabled caching for ConfigMaps entirely after cluster-wide informers caused significant memory growth and OOM kills. That class of fix — scoping caches by namespace, using metadata-only watches, disabling caches for hot types — is unchanged by RangeStream. The client-side counterpart to this server-side streaming is WatchList (KEP-3157), which replaces the informer's initial LIST with a streaming watch; RangeStream and WatchList are complementary halves, not substitutes.

One more v3.7 bonus for metadata-heavy controllers. The same release ships a keys-only range optimization: when a request asks for keys only, etcd answers from its in-memory index without loading serialized values from bbolt. Controllers that enumerate key names — inventory sweeps, existence checks, garbage collection — get cheaper reads with no code change.

Adoption checklist for a small fleet​

For a platform running a handful of owned machines rather than a hyperscaler's fleet, the upgrade is small but order matters:

  1. Upgrade etcd to v3.7 first, one member at a time with health checks between steps, per the release upgrade guide. This release removes legacy v2 components (v2 discovery, v2 request handling, the v2 client) and all deprecated --experimental-* flags — migrate to feature gates or stable flags before upgrading, and note that official images ship multiarch-only now.
  2. Upgrade control planes to Kubernetes v1.37, leaving EtcdRangeStream at its default-on. To opt out explicitly, set --feature-gates=EtcdRangeStream=false on kube-apiserver.
  3. Confirm listStream is non-zero on every management cluster before declaring victory.
  4. Revisit control-plane memory headroom only after observing the new steady state. RangeStream makes peak usage predictable rather than promising a fixed percentage cut — size from your own listStream traffic and etcd memory graphs, not from a blog post's rule of thumb.
  5. Audit Go clients that vendor etcd modules. The v3.7 protobuf overhaul replaces gogo/protobuf with the supported google.golang.org/protobuf, and client creation no longer honors the deprecated blocking dial option. Binaries are unaffected, but anything compiling against etcd's Go SDK may need updates.

Also worth knowing: v3.7 prioritizes lease-revocation requests under overload, speeds up lease keepalives, accelerates concurrent-watch lookups, and adds watch send-loop metrics plus a new etcd_server_request_duration_seconds metric. None of these is the headline, but together they make the "etcd under load" story — the story every growing CAPI fleet eventually lives — noticeably calmer.

The control plane gets quieter as the fleet gets louder​

RangeStream is the rare infrastructure improvement with no new API to learn and no migration to plan: the same request, the same result set, delivered in chunks that keep both sides of the connection out of OOM territory. For a Cluster API fleet, it directly attacks the cost curve that machine growth imposes on the management cluster's etcd and API server — watch-cache rebuilds, controller resync storms, and uncached LIST fallbacks all stop being memory cliffs. Pair it with WatchList on the client side and disciplined cache scoping in your own operators, and the control plane stays boring while the fleet scales. That is the whole game in fleet operations: keep the machinery quiet so the machines can multiply.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex