Skip to main content

Your Build Cache Is Invisible to Kubelet: Sizing Image GC So Tenant Builds Stop Evicting Tenant Pods

9 min readDora NodaDora Noda
Share
On this page

A tenant pushes a big build at 3am. An hour later, a different tenant's pods are evicted with DiskPressure — while the node still holds gigabytes of build cache nobody asked it to keep. If your build nodes double as workload capacity, you have probably lived this incident: the kubelet evicted running workloads to relieve disk pressure it could not relieve, because the bytes filling the disk were never images in the first place.

Here is the one-line policy this post works into numbers: cap your build cache below the kubelet's image-GC low watermark, and derive every number from measured per-node churn instead of kubelet defaults. On a typical 80GB node that means a ~25GB BuildKit ceiling, GC thresholds tightened to 75/70, and ~21GB of burst headroom between steady state and hard eviction. The rest of this post shows exactly how those numbers are derived, the two config files that implement them, and the five gotchas that bite after tuning.

How kubelet image GC actually works​

The kubelet — not any cluster-level controller — owns image garbage collection. Every five minutes it checks filesystem usage on a fixed cadence you cannot configure, and only between two watermarks does it act at all:

KnobDefaultMeaning
imageGCHighThresholdPercent85%Above this disk usage, GC always runs
imageGCLowThresholdPercent80%GC deletes until usage falls to this
imageMinimumGCAge2mAn image must be this old before it is eligible
imageMaximumGCAge0 (disabled)If set, any image unused longer than this is collected (alpha since 1.29)
Deletion orderleast recently used firstOldest-unused images go first
Eligibilityunused images onlyAnything referenced by a pod is never touched

Three consequences fall out of that table. First, the default reclaim band is only 5% of the disk — on an 80GB node, GC starts at 68GB used and stops at 64GB. Second, GC is a slow drain: it runs every five minutes, while a single uncached multi-arch build can write 10GB in seconds. Third, and most important for this post: only container images are eligible. Build cache, layer working directories, container writable layers, and logs are all invisible to image GC. The kubelet documentation explicitly warns against bolting on external GC tools, so whatever image GC cannot see, nothing reclaims — until eviction starts killing pods.

Why the defaults lose on shared build-plus-workload nodes​

On a dedicated workload node, the defaults are fine: image churn is slow, the 5% band is plenty, and the five-minute cadence keeps up. Put tenant builds on the same node pool and three things break.

1. The reclaim band fills with unreclaimable bytes. Every tenant build leaves pulled base images and built layers plus BuildKit cache metadata on the node. The base images are GC-eligible; the build cache is not. So disk usage climbs past 85%, GC runs, deletes a few stale base images, usage barely moves because the bulk is cache — and the next build pushes it higher anyway. GC is bailing a boat with a hole in it.

2. The backstop is pod eviction, and it is closer than it looks. When nodefs and imagefs share one partition — the default on most self-installed nodes — the kubelet only enforces the nodefs signals. The default hard eviction threshold is nodefs.available<10%, i.e. eviction starts at 72GB used on an 80GB disk. GC's floor is 64GB. That leaves an 8GB margin between "GC has done all it can" and "start evicting BestEffort pods, then Burstable ones." One dependency-heavy build with cold cache crosses 8GB without noticing.

3. GC cadence cannot keep up with build bursts. Even if every byte were reclaimable, GC checks every five minutes. A build burst fills the margin in seconds; the kubelet's eviction manager, watching nodefs.available continuously, fires first. Image GC was designed for the slow trickle of rolling deploys, not the firehose of concurrent tenant builds.

The incident writes itself: builds fill the disk with cache GC cannot see, usage crosses 90%, DiskPressure taints the node, running tenant pods get evicted — and the eviction does not even fix the pressure, because evicting a pod does not delete the build cache that caused it.

The sizing method: measure first, then budget the disk​

Thresholds are the last step, not the first. Size from measurements:

Measure the three residents. On a representative node after a typical week:

bash
# Pinned workload images: what running pods actually reference (IMAGE, TAG, ID, SIZE)
crictl images
 
# Build cache footprint (BuildKit worker state)
buildctl du --verbose | tail -5
du -sh /var/lib/buildkit /var/lib/containerd 2>/dev/null
 
# Filesystem headroom and inode headroom, the two eviction signals
df -h /var/lib/kubelet && df -i /var/lib/kubelet

Watch the churn, not just the stock. The number that sizes your policy is gigabytes written per build-heavy day, not current usage. Sample nodefs usage from kubelet metrics across a week of CI traffic, note the largest single-day delta, and treat that as your minimum burst headroom. If your heaviest day writes 15GB of layers, any margin under 15GB between steady state and eviction is a future incident.

Partition the disk on paper before touching config. The rule: sustained usage (OS + pinned images + build-cache ceiling) must sit below the GC low watermark, and the GC low watermark must sit a full burst-day below hard eviction. Here is the worked budget for an 80GB disk with thresholds tightened to 75/70:

ResidentBudgetNotes
OS, kubelet, logs, container writable layers12GBMeasured, grows with log retention
Pinned workload images (running pods' bases)14GBFrom crictl images output
BuildKit cache ceiling (maxUsedSpace)25GBThe cap you enforce in buildkitd.toml
Sustained total51GB (64%)Must stay under the 70% / 56GB low watermark
GC band (70% → 75%)56–60GBGC reclaims here; bursts pass through
Burst headroom (75% → 90% eviction)60–72GB12GB+ for a heavy build day
Hard eviction (nodefs.available<10%)72GBPods start dying here

Sustained 51GB against a 56GB floor leaves 5GB of everyday slack; a 12GB burst day still stops at 63GB, inside the GC band where the next five-minute pass starts draining eligible images. That is the whole method: measure the rows, set the ceiling so the total clears the floor, and keep one burst-day between the floor and eviction.

The configs: kubelet thresholds plus a BuildKit ceiling​

Thresholds alone cannot fix this — the cache needs its own cap, because kubelet GC will never touch it. The two configs work as a pair.

KubeletConfiguration: tighten the band and age out CI churn. The 85/80 defaults assume slow churn; a build node wants GC to start earlier and finish lower, plus a maximum age so per-commit CI tags do not accumulate forever:

yaml
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
imageGCHighThresholdPercent: 75
imageGCLowThresholdPercent: 70
imageMinimumGCAge: 5m
imageMaximumGCAge: 168h  # collect anything unused for 7 days (1.29+)
evictionHard:
  nodefs.available: "10%"
  nodefs.inodesFree: "5%"
evictionMinimumReclaim:
  nodefs.available: "1Gi"

The 75/70 pair mirrors documented managed-platform practice — Red Hat's OpenShift image-GC guidance uses 75/70 with a ~5m30s minimum age — and the 7-day maximum age bounds the "thousand stale CI tags" failure without threatening warm base-image caches. (Set thresholds via the config file: the --image-gc-high-threshold CLI flags are deprecated.)

buildkitd.toml: cap the cache the kubelet cannot see. This is the actual fix; the kubelet tuning just buys margin around it:

toml
[worker.oci]
  gc = true
 
  [[worker.oci.gcpolicy]]
    # Normal operation: keep ~10GB of recent cache for fast rebuilds
    reservedSpace = "10GB"
    maxUsedSpace = "15GB"
    filter = ["unused-for=24h"]
 
  [[worker.oci.gcpolicy]]
    # Backstop: never let total cache exceed the budgeted ceiling
    all = true
    reservedSpace = "25GB"
    maxUsedSpace = "25GB"

Note the field names: BuildKit ≥0.16 (Docker Engine 28+) uses reservedSpace/maxUsedSpace; the older keepStorage/gckeepstorage keys are deprecated and silently ignored, so a policy copied from a 2023 blog post may currently be doing nothing. reservedSpace is the floor BuildKit keeps for build speed, maxUsedSpace the ceiling it enforces — set the ceiling to your budgeted row and the first rule keeps hot cache under it.

Verify the pair is holding. Three checks, cheapest first:

bash
# BuildKit honors the ceiling?
buildctl debug workers | grep -A5 -i gc
 
# Any recent pressure or eviction events?
kubectl get events -A | grep -E 'Evicted|DiskPressure'
 
# Node ephemeral-storage headroom after a build-heavy day?
kubectl describe node <node> | grep -i ephemeral

If DiskPressure events persist with the ceiling in place, the budget is wrong — re-measure the OS/writable-layer row, which is the one teams most often guess instead of measuring.

Five gotchas that bite after tuning​

Re-pull storms from an over-tight low threshold. Set the low watermark at 50% on a node whose warm base images alone are 40GB and every deploy re-pulls gigabytes — you traded evictions for slow deploys and registry egress. Watch image-pull latency (kubelet image_pull_duration_seconds) after tightening; if p99 pull time climbs, the floor is below your working set and needs to rise.

Per-commit tags defeat minimum age. CI systems that tag every commit (app:abc1234) create a stream of images that are unused within minutes but immortal until something collects them — that something is imageMaximumGCAge, which defaults to off. If your registry shows thousands of node-local tags, the max-age knob, not the thresholds, is the fix.

A separate imagefs partition is the structural fix. Everything above assumes one partition. Mounting /var/lib/containerd (or /var/lib/docker) on its own volume splits the signals: imagefs.available<15% guards images while nodefs guards everything else, and image pressure evicts by image reclaim instead of pod eviction. It is the right answer for a fleet you are still designing; the tuning in this post is the right answer for the fleet you already run.

Inodes evict too. nodefs.inodesFree<5% fires on file count, not bytes — and build cache plus extracted node_modules layers are inode-hungry. A node can show 60% byte usage and still evict on inodes. The df -i check in the measurement step exists for exactly this; if inode usage climbs with build traffic, raise the inode reservation or shorten cache retention.

Never run external pruning alongside the kubelet. A cronjobbed crictl rmi --prune or docker system prune races the kubelet's image manager: it can delete an image the kubelet is about to start a pod from, or fight the LRU accounting GC relies on. Upstream's guidance is explicit — do not use external GC tools. If the two-config setup above is insufficient, the answer is a bigger disk or a separate build pool, not a second garbage collector.

Cap the cache, then tune the thresholds​

The failure mode is a category error: image GC manages images, builds produce cache, and on a shared node the unreclaimable category grows until eviction punishes the wrong tenant. The fix order matters — measure churn, cap BuildKit below the GC low watermark, tighten 85/80 to something like 75/70 so GC starts earlier, then verify with events and pull latency. Revisit the budget whenever build traffic changes shape; a threshold tuned for last quarter's churn is next quarter's 3am page.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex