Every kubelet you run today is root on its host. Every container breakout of the last eight years — CRI-O tricked into setting kernel.core_pattern, runc writing to the host's procfs, the kubelet itself executing commands through a gitRepo volume — ended the same way: arbitrary code execution as host root, on a box running other tenants' pods. On September 4, 2026, Kubernetes promoted the fix it has been incubating since 2018 to beta: KubeletInUserNamespace, aka rootless mode, which runs the kubelet, the CRI and OCI runtimes, the CNI plugins, and kube-proxy as an unprivileged host user inside a Linux user namespace.
Here is the verdict first, because the headline overpromises if you skim it: beta means the feature gate is now on by default, but no node changes behavior until you deliberately put it in a user namespace. Nothing about your existing rootful clusters moves. The real question for a team that builds its own node images on owned hardware is what it takes to make rootless the default — and the answer is a compatibility checklist across cgroups, the runtime, networking, storage, and your audit tooling. This post is that checklist, with the CVE receipts for why it is worth clearing and a rollout order that starts where the risk is lowest.
What actually changed in v1.37
Rootless mode is not new. It started as Akihiro Suda's experiment in 2018, merged as an alpha feature gate in Kubernetes v1.22 (2021) under KEP-2033, and sat there for five years while the ecosystem caught up. The v1.37 "Garhwal" release — shipped August 26, 2026, with 67 enhancements — promotes it to beta, and Suda's announcement lists exactly three things that changed with the promotion:
- The gate is on by default.
KubeletInUserNamespaceno longer needs explicit enabling. But enabling the gate does not place the kubelet into a user namespace — that setup remains the operator's job — so existing rootful clusters see zero behavior change on upgrade. - Nodes now report their mode.
kubectl get nodes -o yamlexposesrunningInUserNamespaceinNodeSystemInfo, so you can label or taint rootless nodes and keep workloads that need real host root (some CNI plugin installers, for example) off them with ordinary scheduling constraints. - Upstream CI believes in it. Kubernetes' own node-conformance end-to-end suite now runs on a rootless cluster (
ci-kubernetes-e2e-kind-rootless), which is the difference between "a contributor's side project" and "a configuration a release blocks on."
Three outside-the-gate improvements made the timing work: kernel 6.3's idmapped tmpfs (2023), UserNamespacesSupport defaulting on in v1.33 so hostUsers: false pods need no extra flags, and containerd 2.1's writable cgroups (2025). Together they unlock rootless clusters nested inside user-namespaced pods — Kubernetes-in-Kubernetes without a single privileged: true.
One distinction the announcement stresses, because everyone conflates them: pod user namespaces (hostUsers: false, GA since v1.36) isolate pods while node components still run as root. Rootless mode isolates the node components. They compose rather than conflict.
The CVE receipts: why breakout-to-root is the threat
The announcement justifies the feature with five CVEs, one per year from 2022 to 2026, and the pattern is the point — every layer of the node stack has handed out host root:
| CVE | Component | What the attacker got |
|---|---|---|
| CVE-2022-0811 ("cr8escape") | CRI-O | Arbitrary sysctls such as kernel.core_pattern → code execution as host root |
| CVE-2023-27561 | runc | Masked-path bypass via a volume-mount race → host procfs exposure (a regression of CVE-2019-19921) |
| CVE-2024-10220 | kubelet | gitRepo volumes → arbitrary commands as root (echoing CVE-2018-11235 from 2018) |
| CVE-2025-31133 | runc | Attacker-controlled bind mounts → writes to host procfs (/proc/sysrq-trigger, core_pattern) |
| CVE-2026-53488 | containerd | Crafted image labels → arbitrary host commands |
Under rootless mode, each of these is still a vulnerability — the attacker escapes the container — but the prize shrinks from "host root" to "an unprivileged UID's account." They cannot rewrite the kernel, the bootloader, or the firmware to persist. For a multi-tenant PaaS packing several customers' apps per node, that is the difference between one tenant's breach and every tenant's breach.
How it works in 200 words
A Linux user namespace maps a host-level non-root user (say UID 1000) to UID 0 inside the namespace — a fake root whose privileges stop at the namespace boundary. That fake root is sufficient for nearly everything node components do: mounting volumes, creating cgroups, configuring pod network namespaces.
Two mechanics matter for planning. First, Kubernetes does not create the user namespace; something outside the cluster must — rootless Docker, Podman, nerdctl, LXC, or a raw unshare(2) with CLONE_NEWUSER plus helpers like RootlessKit. Your node image has to own that bootstrap.
Second, the feature gate itself is deliberately boring: it just teaches the kubelet to tolerate the permission errors a non-root process hits when setting node-level sysctls (vm.overcommit_memory, kernel.panic, and friends) and when watching kernel messages via /dev/kmsg, and teaches kube-proxy to tolerate the RLIMIT_NOFILE error. Everything else — the namespace, the cgroup delegation, the port forwarding — is ordinary Linux plumbing you assemble around it.
The adoption checklist for a self-built node image
This is the core deliverable: every item below comes from the official runbook for non-root node components. Clear all of them before a rootless node joins your fleet; any single miss is a node that fails in confusing ways.
1. Cgroups: v2 plus a delegated tree. Rootless requires cgroup v2 — v1 is explicitly unsupported — and a writable cgroup subtree delegated to the unprivileged user via a systemd unit with Delegate=yes. If your node image still boots cgroup v1 anywhere in the fleet, that migration is a prerequisite, not a parallel task.
2. Container runtime: fuse-overlayfs and the cgroupfs driver. For containerd (rootless CRI support since 1.4), the runbook requires disable_apparmor = true, restrict_oom_score_adj = true, disable_hugetlb_controller = true (systemd cannot delegate the hugetlb controller), the fuse-overlayfs snapshotter, and SystemdCgroup = false. Plain overlayfs without FUSE is possible on kernel 5.11+ only with SELinux disabled — note the tension if your image is SELinux-enforcing. For CRI-O (supported since 1.22), set _CRIO_ROOTLESS=1, use the overlay driver with the fuse-overlayfs mount program, and cgroup_manager = "cgroupfs".
3. Kubelet: cgroupfs and tolerated sysctls. Set cgroupDriver: "cgroupfs" in KubeletConfiguration and accept that node-level sysctl tuning (vm.overcommit_memory, vm.panic_on_oom, kernel.panic, kernel.panic_on_oops, kernel.keys.root_maxkeys/maxbytes) plus /dev/kmsg access silently do nothing. If your hardening baseline asserts those sysctls, move the assertion to the host outside the namespace.
4. kube-proxy: iptables or userspace, no conntrack tuning. Configure mode: "iptables" (or "userspace") with conntrack.maxPerCore: 0 and zeroed TCP timeouts so it skips the net.netfilter sysctls it cannot set. If you run eBPF-based networking (Cilium in kube-proxy replacement mode, for instance), that is a per-driver audit item, not a default — the runbook only blesses the two classic modes.
5. Networking: Flannel VXLAN is the known-good path. The node network namespace needs a non-loopback interface via slirp4netns, VPNKit, or lxc-user-nic, and the kubelet port (10250/TCP) plus every NodePort must be forwarded from the node namespace to the host with RootlessKit, slirp4netns, or socat. For pod networking the docs say plainly that some CNI plugins may not work and Flannel over VXLAN (8472/UDP) is known to work — Usernetes runs multi-node rootless clusters exactly this way. Privileged ports below 1024 need CAP_NET_BIND_SERVICE on the forwarder binary.
6. Storage: local volumes yes, network and block volumes no. Inside a user namespace only tmpfs, bind, and FUSE mounts work, so NFS, iSCSI, and block-device volumes are out, as is kernel-mode NFS. local, hostPath, emptyDir, configMap, secret, and downwardAPI are known good. FUSE-based CSI volumes are theoretically possible but explicitly not recommended. If your tenants expect PVC-backed Postgres on network storage, rootless nodes cannot serve those pods today — which is precisely what the new runningInUserNamespace taint is for.
7. SELinux interplay. The fuse-overlayfs requirement exists partly because non-FUSE overlayfs in a user namespace needs SELinux disabled. If your node image derives from a SELinux-enforcing distro, budget a test pass for the storage driver under your actual policy rather than assuming the containerd defaults transfer.
8. CIS and kube-bench expectations. The CIS Kubernetes Benchmark's node section (§4.1) asserts kubelet files — the service file, kubelet.conf, the config file — are owned by root:root. A rootless kubelet's files are necessarily owned by the unprivileged user, so a stock kube-bench run will flag what is actually your intended design. Treat those findings as "document the exception and assert the replacement control" (the kubelet user owns nothing else on the host), not as failures to chase to zero.
9. Taint and label by mode. Wire runningInUserNamespace into your node provisioning so rootless nodes carry a taint (or at minimum a label) from birth. Stateful tenants, CNI installers, and anything mounting host paths stays on rootful nodes by default until each workload class is individually cleared.
Clear those nine and a rootless node is a stricter, quieter kubelet. Skip the audit and you get the worst of both: neither the compatibility of rootful nor the confinement of rootless.
Where to run it first
Do not start with the fleet default. Start where the blast radius is smallest and the threat model fits best:
- Cluster API bootstrap clusters first. The announcement lists bootstrapping as a first-class use case: a temporary unprivileged cluster that mints your real clusters. An ephemeral bootstrap control plane that never needs host root is the cheapest possible pilot — it exercises your user-namespace plumbing end to end with zero tenant traffic at stake.
- Untrusted-workload pools second. CI runners, AI coding-agent sandboxes (the announcement explicitly names an agent deceived by malicious internet content as the scenario), and preview environments are exactly the workloads whose breakout you most want confined to a UID. Taint these nodes rootless-only and let scheduling do the segregation.
- Try it locally with zero commitment.
kindon rootless Docker/Podman/nerdctl, minikube with the Docker driver, multi-node Usernetes over Flannel VXLAN, or experimental rootless k3s (the only path that needs no external runtime) each get you a rootless cluster in an afternoon. Validate your CNI choice there before touching a node image.
The fleet-default conversation comes after the per-driver CNI/CSI audit and only once your storage story has an answer for network volumes — for most self-hosted platforms, that makes rootless a 2027 default, not a 2026 one.
What beta does not promise
Three honest limits, all from the announcement itself. User namespaces do not mitigate kernel vulnerabilities — if the exploit is in the kernel, the namespace boundary is the thing being exploited — so rootless composes with seccomp, AppArmor/SELinux, and read-only root filesystems rather than replacing them. The user-namespace lifecycle (creation, cgroup delegation, port forwarding, upgrades of the plumbing) is yours to operate; the gate only covers kubelet's and kube-proxy's error tolerance.
And beta is not GA: the project says general availability depends on feedback and adoption, with KEP-5474 (writable cgroups for unprivileged containers) and KEP-5714 (cgroup namespace control) shaping the Kubernetes-in-Kubernetes future. File issues when the plumbing bites — that feedback is literally the GA criterion.
Eight years from experiment to beta is slow. It is also what "this changes the security boundary of every node" deserves: the careful version, with conformance CI behind it and a checklist instead of a flag flip. Build the checklist into your node image pipeline now, and the GA announcement becomes a taint removal rather than a project.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



