Until this month, no Kubernetes policy could stop a compromised container from downloading a binary to a writable volume, running chmod +x, and executing it. Not an OPA rule, not a Kyverno policy, not Pod Security Standards. The enforcement point simply did not exist: volumes were bind-mounted into containers without noexec, nosuid, or nodev, and no API field could say otherwise. On September 16, 2026, Kubernetes v1.37 created that enforcement point — bindMountOptions on volume mounts and a mode field on emptyDir — and moved a whole class of volume rules from "approximate it in admission policy" to "declare it in the Pod spec and let the kernel enforce it."
The short version: what moves from policy to platform
To be precise about the title's claim: "delete from OPA" means the enforcement mechanism for these rules moves out of admission-time policy and into kubelet-and-kernel enforcement. The policy engine does not disappear — defaults are unchanged, so a platform still needs admission rules that require the new fields. What disappears is the machinery that tried, and failed, to approximate kernel guarantees from outside the kernel.
| Rule | Before v1.37 | After v1.37 |
|---|---|---|
No binary execution from writable /tmp | Not expressible: no API field; readOnlyRootFilesystem covers only the image filesystem | bindMountOptions: [noexec, nosuid]; the kernel refuses execution |
| One container cannot delete another's files in shared scratch | Init container running chmod, or deny shared emptyDir by policy | emptyDir with mode: 01777; sticky-bit semantics, kernel-enforced |
| Least-privilege data dir, owner and group only | Init-container chmod 0750; hard to verify for compliance | emptyDir with mode: 0750, declared in the volume source |
| Same protections on PVC-backed volumes | PV mountOptions apply at the CSI layer and never reach the container bind mount | bindMountOptions applies to PV, CSI, and projected mounts too |
The rest of this post substantiates every cell: what shipped, why policy could never express it, the three manifests that replace policy code, what stays in the policy engine, and the five rollout gotchas.
What actually shipped on September 16
The September 16 blog post by Nispriha Jagan and Neeraj Krishna Gopalakrishna (both Red Hat) introduces two alpha features in v1.37, each behind its own gate, enabled on both the API server and the kubelet:
- KEP-5855, gate
VolumeBindMountOptions: abindMountOptionsfield onvolumeMountsacceptingnoexec,nosuid, andnodev. These become VFS flags on the bind mount the container runtime creates inside the container —noexecrefuses direct execution of binaries,nosuidignores setuid/setgid bits,nodevrefuses to interpret device nodes. - KEP-5502, gate
EmptyDirVolumeMode: amodefield onemptyDiraccepting values from0000to01777, replacing the hardcoded0777the volume type has created directories with since forever. It works with all threeemptyDirmediums: disk-backed,Memory(tmpfs), andHugePages.
Coverage is broad: bindMountOptions works with emptyDir, PersistentVolumes, CSI volumes, projected volumes, ConfigMaps, and Secrets. The one explicit exclusion is image volumes. Both features are Linux-only — noexec-style flags and Unix permission modes have no meaning on Windows nodes, where the fields are accepted and skipped. And both are strictly additive: omit the fields and you get exactly yesterday's behavior, 0777 included.
Context from Sysdig's 1.37 security rundown: the Garhwal release (GA August 26, 2026) carries 67 enhancements, 19 with security implications — including SELinuxMount graduating to stable, a sibling storage-hardening change that mounts volumes with -o context= instead of recursively relabeling them. Volume security is having a release.
The "before": why no policy engine could express noexec
This is the part worth sitting with, because it explains why the gap survived every policy tool Kubernetes ever grew. OPA Gatekeeper, Kyverno, and custom admission webhooks all operate on the same input: the Pod object submitted to the API. They can require fields, deny values, and mutate defaults — but only for fields that exist. Pre-v1.37, no field in the Pod spec described bind-mount flags, so there was nothing for a policy to require. The strongest available postures were all approximations:
readOnlyRootFilesystem: trueis genuinely valuable, but it covers the container image filesystem, not writable volumes. The attack the k8s blog names — download payload toemptyDir,chmod +x, execute — sails past it.- Pod Security Standards' restricted profile limits which volume types a pod may use, not how they are mounted. An
emptyDiris allowed; anemptyDirwithoutnoexecis the only kind that existed. - PV
mountOptionslook like the answer until you read the layering: those options are filesystem-level flags applied by the CSI driver at the node, and they do not reliably translate into bind-mount flags inside the container. Different layer, no conflict — and no help. - For
emptyDirpermissions specifically, the workaround was an init container runningchmod— extra complexity the blog calls out as hard to verify for compliance, since the declared spec says0777while the runtime state says whatever the init container did last.
External auditors had been writing this up for years. Issue #48912 flagged the missing mount options after an audit; issue #119627 records the Kubernetes 1.24 security audit finding (NCC-E003660-7HM) that called the inability to mount emptyDir with noexec a security failure outright.
And the admission layer had its own structural weakness. A policy enforced by a webhook is only as available as the webhook: leave failurePolicy: Ignore (the default for custom webhooks) and a downed policy service fails open, admitting unscanned pods silently; set failurePolicy: Fail and the same outage blocks every deployment cluster-wide. Platforms pick their poison per policy. Kernel-enforced mount flags have no such dial — the kubelet applies them or rejects the pod.
The "after": three manifests that delete policy code
Each manifest below replaces a specific piece of platform machinery. All three assume the v1.37 alpha gates enabled.
1. No-exec, no-suid scratch space. The direct fix for the audit finding: an emptyDir at /tmp that the kernel will not execute from, paired with a read-only root filesystem so the writable volume is the only place a payload could land — and it cannot run there.
apiVersion: v1
kind: Pod
metadata:
name: hardened-bindmount-pod
spec:
os:
name: linux
containers:
- name: hardened-app
image: alpine:latest
command: ["sleep", "3600"]
securityContext:
readOnlyRootFilesystem: true
volumeMounts:
- name: temp-storage
mountPath: /tmp
bindMountOptions:
- noexec
- nosuid
volumes:
- name: temp-storage
emptyDir: {}What it deletes: any detection-side compensating control (a Falco rule watching for chmod +x under /tmp, say) gets demoted from enforcement to telemetry, and the "writable volume" exception every tenant's security review used to carry goes away. Verification is one kubectl exec: write a script to /tmp, try to run it, and the kernel answers Permission denied because MS_NOEXEC is enforced at the bind-mount level.
2. Sticky-bit shared scratch for multi-container pods. The CI-shaped case from the k8s blog: a builder container and a sidecar sharing a workspace, where a compromise in one must not let it delete the other's artifacts. Mode 01777 reproduces classic /tmp semantics — anyone may write, only the owner (or root) may delete or rename.
apiVersion: v1
kind: Pod
metadata:
name: hardened-emptydir-pod
spec:
os:
name: linux
containers:
- name: builder
image: alpine:latest
command: ["sleep", "3600"]
volumeMounts:
- name: shared-tmp
mountPath: /tmp
- name: sidecar-logger
image: alpine:latest
command: ["sleep", "3600"]
volumeMounts:
- name: shared-tmp
mountPath: /tmp
volumes:
- name: shared-tmp
emptyDir:
mode: 01777What it deletes: the init container that used to chmod 1777 the shared directory before the app containers started — along with the ordering dependency and the compliance gap between declared spec and runtime state. ls -ld /tmp now shows drwxrwxrwt straight from the declared manifest, and a cross-user rm fails with Operation not permitted.
3. Least-privilege application data. A database pod locking its temporary storage to one user and group, denying even fellow sidecars in the same pod:
apiVersion: v1
kind: Pod
metadata:
name: db-tmp-lockdown
spec:
os:
name: linux
containers:
- name: db
image: postgres:17
command: ["sleep", "3600"]
volumeMounts:
- name: db-tmp
mountPath: /var/run/db-tmp
volumes:
- name: db-tmp
emptyDir:
mode: 0750What it deletes: the per-workload chmod 0750 init container and, more importantly, the policy exception process around it — no more "this team's init container needs to run as root to fix up permissions" carve-outs.
One thing all three manifests share: nothing here is default-on. The platform must still require these fields, which means admission policy of a new, smaller shape — a Kyverno mutate rule that defaults bindMountOptions: [noexec, nosuid] onto tenant volume mounts, or a validate rule rejecting emptyDir without an explicit mode. The policy engine stops approximating kernel behavior and starts mandating declarations. Smaller, dumber, and far easier to audit.
What mount options cannot express (the policy engine stays)
Kernel mount flags are a narrow, deep enforcement point — they say nothing about anything above the mount. A platform deleting policy code needs this keep-list taped next to the delete-list:
- Image provenance.
noexecstops execution from a volume; it says nothing about which images may run. Registry allow-lists, digest pinning, and signature verification (the Sigstore / Ratify-shaped stack) remain admission-shaped problems. - Syscall filtering. A binary that never touches a writable volume is unaffected by mount flags. Seccomp profiles, and policies requiring them, stay.
- Tenant exception management. One tenant's legacy job genuinely needs an executable scratch volume. The grant/deny workflow for that exception — who approves, how long it lasts, where it is recorded — is policy-shaped and always will be.
- The require-the-flag rules themselves. As the manifests above show, defaults are unchanged, so "every tenant mount carries
noexec" is still an admission rule. The difference is that it now enforces something real instead of approximating something unrepresentable.
This is the honest reading of "shrinks but doesn't disappear": the policy repository gets shorter and each remaining rule gets more legible, because every rule now points at a field that exists.
Operator's rollout checklist: five gotchas before you enable the gates
Alpha features with kernel-level effects deserve a careful rollout. All five of these come straight from the k8s blog's "things to know" section:
- Enable both gates on both components.
VolumeBindMountOptionsandEmptyDirVolumeModemust be on for the API server and the kubelet. Half-enabled is the worst state, which leads to gotcha 3. - Check runtime support before scheduling. For
bindMountOptions, the container runtime must support the CRImount_optionsfield and advertise it viaruntimeFeatures. The scheduler uses Node Declared Features (stable in 1.37) to keep flagged pods off incapable nodes; if one arrives anyway, the kubelet rejects it. There is no silent degradation — for this feature. Upgrade runtimes before enabling the gate. - Mind the skew asymmetry.
bindMountOptionsfails loud (rejection), butemptyDirmodefails silent: if the API server accepts the field and the kubelet doesn't know it, the kubelet falls back to0777without complaint. During a rolling upgrade, audit actual directory modes on mixed-version node pools — declared0750with effective0777is exactly the compliance gap this feature was meant to close. fsGroupoverridesmode. If the pod's security context setsfsGroup, its group-permission handling wins over theemptyDirmode, mirroring the long-standing behavior ofdefaultModeon Secret and ConfigMap volumes. Platforms that setfsGroupbroadly should test the interaction before promising tenants a mode.- Linux-only, and image volumes excluded.
bindMountOptionshas no effect on Windows nodes,modeis skipped there too, and image volumes are explicitly unsupported. If your fleet is mixed-OS, the require-the-flag admission rule needs a Windows carve-out, not a global mandate.
Work through those five and the upgrade is boring in the best way: additive fields, unchanged defaults, enforcement that either applies or says so.
Enforcement moves down the stack
Step back and the pattern is bigger than two fields. Kubernetes spent years pushing security opinions up into admission control — admission is flexible, but it is also distant from the thing being protected and coupled to webhook availability. bindMountOptions and emptyDir mode push two opinions back down, to the kubelet and the kernel, where they hold regardless of which policy service is having a bad day. For a self-hosted PaaS operator, that is the actionable frame for upgrade planning: every rule that moves down the stack is a rule that stops paging you when the policy tier degrades.
v1.37 is fresh — Garhwal went GA on August 26, these gates are alpha, and no prudent platform enables alpha gates on tenant fleets on day one. But the direction is set, the audit findings are answered, and the manifests above are the shape of the steady state. The right move this quarter is a lab cluster with both gates on, the require-the-flag rules drafted, and the delete-list for the policy repo started. By the time these graduate to beta, the only remaining work should be deleting code.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



