Skip to main content

Kubernetes v1.37 Puts No-Exec Volumes in the Pod Spec: What a PaaS Can Delete From OPA

10 min readDora NodaDora Noda
Share
On this page

Until this month, no Kubernetes policy could stop a compromised container from downloading a binary to a writable volume, running chmod +x, and executing it. Not an OPA rule, not a Kyverno policy, not Pod Security Standards. The enforcement point simply did not exist: volumes were bind-mounted into containers without noexec, nosuid, or nodev, and no API field could say otherwise. On September 16, 2026, Kubernetes v1.37 created that enforcement point — bindMountOptions on volume mounts and a mode field on emptyDir — and moved a whole class of volume rules from "approximate it in admission policy" to "declare it in the Pod spec and let the kernel enforce it."

The short version: what moves from policy to platform​

To be precise about the title's claim: "delete from OPA" means the enforcement mechanism for these rules moves out of admission-time policy and into kubelet-and-kernel enforcement. The policy engine does not disappear — defaults are unchanged, so a platform still needs admission rules that require the new fields. What disappears is the machinery that tried, and failed, to approximate kernel guarantees from outside the kernel.

RuleBefore v1.37After v1.37
No binary execution from writable /tmpNot expressible: no API field; readOnlyRootFilesystem covers only the image filesystembindMountOptions: [noexec, nosuid]; the kernel refuses execution
One container cannot delete another's files in shared scratchInit container running chmod, or deny shared emptyDir by policyemptyDir with mode: 01777; sticky-bit semantics, kernel-enforced
Least-privilege data dir, owner and group onlyInit-container chmod 0750; hard to verify for complianceemptyDir with mode: 0750, declared in the volume source
Same protections on PVC-backed volumesPV mountOptions apply at the CSI layer and never reach the container bind mountbindMountOptions applies to PV, CSI, and projected mounts too

The rest of this post substantiates every cell: what shipped, why policy could never express it, the three manifests that replace policy code, what stays in the policy engine, and the five rollout gotchas.

What actually shipped on September 16​

The September 16 blog post by Nispriha Jagan and Neeraj Krishna Gopalakrishna (both Red Hat) introduces two alpha features in v1.37, each behind its own gate, enabled on both the API server and the kubelet:

  • KEP-5855, gate VolumeBindMountOptions: a bindMountOptions field on volumeMounts accepting noexec, nosuid, and nodev. These become VFS flags on the bind mount the container runtime creates inside the container — noexec refuses direct execution of binaries, nosuid ignores setuid/setgid bits, nodev refuses to interpret device nodes.
  • KEP-5502, gate EmptyDirVolumeMode: a mode field on emptyDir accepting values from 0000 to 01777, replacing the hardcoded 0777 the volume type has created directories with since forever. It works with all three emptyDir mediums: disk-backed, Memory (tmpfs), and HugePages.

Coverage is broad: bindMountOptions works with emptyDir, PersistentVolumes, CSI volumes, projected volumes, ConfigMaps, and Secrets. The one explicit exclusion is image volumes. Both features are Linux-only — noexec-style flags and Unix permission modes have no meaning on Windows nodes, where the fields are accepted and skipped. And both are strictly additive: omit the fields and you get exactly yesterday's behavior, 0777 included.

Context from Sysdig's 1.37 security rundown: the Garhwal release (GA August 26, 2026) carries 67 enhancements, 19 with security implications — including SELinuxMount graduating to stable, a sibling storage-hardening change that mounts volumes with -o context= instead of recursively relabeling them. Volume security is having a release.

The "before": why no policy engine could express noexec​

This is the part worth sitting with, because it explains why the gap survived every policy tool Kubernetes ever grew. OPA Gatekeeper, Kyverno, and custom admission webhooks all operate on the same input: the Pod object submitted to the API. They can require fields, deny values, and mutate defaults — but only for fields that exist. Pre-v1.37, no field in the Pod spec described bind-mount flags, so there was nothing for a policy to require. The strongest available postures were all approximations:

  • readOnlyRootFilesystem: true is genuinely valuable, but it covers the container image filesystem, not writable volumes. The attack the k8s blog names — download payload to emptyDir, chmod +x, execute — sails past it.
  • Pod Security Standards' restricted profile limits which volume types a pod may use, not how they are mounted. An emptyDir is allowed; an emptyDir without noexec is the only kind that existed.
  • PV mountOptions look like the answer until you read the layering: those options are filesystem-level flags applied by the CSI driver at the node, and they do not reliably translate into bind-mount flags inside the container. Different layer, no conflict — and no help.
  • For emptyDir permissions specifically, the workaround was an init container running chmod — extra complexity the blog calls out as hard to verify for compliance, since the declared spec says 0777 while the runtime state says whatever the init container did last.

External auditors had been writing this up for years. Issue #48912 flagged the missing mount options after an audit; issue #119627 records the Kubernetes 1.24 security audit finding (NCC-E003660-7HM) that called the inability to mount emptyDir with noexec a security failure outright.

And the admission layer had its own structural weakness. A policy enforced by a webhook is only as available as the webhook: leave failurePolicy: Ignore (the default for custom webhooks) and a downed policy service fails open, admitting unscanned pods silently; set failurePolicy: Fail and the same outage blocks every deployment cluster-wide. Platforms pick their poison per policy. Kernel-enforced mount flags have no such dial — the kubelet applies them or rejects the pod.

The "after": three manifests that delete policy code​

Each manifest below replaces a specific piece of platform machinery. All three assume the v1.37 alpha gates enabled.

1. No-exec, no-suid scratch space. The direct fix for the audit finding: an emptyDir at /tmp that the kernel will not execute from, paired with a read-only root filesystem so the writable volume is the only place a payload could land — and it cannot run there.

yaml
apiVersion: v1
kind: Pod
metadata:
  name: hardened-bindmount-pod
spec:
  os:
    name: linux
  containers:
    - name: hardened-app
      image: alpine:latest
      command: ["sleep", "3600"]
      securityContext:
        readOnlyRootFilesystem: true
      volumeMounts:
        - name: temp-storage
          mountPath: /tmp
          bindMountOptions:
            - noexec
            - nosuid
  volumes:
    - name: temp-storage
      emptyDir: {}

What it deletes: any detection-side compensating control (a Falco rule watching for chmod +x under /tmp, say) gets demoted from enforcement to telemetry, and the "writable volume" exception every tenant's security review used to carry goes away. Verification is one kubectl exec: write a script to /tmp, try to run it, and the kernel answers Permission denied because MS_NOEXEC is enforced at the bind-mount level.

2. Sticky-bit shared scratch for multi-container pods. The CI-shaped case from the k8s blog: a builder container and a sidecar sharing a workspace, where a compromise in one must not let it delete the other's artifacts. Mode 01777 reproduces classic /tmp semantics — anyone may write, only the owner (or root) may delete or rename.

yaml
apiVersion: v1
kind: Pod
metadata:
  name: hardened-emptydir-pod
spec:
  os:
    name: linux
  containers:
    - name: builder
      image: alpine:latest
      command: ["sleep", "3600"]
      volumeMounts:
        - name: shared-tmp
          mountPath: /tmp
    - name: sidecar-logger
      image: alpine:latest
      command: ["sleep", "3600"]
      volumeMounts:
        - name: shared-tmp
          mountPath: /tmp
  volumes:
    - name: shared-tmp
      emptyDir:
        mode: 01777

What it deletes: the init container that used to chmod 1777 the shared directory before the app containers started — along with the ordering dependency and the compliance gap between declared spec and runtime state. ls -ld /tmp now shows drwxrwxrwt straight from the declared manifest, and a cross-user rm fails with Operation not permitted.

3. Least-privilege application data. A database pod locking its temporary storage to one user and group, denying even fellow sidecars in the same pod:

yaml
apiVersion: v1
kind: Pod
metadata:
  name: db-tmp-lockdown
spec:
  os:
    name: linux
  containers:
    - name: db
      image: postgres:17
      command: ["sleep", "3600"]
      volumeMounts:
        - name: db-tmp
          mountPath: /var/run/db-tmp
  volumes:
    - name: db-tmp
      emptyDir:
        mode: 0750

What it deletes: the per-workload chmod 0750 init container and, more importantly, the policy exception process around it — no more "this team's init container needs to run as root to fix up permissions" carve-outs.

One thing all three manifests share: nothing here is default-on. The platform must still require these fields, which means admission policy of a new, smaller shape — a Kyverno mutate rule that defaults bindMountOptions: [noexec, nosuid] onto tenant volume mounts, or a validate rule rejecting emptyDir without an explicit mode. The policy engine stops approximating kernel behavior and starts mandating declarations. Smaller, dumber, and far easier to audit.

What mount options cannot express (the policy engine stays)​

Kernel mount flags are a narrow, deep enforcement point — they say nothing about anything above the mount. A platform deleting policy code needs this keep-list taped next to the delete-list:

  • Image provenance. noexec stops execution from a volume; it says nothing about which images may run. Registry allow-lists, digest pinning, and signature verification (the Sigstore / Ratify-shaped stack) remain admission-shaped problems.
  • Syscall filtering. A binary that never touches a writable volume is unaffected by mount flags. Seccomp profiles, and policies requiring them, stay.
  • Tenant exception management. One tenant's legacy job genuinely needs an executable scratch volume. The grant/deny workflow for that exception — who approves, how long it lasts, where it is recorded — is policy-shaped and always will be.
  • The require-the-flag rules themselves. As the manifests above show, defaults are unchanged, so "every tenant mount carries noexec" is still an admission rule. The difference is that it now enforces something real instead of approximating something unrepresentable.

This is the honest reading of "shrinks but doesn't disappear": the policy repository gets shorter and each remaining rule gets more legible, because every rule now points at a field that exists.

Operator's rollout checklist: five gotchas before you enable the gates​

Alpha features with kernel-level effects deserve a careful rollout. All five of these come straight from the k8s blog's "things to know" section:

  1. Enable both gates on both components. VolumeBindMountOptions and EmptyDirVolumeMode must be on for the API server and the kubelet. Half-enabled is the worst state, which leads to gotcha 3.
  2. Check runtime support before scheduling. For bindMountOptions, the container runtime must support the CRI mount_options field and advertise it via runtimeFeatures. The scheduler uses Node Declared Features (stable in 1.37) to keep flagged pods off incapable nodes; if one arrives anyway, the kubelet rejects it. There is no silent degradation — for this feature. Upgrade runtimes before enabling the gate.
  3. Mind the skew asymmetry. bindMountOptions fails loud (rejection), but emptyDir mode fails silent: if the API server accepts the field and the kubelet doesn't know it, the kubelet falls back to 0777 without complaint. During a rolling upgrade, audit actual directory modes on mixed-version node pools — declared 0750 with effective 0777 is exactly the compliance gap this feature was meant to close.
  4. fsGroup overrides mode. If the pod's security context sets fsGroup, its group-permission handling wins over the emptyDir mode, mirroring the long-standing behavior of defaultMode on Secret and ConfigMap volumes. Platforms that set fsGroup broadly should test the interaction before promising tenants a mode.
  5. Linux-only, and image volumes excluded. bindMountOptions has no effect on Windows nodes, mode is skipped there too, and image volumes are explicitly unsupported. If your fleet is mixed-OS, the require-the-flag admission rule needs a Windows carve-out, not a global mandate.

Work through those five and the upgrade is boring in the best way: additive fields, unchanged defaults, enforcement that either applies or says so.

Enforcement moves down the stack​

Step back and the pattern is bigger than two fields. Kubernetes spent years pushing security opinions up into admission control — admission is flexible, but it is also distant from the thing being protected and coupled to webhook availability. bindMountOptions and emptyDir mode push two opinions back down, to the kubelet and the kernel, where they hold regardless of which policy service is having a bad day. For a self-hosted PaaS operator, that is the actionable frame for upgrade planning: every rule that moves down the stack is a rule that stops paging you when the policy tier degrades.

v1.37 is fresh — Garhwal went GA on August 26, these gates are alpha, and no prudent platform enables alpha gates on tenant fleets on day one. But the direction is set, the audit findings are answered, and the manifests above are the shape of the steady state. The right move this quarter is a lab cluster with both gates on, the require-the-flag rules drafted, and the delete-list for the policy repo started. By the time these graduate to beta, the only remaining work should be deleting code.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex