AI agents went from chat demos to millions of production sandboxes in under five months — and the runtime underneath them just became a Kubernetes primitive you can install on any cluster. GKE Agent Sandbox went generally available on May 20, 2026, after growing 16x in sandboxes on GKE in less than five months since its KubeCon NA preview in November 2025. Langchain and Lovable are deploying millions of agents on it. And the whole thing is upstream: kubernetes-sigs/agent-sandbox, a SIG Apps subproject, ships a declarative Sandbox CRD with gVisor isolation that runs anywhere Kubernetes runs — including a Cluster-API-managed fleet on machines you own.
Here is the verdict before the why: if you run a self-hosted PaaS, adopt the upstream Sandbox API instead of building your own agent-runtime layer. Warm pools, pod snapshots, suspend/resume lifecycle, and pluggable gVisor/Kata isolation are now kubectl apply away, with Go and Python SDKs included. What you still own is the part only you can own: node RuntimeClasses, pool sizing, egress policy, and GPU scheduling. This post shows the evidence: what the primitive is, the production numbers behind it, what it replaces on a vendor invoice, and where the roadmap still has gaps.
What the primitive actually is: four CRDs, one RuntimeClass
Agent Sandbox exists because an AI agent runtime is a workload shape Kubernetes' built-in types describe badly. An agent needs a singleton (one live instance, not a scaled replica set), stateful (a persistent terminal, filesystem, and identity across tool calls), isolated execution environment for running untrusted, LLM-generated code. A Deployment gives you replicas you don't want; a bare Pod gives you no stable identity, no lifecycle, and shared-kernel execution. The Sandbox API closes exactly that gap.
The API surface is small — one core CRD plus three extensions:
| CRD | API group | What it does |
|---|---|---|
Sandbox | agents.x-k8s.io/v1beta1 | One stateful, singleton, pod-backed workload: stable name and hostname, optional persistent storage, pause/resume, scheduled deletion |
SandboxTemplate | extensions.agents.x-k8s.io/v1beta1 | Reusable blueprint: pod spec plus network policy and injection policies |
SandboxWarmPool | extensions.agents.x-k8s.io/v1beta1 | Keeps N pre-booted sandboxes from a template ready for instant claim |
SandboxClaim | extensions.agents.x-k8s.io/v1beta1 | Checks one sandbox out of a pool without touching template details |
Two design decisions matter. First, isolation is delegated, not implemented: the controller is a sandbox orchestrator, and the actual boundary comes from the RuntimeClass on the pod spec — gVisor by default, Kata Containers when you want a full guest kernel. Second, allocation is separated from definition: templates define the environment once, warm pools pre-boot it, and claims hand out ready instances. The consumer-facing flow looks like this:
apiVersion: extensions.agents.x-k8s.io/v1beta1
kind: SandboxClaim
metadata:
name: agent-session-7f3a
spec:
templateRef:
name: python-agent-runtime
ttlSecondsAfterClaimed: 3600Claim it, get a bound, already-booted sandbox with a stable identity — no image pull, no cold boot, no hand-rolled pool manager. The core API graduated to v1beta1 with the v0.5.0 release in June 2026, which is the stability signal a platform team wants before building on someone else's CRD.
The production evidence, in numbers
Upstream CRDs are cheap; production evidence is not. Agent Sandbox has both, because GKE ran it as a managed add-on first and published the numbers:
| Signal | Number | Source |
|---|---|---|
| Growth since KubeCon NA Nov 2025 preview | 16x sandboxes on GKE in under 5 months | Google Cloud blog, May 2026 |
| Production agents | Millions, incl. Langchain and Lovable | Google Cloud blog, May 2026 |
| Warm-pool allocation rate | 300 sandboxes/sec per cluster, sub-second; p90 at 200 ms | GKE Agent Sandbox GA notes |
| Warm pool vs cold boot | 10–15x faster allocation | Upstream quickstart |
| Suspend/resume | Idle agents suspended via pod snapshots, resumed in seconds | GKE Agent Sandbox docs |
| Price-performance | Up to 30% better on Axion vs comparable hyperscaler instances | Google Cloud blog, May 2026 |
Two details deserve emphasis. First, the warm-pool economics: keeping pre-booted replicas ready sounds expensive until you see the second mechanism — pools are replenished from standby capacity buffers (suspended VMs), a cold pool of suspended sandboxes that refills the warm pool for a fraction of the cost of running everything hot. Second, the scale that forced the performance work: one community report puts Lovable alone at roughly 200,000 sandboxes a day. That is the load profile — bursty and mostly idle between tool calls — that cold-start-per-request architectures cannot serve and that pod snapshots plus warm pools were built for.
Isolation you can audit: gVisor by default, Kata when you need a kernel
The threat model is stated plainly in the project docs: isolated environments for executing untrusted, LLM-generated code. The default answer is gVisor — a userspace kernel that intercepts syscalls — paired with a default-deny Kubernetes network policy, so a compromised agent gets neither host syscalls nor open egress. For tenants that need stronger separation, the pluggable interface swaps in Kata Containers: each sandbox becomes a lightweight VM with its own guest kernel.
This layering is the part a platform team should appreciate most. The controller never hard-codes an isolation technology; it enforces whatever RuntimeClass the template names. The repo already ships a Firecracker sandbox example running microVMs through Kata (kata-fc) behind a SandboxWarmPool — two warm microVMs ready for instant claim. So the upgrade path from "syscall filter is enough for this tenant" to "this tenant gets a dedicated microVM kernel" is a template change, not a re-architecture.
Compare that to the status quo it replaces: agent code in plain pods sharing the host kernel, with isolation bolted on via ad-hoc seccomp profiles and prayer. The upstream primitive moves the boundary from tribal knowledge to a declared, auditable field in a CRD.
What a self-hosted PaaS gets for free — and what vendors still charge for
Here is the build-vs-adopt ledger. The left column is what kubectl apply of the upstream manifests plus a RuntimeClass gives you; the right column is what the managed sandbox vendors sell:
| Capability | Upstream agent-sandbox (free) | E2B / Daytona / Modal (metered) |
|---|---|---|
| Isolated execution | gVisor default, Kata pluggable | Same class of isolation (Modal is gVisor-based too) |
| Fast allocation | Warm pools, 10–15x vs cold | ~150 ms creates (E2B), sub-90 ms from snapshot (Daytona), sub-second (Modal) |
| Idle efficiency | Pod snapshots + standby buffers | E2B/Daytona bill full alive duration; Modal and Fly Sprites don't bill idle |
| Lifecycle API | Sandbox/Claim/TTL/scheduled deletion + Go/Python SDKs | Proprietary SDK per vendor |
| List price anchor | Your node cost | ~$0.0504/vCPU-hr + $0.0162/GiB-hr (E2B/Daytona) |
The price anchor matters because agent workloads are bursty and idle-heavy — exactly the profile where per-second metering compounds. One community cost model puts 1,000 sustained concurrent runs at roughly $73k/month at E2B/Daytona rates versus $5–15k/month on self-hosted bare metal. Treat any single cost model skeptically, but the order of magnitude is the point: at fleet scale, the sandbox line item dwarfs the node cost, and the upstream controller deletes the line item while keeping the capability.
Self-hosting the vendors instead is a narrower door than it looks. E2B's self-hosted stack wants Firecracker plus Nomad on GCP (GA) or AWS (beta) with nested virtualization — not your Hetzner fleet. Daytona's open-source edition was frozen in June 2026, leaving hosted-only. The upstream project, by contrast, installs from release manifests onto any cluster with a RuntimeClass-capable runtime — including the Pulumi-documented pattern of a GKE Standard cluster with a dedicated gVisor node pool, which maps directly onto a Cluster-API-managed pool of owned machines.
Be honest about what you still own, though. Upstream gives you orchestration, not operations: node images with gVisor or Kata installed, RuntimeClass definitions, warm-pool sizing against your actual claim rate, egress enforcement (FQDN allowlists via your CNI of choice — the EKS reference solution pairs the controller with Cilium chaining), and GPU scheduling once sandboxes want accelerators. That is ordinary fleet work, not a second product to build.
Roadmap and honest limits
The upstream roadmap shows where the 2026–2027 work is going, and it is all in the direction of less code for you to write:
- Agent and RL framework plugins (in progress): native runtime plugins for LangChain, CrewAI, kAgent, OpenEnv, and Ray RLlib — "plug-and-play" execution environments instead of per-framework sandbox glue.
- More isolation backends: QEMU and Firecracker API support beyond today's gVisor/Kata RuntimeClass path, for hardware-level microVM isolation through the same Sandbox API.
- Agent Substrate (new OSS project): for the scale tier above what a standard control plane handles — Google frames it as millions of sub-second tool calls that would overwhelm etcd-backed controllers. Substrate keeps the Sandbox runtime and snapshotting but adds a minimal control plane that moves agents onto ready compute in real time, with data-locality-aware scheduling. Standard Kubernetes is optimized for thousands of long-running services; Substrate is the answer for the chatter tier.
And the limits, stated plainly. First, the API targets stateful singleton runtimes — long-lived agent sessions with identity and storage. If your workload is one-shot fire-and-forget execution (a trigger fires, code runs once, nothing persists), teams have reasonably concluded the Sandbox shape is heavier than they need and built a thinner executor instead. Match the primitive to the workload. Second, control-plane ceilings still apply: warm pools at 300 claims a second are a GKE-tuned number, and your small fleet's etcd will have its own opinion about claim churn — size pools from measured claim rates, not marketing numbers. Third, framework plugins are in progress, not shipped; today you wire LangChain or CrewAI to the Sandbox API yourself via the SDKs.
None of those limits is an argument for building your own sandbox orchestrator. They are scoping notes for adopting the shared one.
What this means if you run the fleet
A year ago, "we need to run untrusted agent code" meant a vendor contract or a homegrown isolation project. Now it means installing a SIG Apps controller, defining a template and a warm pool, and pointing your agent framework at a SandboxClaim. 16x growth and millions of production agents are the de-risking; the v1beta1 API graduation is the stability commitment; the roadmap is all leverage in your direction.
For a Cluster-API-based platform, the fit is almost embarrassingly direct: the same declarative reconciliation you use for machines now covers agent runtimes — a pool size in a manifest, controllers converging reality toward it, no pets. The sandbox stops being a product decision and becomes a fleet configuration.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. AI agents are first-class operators there, and upstream primitives like the Sandbox API are exactly what a self-hosted agent platform should be built on. Star the repo on GitHub or deploy your first app today.



