The most interesting thing about E2B's sandbox limits is the one that disappeared. Third-party guides still warn that E2B auto-deletes paused sandboxes after 30 days — but E2B's current docs promise the opposite: a paused sandbox "is kept indefinitely," with "no time-to-live and no automatic deletion."
That quiet reversal matters because E2B is the category's reference point. The company raised a $21M Series A led by Insight Partners in July 2025, reporting that 88% of Fortune 100 companies had signed up for its Firecracker-microVM sandboxes. When the market leader deletes a limit, it is conceding that the limit cost more legitimacy than it saved in infrastructure. The limits it kept deserve the same scrutiny.
The 30-day deletion is gone. The sharper limit stayed.
Here is the current shape of E2B's sandbox lifecycle, straight from its docs and pricing page. Continuous runtime is capped by plan: 1 hour on Hobby, 24 hours on Pro, custom on Enterprise. Every sandbox also has a timeout that defaults to 5 minutes if you never set one. And what happens when the timeout expires is your choice — with a destructive default: onTimeout: 'kill' terminates the sandbox and its state is gone, while onTimeout: 'pause' preserves filesystem and memory.
Pausing is the escape hatch from the runtime cap. A pause snapshots filesystem and memory in seconds and resumes in about a second, billing stops while paused, and pausing plus resuming resets the continuous-runtime clock — so a sandbox that pauses and resumes can live as long as you keep choreographing it. E2B even offers autoResume, which wakes a paused sandbox on the next SDK call or HTTP request. The platform gives you every tool to build a long-lived workflow on top of short-lived primitives, as long as you build the choreography yourself.
That is the honest current picture: retention is now infinite, but continuous runtime still has a ceiling, the default timeout is five minutes, and the default expiry action destroys your state. The rest of this post is a concrete read of what those two surviving choices protect against, what they cost a legitimate workflow, and what a self-hosted sandbox primitive — running on capacity you own rather than a metered fleet — should do differently.
Limit 1: the runtime cap — what one hour and 24 hours are actually pricing
Start with the money, because the runtime caps are a pricing decision wearing an engineering costume. E2B bills running sandboxes per second: $0.000014 per vCPU-second ($0.0504/vCPU/hr) and $0.0000045 per GiB-second ($0.0162/GiB/hr), with storage included free. The default sandbox carries 2 vCPUs and 4 GiB of RAM, which works out to about $0.166 per running hour — roughly $4 a day, or around $120 a month for one sandbox that never pauses, before the $150/month Pro plan fee.
Now anchor a typical workflow to those numbers: picture a multi-hour coding-agent session that clones a repo, spends ten minutes installing dependencies, then alternates bursts of test runs with idle gaps while the model reasons or a human reviews a diff. Left running naively for a 48-hour task with overnight idle stretches, that single session burns roughly $8 in compute doing nothing most of the time — and on Pro it cannot even run 48 hours straight, because the 24-hour ceiling forces at least one pause/resume cycle mid-task. E2B's caps exist so one customer's forgotten while True loop cannot convert the provider's multi-tenant fleet into an unbounded compute liability. That is a completely rational thing for a metered provider to defend against.
But notice where the cap binds and where it does not. For short, checkpointed code-interpreter calls — run a snippet, return the result, kill the sandbox — the caps are harmless; the workload finishes in seconds and never notices the ceiling. The caps turn fatal one step up the ladder: long-lived stateful sessions where the working tree, installed dependencies, and in-memory process state are the product. There, a missed setTimeout on the 5-minute default does not pause your work — it kills the sandbox and the state is gone, punishing exactly the users with the most to lose: newcomers who have not yet learned the lifecycle choreography.
And the choreography is real work. Production E2B integrations implement pause-on-idle, resume-by-ID, stream reconnection (SDK streams send keepalive pings but no single connection is guaranteed past an hour), and reattach-to-running-process logic as first-class machinery. None of that is your agent's job; all of it is forced by the shape of the primitive. E2B documents every piece of it well — which is admirable, and also an admission that the primitive's default posture is "prove you still need this compute, continuously, or lose it."
Limit 2: retention — the fight E2B already conceded
The retired 30-day deletion is worth autopsying, because why E2B could afford to drop it tells you what the limit was ever protecting. Paused snapshots cost $0 in compute — billing stops the moment a sandbox pauses — and sandbox storage is bundled free with every plan. Keeping a paused sandbox costs E2B some object storage and nothing else. The 30-day sweep was never load-bearing for fleet capacity; it was hygiene against unbounded free storage accumulation, and E2B evidently decided the trust cost exceeded the storage savings.
The trust cost was real. E2B has no separate volume product — the sandbox is the storage — so a retention sweep does not archive your data somewhere retrievable; it destroys the only copy. One team documented exactly this failure shape: an ordinary config change made resume refuse their paused sandboxes, the recovery path killed the sandboxes holding user files, and workspaces were wiped by a deploy. Their summary deserves quoting verbatim: "On E2B the sandbox is the storage."
The sharper evidence is e2b-dev/E2B#1381, "Sandbox gets deleted with no warning": a running sandbox disappearing well before its configured timeout, with subsequent calls failing on "The sandbox was not found." That report is about a running sandbox, not a paused one past TTL — which makes it worse, not better, for the category. It shows the failure mode users actually fear is not "my 29-day-old snapshot aged out" but "my live work vanished without a warning, callback, or status change I could have reacted to." E2B fixed the retention policy. The deeper design question — whether compute and storage should share a single fate — is still open, and it is the first thing a self-hosted alternative should answer differently.
The category table: same two axes, five different answers
E2B's combination — capped continuous runtime, infinite paused retention, kill-by-default timeout — is a business choice, not physics. Every competitor draws the same two lines (how long can it run, how long does its state survive) in a different place:
| Provider | Continuous runtime cap | Idle / retention policy |
|---|---|---|
| E2B | 1h Hobby / 24h Pro, resets on pause+resume | Paused kept indefinitely, $0 compute |
| Daytona | None; auto-stops after 15 min idle by default | Stopped workspaces persist; storage-metered |
| Modal | 5 min default, max 24h | Memory snapshots expire in 7 days; filesystem snapshots 30-day default TTL |
| Vercel Sandbox | 45 min Hobby / 24h Pro per session | Filesystem-only snapshots, 30-day default expiry, $0.08/GB-mo |
| Cloudflare Sandbox | Sleep after ~10 min idle | Files deleted on sleep — no resume-with-state |
| Fly.io Sprites | None; sleeps when idle | State kept while warm |
Nobody agrees on a runtime ceiling — Daytona and Fly Sprites simply do not have one, trusting idle-sleep instead. Nobody agrees on retention either: Modal and Vercel keep E2B's old answer (30-day snapshot TTLs), while Cloudflare takes the hardest line in the industry (sleep means your files are gone). The axes are genuinely independent, with each vendor's position reflecting its cost structure — so a self-hosted primitive gets to pick its own point, and owned capacity moves the economics enough that the right point is nowhere near any row in this table.
Four things a self-hosted sandbox should do differently
On a metered multi-tenant fleet, every hour of runtime and every gigabyte of retained state is someone else's margin. On owned Cluster-API capacity — machines whose cost you already paid whether they compute or idle — the marginal cost of hour 25 of a paused-then-resumed session, or of keeping a snapshot for day 31, is approximately zero. That single economic fact rewrites all four design decisions:
1. Separate volumes from compute, so storage outlives sandboxes. The "sandbox is the storage" coupling is the root cause of every silent-data-loss story above. A self-hosted primitive should give each workspace a persistent volume as the durable unit and treat the microVM as ephemeral compute attached to it — Blaxel already ships this shape with volumes that persist after sandbox deletion. Pausing, killing, rescheduling, or upgrading the fleet then never threatens user data, because no lifecycle transition on the compute object can destroy the storage object. This is trivially cheap on owned disks and the highest-leverage divergence from E2B's model.
2. Make the defaults non-destructive: pause, never kill, on timeout. E2B's kill-by-default exists to protect a shared fleet from abandoned compute. A fleet you own can afford the opposite default: when a timeout expires, pause — preserving filesystem, memory, and processes — and only release resources after a separate, explicit, warned retention window. Newcomers who never learn the lifecycle API then get safety instead of data loss, and the only cost is disk for paused snapshots you were already paying for.
3. Make retention an explicit operator policy with advance warning, not a silent TTL. The lesson of the 30-day era and issue #1381 is that silent deletion is the unforgivable part, more than any particular duration. A self-hosted control plane should let the operator set retention per workspace class (scratch sandboxes: days; user projects: months or never), emit warnings well before any sweep, and keep a recoverable grace copy after it. Deletion the user saw coming three times is operations; deletion discovered via "sandbox was not found" is a betrayal.
4. Shape quotas by capacity planning, not by tier upsell. The 1-hour/24-hour cliff exists to segment Hobby from Pro from Enterprise — it is a pricing ladder, not a technical necessity, as Daytona's and Sprites' unlimited runtimes prove. On owned capacity the honest replacement is fair-share quotas derived from actual fleet headroom: per-tenant concurrency limits, idle detection with warned reaping, and burst allowances tied to real utilization. The question changes from "which plan unlocks hour 25" to "does the fleet have headroom for this workload right now" — answerable by a scheduler, not a checkout page.
None of these requires inventing new isolation technology. Firecracker microVMs, snapshot/restore, and per-second accounting all work fine; what changes is the policy layer above them, once the policy no longer needs to defend someone else's margin.
Copy the isolation, not the incentives
E2B's limits were never arbitrary — each one defends a real cost on a metered multi-tenant fleet — and the company's quiet move to infinite paused retention shows it revisits them when the trust math changes. But limits shaped by a provider's cost structure make poor defaults for infrastructure you operate yourself. A sandbox primitive on owned capacity should keep E2B's Firecracker-grade isolation and per-second honesty while flipping the four policies above: durable volumes beneath ephemeral compute, pause-by-default timeouts, warned explicit retention, and quotas computed from headroom rather than plan tiers.
Building agent sandboxes on infrastructure you own is on bex's roadmap: Firecracker-grade isolation on Cluster-API capacity, where the lifecycle policy answers to your workloads instead of a metering tier. Star the repo on GitHub or deploy your first app today.



