Skip to main content

E2B vs Modal vs Daytona vs Fly Machines: The Agent Sandbox Market in 2026

9 min readDora NodaDora Noda
Share
On this page

In March 2024, E2B ran 40,000 sandboxes a month. A year later it was running 15 million. That 375x jump is the clearest single number for what happened to AI infrastructure in 2025: agents stopped being demos that answer questions and became programs that write, run, and debug code — and every one of those runs needs somewhere disposable, isolated, and fast to happen in.

That "somewhere" is now a real market with real venture money behind it. E2B raised a $21M Series A led by Insight Partners in July 2025. Daytona raised a $24M Series A led by FirstMark in February 2026. Modal closed a $355M Series C at a $4.65B valuation in May 2026. And Fly.io entered the ring in January 2026 with Sprites, persistent microVMs aimed squarely at agent workloads. Four companies, four different answers to the same question: where should an agent's code execute?

This post compares all four on the dimensions that actually decide a purchase — isolation strength, cold start, statefulness, GPU support, and price — and then draws the line the market keeps blurring: a sandbox is not a deploy platform, and treating one like the other breaks both.

The four-way comparison​

E2BModalDaytonaFly Machines / Sprites
IsolationFirecracker microVMs (own kernel per sandbox)gVisor on KVMContainers (Docker / Kata / Sysbox)Firecracker microVMs
Cold start (vendor claim)~150 msSub-secondSub-90 ms~300 ms warm restore (Sprites); seconds cold
Cold start (independent)Fastest in one 5-way test (~0.5 s end-to-end)~56–65 s full bootstrap in harness tests~0.75 s measured by a third party vs the 90 ms claimVaries by image size
PersistenceEphemeral by default (5-min timeout, up to 24 h on paid plans)Stateful; sandboxes auto-shutdown when the agent finishesStateful snapshots: fork, clone, resume mid-executionMachines are disposable; Sprites persist 100 GB NVMe and auto-idle
GPUNoYes (T4 through H100/B200)YesNo (CPU-only)
Pricing~$0.05/vCPU-hr, per-secondPer-second provisioned CPU/GPUPer-second, similar rates to E2BMachines from ~$0.003/hr (shared CPU); Sprites $0.07/CPU-hr + $0.04375/GB-hr, idle billed as cold storage

Two caveats before anyone shops off this table. First, cold-start numbers are the most gamed metric in this market: vendors measure "time to sandbox object returned" while independent tests measure "time to first command executed" or full dependency bootstrap, and those differ by an order of magnitude. Daytona's sub-90 ms claim and a practitioner's report of 2–5 seconds in production can both be true under different definitions. Second, per-second pricing only means what the meter counts — which brings us to the idle problem below.

What each one actually is​

E2B is the reference agent sandbox. Firecracker microVMs created and destroyed per session, SDKs in Python and JavaScript, custom sandbox templates, and an open-source runtime (Apache-2.0) you can self-host. E2B reports sign-ups from 88% of the Fortune 100, with Perplexity, Hugging Face, Groq, and Vercel named as customers. The honest limits: no GPU inside the sandbox (Firecracker lacks PCIe passthrough), and state is deliberately ephemeral — if your agent needs to remember what it installed last Tuesday, E2B is the wrong tool. One comparison puts the default 2-vCPU sandbox at $0.10/hour, roughly 30x the raw compute price of a minimal Fly Machine.

Modal is a serverless cloud that happens to be the best sandbox for GPU work. Founded in 2021 by ex-Spotify engineer Erik Bernhardsson, Modal's pitch is "just write Python": decorate a function, declare its image and GPU in code, and modal deploy turns it into an autoscaling endpoint. Modal Sandboxes bring that same substrate to agents with auto-shutdown when the run finishes. It is the only option in this comparison with first-class GPU support, from T4s to H100s — required the moment an agent runs inference, image generation, or fine-tuning rather than plain code. The trade-off is generality: you are buying into a full serverless platform, with provisioned-resource pricing to match, not a purpose-built sandbox API.

Daytona is the stateful one. Founded in 2023 as a dev-environments company, Daytona pivoted to agent infrastructure in early 2025 and now sells "programmatic, composable computers": an agent launches a sandbox in milliseconds, forks into parallel branches to explore decision paths, and snapshots mid-execution so state survives failures. That fork/clone/resume model is the genuine differentiator — E2B gives you a fresh box every time, Daytona gives you a box with a memory. Isolation is container-based rather than microVM-based, which is the price of the speed: weaker blast-radius containment than a dedicated guest kernel, partially offset by Kata/Sysbox options.

Fly Machines are the primitive, not the product. Fly gives you Firecracker microVMs you boot and destroy per request over an API, across 35+ regions, at commodity prices — and you build the sandbox abstraction yourself: image management, lifecycle, network allow-listing, eviction. That DIY character is both the appeal (full control, cheapest raw compute, no per-seat anything) and the cost (every sandbox feature is your code to write and operate). Sprites, launched in January 2026, close part of the gap: persistent 100 GB NVMe volumes, ~300 ms checkpoint/restore, scale-to-zero with idle billed only as cold storage at $0.02/GB-month. For async coding agents that tolerate a one-second cold start, Sprites changed the value conversation overnight.

The idle-billing gotcha that decides real costs​

Here is the cost question vendors hope you never ask: what does the meter do while the model is thinking? A real agent loop is mostly waiting — the sandbox sits open, the model reasons, nothing executes. An August 2026 benchmark that priced realistic agent workloads found a ~2.4x spread on identical work, driven largely by idle policy: Vercel's active-CPU metering came in cheapest at $16.27 for the reference workload, E2B and Daytona landed at $27.60, and Modal at $39.66. Same code executed, wildly different bills — because one vendor meters provisioned time and another meters active CPU.

This is the sensitivity analysis that matters more than any per-vCPU sticker price. If your agents are bursty (short executions separated by long reasoning pauses), idle-metering policy dominates your bill. Ask every vendor how they meter an open-but-idle sandbox before you compare hourly rates — and treat any comparison that only quotes $/vCPU-hr as incomplete.

A sandbox is not a deploy platform​

Now the line the market keeps blurring. All four products above answer "where does the agent try code." None of them answers "where does the resulting app live." Those are two different primitives with opposite requirements:

Agent sandboxDeploy platform
LifecycleSeconds to hours, then destroyedMonths to years, then upgraded
Trust postureRuns untrusted, model-generated code; assumes compromiseRuns reviewed, committed code; assumes integrity
DurabilityDisposable by design; persistence is a checkpoint, not a promisePersistent by contract; data loss is an incident
Billing shapePer-second, scale-to-zero, idle is wasteAlways-on or autoscaled with a floor; idle is headroom
Success metricFast boot, strong isolation, cheap coworker-parallelismUptime, deploy velocity, rollback safety

Bolting sandbox execution directly onto the deploy platform conflates workloads with opposite durability and trust requirements, and each direction of the confusion fails concretely:

  • Running untrusted agent code on the deploy substrate means the thing that serves your users shares a trust domain with the thing that executes model hallucinations. A prompt-injected agent that escapes a weak sandbox should land in an empty microVM scheduled for deletion — not on a node that also hosts production traffic.
  • Treating sandbox output as deployed skips everything a deploy platform does: health checks, traffic shifting, rollback, domain and TLS management, log retention. Code that ran once in a sandbox is a draft, not a release; the path from one to the other runs through git, review, and a deploy pipeline — not through keeping the sandbox warm.
  • Merging the bills hides both cost signals. Sandbox spend should scale with agent activity (runs per day); platform spend should scale with serving load (requests per second). One meter for both means you can optimize neither.

The healthy architecture is a handoff, not a merger: the agent iterates in a disposable sandbox, commits working code to git, and the deploy platform picks it up from there. Sandbox vendors know this — it is why Daytona's unit of work is a branchable snapshot and E2B's is a session, not a service. Nobody in this market sells you an SLA on request latency or a custom domain with auto-TLS, because that was never the job.

What self-hosters should borrow​

If you operate your own platform, the sandbox market still has plenty to teach you — as patterns to adopt, not as a second product to staple on:

  • Per-run Firecracker isolation is the ceiling for tenant-executed code. E2B proved the unit economics work at 15M runs a month; Kata Containers or gVisor are the container-native fallback when KVM-per-tenant is too heavy for your fleet.
  • Snapshot/restore beats rebuild for agent iteration. Daytona's fork-and-resume and Sprites' 300 ms checkpoint/restore both say the same thing: make the sandbox's unit of state a snapshot, and agents get branching workflows instead of cold boots.
  • Meter active execution, not wall-clock tenancy. The idle-billing spread above is a pricing lesson for anyone who will ever bill agent workloads: if your meter charges for thinking time, your heaviest users are subsidizing your pricing model instead of your infrastructure.
  • Keep the trust domains physically separate. The cheapest version of this lesson is E2B's open-source runtime: self-hostable Apache-2.0 sandbox infrastructure you can run on your own machines, on a separate network segment from anything that serves traffic, without paying anyone per second.

The through-line: sandbox execution is a genuinely separate infrastructure layer from deployment, with its own lifecycle, trust model, and billing shape. The four vendors above disagree about isolation technology, statefulness, and price — but they agree, by what none of them sells, that running an agent's experiments and running your users' app are two different jobs.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Agent sandboxes are where code gets tried; bex is where it goes to live. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide