Skip to main content

One Sandbox API for E2B, Modal, Vercel, and Kubernetes: What Capability Negotiation Actually Buys You

10 min readDora NodaDora Noda
Share
On this page

Daytona took its core closed-source in June 2026 — and every team that had built directly on its SDK learned the same lesson at the same time: a sandbox vendor is a single point of failure you chose on purpose. Open Sandbox Router, a provider-agnostic control plane for code-execution sandboxes, proposes the fix: one Sandbox handle, one TypeScript SDK, and a router that places each sandbox on E2B, Modal, Vercel Sandbox, Daytona, or your own Kubernetes cluster based on declared capabilities, cost, and live health.

The verdict up front: the design's core insight is correct — sandboxes are stateful, so routing is a create-time decision with a durable sandbox-to-provider binding — and create-time failover plus capability filtering are genuine wins. But mid-session migration across heterogeneous isolation tech is honestly best-effort, and no router can paper over the fact that Firecracker and a shared-kernel container are different security postures. Here is the concrete evaluation: the fragmentation in numbers, the mechanism, the cost table, the two failover regimes, and what leaks.

The fragmentation it attacks, in numbers​

The pain is easy to quantify. Every provider ships its own SDK, resource model, template format, and lifecycle semantics: execution lives on Sandbox for Vercel and Modal, under Sandbox.commands for E2B, and under Sandbox.process for Daytona. Switching providers is a rewrite, not a config change.

Even the OpenAI Agents SDK ships separate sandbox backend extensions for E2B, Modal, Daytona, Cloudflare, Runloop, Blaxel, and Vercel — fragmentation so thorough that the ecosystem's center of gravity just carries all of it.

Billing is three different philosophies wearing the same "per-second" costume: Vercel meters active CPU (idle time waiting on the network or the LLM is free), E2B meters session time, Modal meters per-second provisioned. And cold starts span an order of magnitude. The August 2026 single-harness benchmark from MarkTechPost — time-to-interactive from create() to first successful command, 100 iterations per provider in one concurrent burst — measured:

Providerp50 TTIp95 TTISuccess
Daytona0.27s0.43s37%
Vercel Sandbox0.67s1.04s100%
Modal0.88s1.00s100%
E2B1.61s1.77s100%

That table is the whole argument for measuring once through one SDK and routing forever: the fastest provider in the run also failed 63% of creations under burst load. A router with live health signals would have steered around that bad day automatically; a hardcoded SDK choice eats it. Note the self-hosted row the spec adds for contrast — 10–150ms on your own Kubernetes, where the only queue is yours.

Then there is the exhibit A for portability: Daytona's core development moved to a private codebase in June 2026, freezing the public AGPL repo at v0.190.0 with no further updates. The hosted product continues (managed control plane plus bring-your-own-compute), and the community answered with the Nightona fork.

But anyone who treated "open source, self-hostable" as a permanent property of their Daytona dependency got a forced migration. We covered Daytona's customer-managed compute story before the freeze; the freeze is what happens when your abstraction layer is somebody else's repo banner.

Capability negotiation, concretely​

The mechanism that makes one API span five backends is capability negotiation, and the spec treats it as first-class rather than a filter bolted on later. A create request declares required capabilities — { gpu, snapshot, ports, minMemory, region } — and the router runs a pipeline: filter the candidate set to providers that can comply, score the survivors on cost/latency/region/reliability, order them by policy, and attempt in order with failover. The principle is typed capabilities over silent degradation: if nothing can meet the request, you get a typed error, not a sandbox that quietly lacks the GPU you asked for.

Two pieces of supporting machinery make this more than a load balancer. First, a portable template spec: one Dockerfile-based osr.template.yaml (OCI base, setup steps, env, workdir, exposed ports) that osr template build compiles into each provider's native format — E2B template, Modal image, OCI image for self-hosted — with refs tracked in a content-hash-keyed registry. Content hashing means "same declared environment" is verifiable, not aspirational.

Templates are the second-biggest portability pain after the API itself, which is why the spec tackles them before any higher-level orchestration.

Second, the durable binding: sandbox → { provider, region, capabilities, creds ref } persisted in Postgres (gateway mode) or an embedded store (library mode). Every non-create operation resolves the binding and dispatches to the home provider — there is deliberately no re-routing a live sandbox. A reaper enforces TTLs and reconciles orphaned provider resources, and idempotency keys on create prevent duplicate provisioning on client retries. The overhead budget is explicit: under 50ms of added control-plane latency, with a direct-connect mode planned so the gateway leaves the exec hot path entirely.

The cost table no single vendor will show you​

Here is the normalized comparison the spec's economics section is built to produce continuously, assembled from 2026 provider pricing:

BackendCPUMemoryBilling basis
E2B$0.0504 / vCPU-hr$0.0162 / GiB-hrSession time; $150/mo Pro tier; BYOC available
Modal$0.1419 / core-hr$0.0242 / GiB-hrPer-second provisioned; strongest GPU story (T4–H100)
Vercel Sandbox$0.128 / vCPU-hr$0.0212 / GB-hrActive CPU only + provisioned memory; GA January 2026
Daytona$0.0504 / vCPU-hr$0.0162 / GiB-hrUsage-based; hosted product after the June 2026 OSS freeze
Self-hosted KubernetesYour hardwareYour hardwareFlat infra cost + your ops time; 10–150ms cold starts

But "cheapest" depends entirely on workload shape, which is exactly why a static table goes stale and a routing layer earns its keep:

  • Bursty agent loops (short exec bursts separated by LLM think time) favor Vercel's active-CPU model — the think time is unbilled — while session-time and provisioned models charge through the idle gaps.
  • Long-lived coding-agent workspaces (repo checked out, dev server on an exposed port, multi-hour sessions) favor flat-rate or self-hosted capacity, where the marginal hour is near zero.
  • GPU inference or training bursts collapse the candidate set to Modal (or your own GPU nodes) regardless of CPU pricing — capability filtering does its job before cost scoring even runs.

The spec is admirably honest that cost normalization is approximate: active-CPU vs session vs provisioned cannot be reduced to one number without assumptions. The router's dollar figures are estimates for routing and reporting, reconciled against provider invoices where APIs allow — never presented as the invoice. That humility is a feature. Any multi-cloud cost tool that shows you a single blended number with no error bars is lying; OSR at least labels the lie.

Failover has two regimes, and only one is cheap​

This is the section that separates a serious design from a pitch deck. OSR splits failover in two, and prices them differently:

Create-time failover: cheap, always on. No state exists yet, so a CapacityError or provider outage just means trying the next scored candidate. This is the regime that would have absorbed Daytona's burst-load bad day in the benchmark above, or any regional outage, with zero workload changes. It is also the regime that makes the Daytona closed-source episode survivable: new sandboxes route elsewhere by policy while you migrate old ones deliberately.

Mid-session migration: opt-in, best-effort, capability-gated. Moving a running sandbox requires snapshot on the source and restore on the target, plus a rebind of the durable binding. The spec documents the hard limit explicitly: filesystem state can move (with a portable tar export/import fallback), but live in-memory process state is not guaranteed across heterogeneous providers. A checkpoint taken inside a Firecracker microVM does not resume inside a gVisor sandbox, and no control plane can fix that — it is a property of the isolation technology, not the API. Automatic migration on provider degradation exists as a policy (onProviderDegraded: "migrate" | "fail"), but the default posture is fail loudly rather than migrate lossily.

That honesty extends to the roadmap: mid-session migration is scoped to filesystem-state portability first, MVP is create-time failover only, and the spec lists divergent capabilities as its number-one risk. Believe a roadmap that names its hardest problem first.

What the abstraction leaks​

Three things leak through every provider-agnostic sandbox API, and OSR handles each differently:

1. Isolation strength. E2B and Vercel run Firecracker microVMs (separate kernel per sandbox), Modal runs gVisor on KVM, Daytona runs OCI containers with optional Kata/Sysbox hardening. These are genuinely different security postures for mutually untrusted tenant code, and the spec refuses to abstract them away: isolation is a policy floor (Firecracker > gVisor > container), and the router never places below it. If your threat model says microVM-or-nothing, half the candidate set is correctly excluded on every request. A router that hid this would be a vulnerability laundering machine.

2. Adapter maintenance. Provider APIs churn — the spec names Modal's filesystem API deprecation and Daytona's license change as the kind of drift adapters must absorb. The mitigation is thin adapters, per-provider contract tests, community-owned adapters, and capability probing to catch drift early. This is the project's real ongoing cost, and its open-core model (gateway, adapters, SDKs, and routing all OSS; hosted extras paywalled but no core capability gated) is structured to distribute it. Still: count the adapters before you bet the platform. MVP promises E2B, Modal, Vercel, plus the self-hosted Kubernetes reference.

3. The self-hosted default is the whole game. The Kubernetes adapter — each sandbox a Pod behind a RuntimeClass you choose (gVisor, Kata, Firecracker via something like Firekube, or plain runc for trusted workloads) — is what makes the router a portability story instead of a meta-vendor story. Keep it as the default backend and the commercial providers become burst capacity and failover targets rather than load-bearing dependencies. This is converging with real code elsewhere: AgentENV already runs Firecracker microVMs behind an E2B-compatible HTTP API so existing E2B SDK code works unmodified, kubernetes-sigs/agent-sandbox is standardizing SandboxClaim CRDs with warm pools, and e2bgateway already routes an E2B-shaped API across E2B Cloud, agent-sandbox, and Alibaba's OpenSandbox. The E2B-compatible contract is becoming the portable narrow waist whether or not OSR wins — the router just makes it routable.

Benchmark once, fail over, self-host the default​

For a deploy-from-chat platform — or any agent product executing untrusted model output — the adoption shape the spec points to is concrete: declare capabilities per workload class, let the router benchmark and place, keep the self-hosted adapter as the default so your unit economics and your threat model both rest on infrastructure you control, and treat commercial backends as scored, policy-gated overflow. The durable binding plus audit log gives platform teams the per-tenant attribution that raw provider SDKs never provide.

Two caveats before you npm install anything. First, the spec is draft v0.1: the routing pipeline, capability model, and cost normalizer are designed, not shipped, so adopt the shape (capability declarations, create-time failover, self-hosted default) even if you hand-roll the first version. Second, E2B's repeated appearances across every agent framework show where the gravity is: build against the E2B-compatible contract and you inherit portability across AgentENV, agent-sandbox, e2bgateway, and OSR's future adapters alike.

The sandbox market spent 2026 fragmenting — new entrants, diverging licenses, three billing philosophies. A router that is honest about what it can and cannot abstract is exactly the right response: fail over what is stateless, declare what is required, and own the default.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide