Daytona took its core closed-source in June 2026 — and every team that had built directly on its SDK learned the same lesson at the same time: a sandbox vendor is a single point of failure you chose on purpose. Open Sandbox Router, a provider-agnostic control plane for code-execution sandboxes, proposes the fix: one Sandbox handle, one TypeScript SDK, and a router that places each sandbox on E2B, Modal, Vercel Sandbox, Daytona, or your own Kubernetes cluster based on declared capabilities, cost, and live health.
The verdict up front: the design's core insight is correct — sandboxes are stateful, so routing is a create-time decision with a durable sandbox-to-provider binding — and create-time failover plus capability filtering are genuine wins. But mid-session migration across heterogeneous isolation tech is honestly best-effort, and no router can paper over the fact that Firecracker and a shared-kernel container are different security postures. Here is the concrete evaluation: the fragmentation in numbers, the mechanism, the cost table, the two failover regimes, and what leaks.
The fragmentation it attacks, in numbers
The pain is easy to quantify. Every provider ships its own SDK, resource model, template format, and lifecycle semantics: execution lives on Sandbox for Vercel and Modal, under Sandbox.commands for E2B, and under Sandbox.process for Daytona. Switching providers is a rewrite, not a config change.
Even the OpenAI Agents SDK ships separate sandbox backend extensions for E2B, Modal, Daytona, Cloudflare, Runloop, Blaxel, and Vercel — fragmentation so thorough that the ecosystem's center of gravity just carries all of it.
Billing is three different philosophies wearing the same "per-second" costume: Vercel meters active CPU (idle time waiting on the network or the LLM is free), E2B meters session time, Modal meters per-second provisioned. And cold starts span an order of magnitude. The August 2026 single-harness benchmark from MarkTechPost — time-to-interactive from create() to first successful command, 100 iterations per provider in one concurrent burst — measured:
| Provider | p50 TTI | p95 TTI | Success |
|---|---|---|---|
| Daytona | 0.27s | 0.43s | 37% |
| Vercel Sandbox | 0.67s | 1.04s | 100% |
| Modal | 0.88s | 1.00s | 100% |
| E2B | 1.61s | 1.77s | 100% |
That table is the whole argument for measuring once through one SDK and routing forever: the fastest provider in the run also failed 63% of creations under burst load. A router with live health signals would have steered around that bad day automatically; a hardcoded SDK choice eats it. Note the self-hosted row the spec adds for contrast — 10–150ms on your own Kubernetes, where the only queue is yours.
Then there is the exhibit A for portability: Daytona's core development moved to a private codebase in June 2026, freezing the public AGPL repo at v0.190.0 with no further updates. The hosted product continues (managed control plane plus bring-your-own-compute), and the community answered with the Nightona fork.
But anyone who treated "open source, self-hostable" as a permanent property of their Daytona dependency got a forced migration. We covered Daytona's customer-managed compute story before the freeze; the freeze is what happens when your abstraction layer is somebody else's repo banner.
Capability negotiation, concretely
The mechanism that makes one API span five backends is capability negotiation, and the spec treats it as first-class rather than a filter bolted on later. A create request declares required capabilities — { gpu, snapshot, ports, minMemory, region } — and the router runs a pipeline: filter the candidate set to providers that can comply, score the survivors on cost/latency/region/reliability, order them by policy, and attempt in order with failover. The principle is typed capabilities over silent degradation: if nothing can meet the request, you get a typed error, not a sandbox that quietly lacks the GPU you asked for.
Two pieces of supporting machinery make this more than a load balancer. First, a portable template spec: one Dockerfile-based osr.template.yaml (OCI base, setup steps, env, workdir, exposed ports) that osr template build compiles into each provider's native format — E2B template, Modal image, OCI image for self-hosted — with refs tracked in a content-hash-keyed registry. Content hashing means "same declared environment" is verifiable, not aspirational.
Templates are the second-biggest portability pain after the API itself, which is why the spec tackles them before any higher-level orchestration.
Second, the durable binding: sandbox → { provider, region, capabilities, creds ref } persisted in Postgres (gateway mode) or an embedded store (library mode). Every non-create operation resolves the binding and dispatches to the home provider — there is deliberately no re-routing a live sandbox. A reaper enforces TTLs and reconciles orphaned provider resources, and idempotency keys on create prevent duplicate provisioning on client retries. The overhead budget is explicit: under 50ms of added control-plane latency, with a direct-connect mode planned so the gateway leaves the exec hot path entirely.
The cost table no single vendor will show you
Here is the normalized comparison the spec's economics section is built to produce continuously, assembled from 2026 provider pricing:
| Backend | CPU | Memory | Billing basis |
|---|---|---|---|
| E2B | $0.0504 / vCPU-hr | $0.0162 / GiB-hr | Session time; $150/mo Pro tier; BYOC available |
| Modal | $0.1419 / core-hr | $0.0242 / GiB-hr | Per-second provisioned; strongest GPU story (T4–H100) |
| Vercel Sandbox | $0.128 / vCPU-hr | $0.0212 / GB-hr | Active CPU only + provisioned memory; GA January 2026 |
| Daytona | $0.0504 / vCPU-hr | $0.0162 / GiB-hr | Usage-based; hosted product after the June 2026 OSS freeze |
| Self-hosted Kubernetes | Your hardware | Your hardware | Flat infra cost + your ops time; 10–150ms cold starts |
But "cheapest" depends entirely on workload shape, which is exactly why a static table goes stale and a routing layer earns its keep:
- Bursty agent loops (short exec bursts separated by LLM think time) favor Vercel's active-CPU model — the think time is unbilled — while session-time and provisioned models charge through the idle gaps.
- Long-lived coding-agent workspaces (repo checked out, dev server on an exposed port, multi-hour sessions) favor flat-rate or self-hosted capacity, where the marginal hour is near zero.
- GPU inference or training bursts collapse the candidate set to Modal (or your own GPU nodes) regardless of CPU pricing — capability filtering does its job before cost scoring even runs.
The spec is admirably honest that cost normalization is approximate: active-CPU vs session vs provisioned cannot be reduced to one number without assumptions. The router's dollar figures are estimates for routing and reporting, reconciled against provider invoices where APIs allow — never presented as the invoice. That humility is a feature. Any multi-cloud cost tool that shows you a single blended number with no error bars is lying; OSR at least labels the lie.
Failover has two regimes, and only one is cheap
This is the section that separates a serious design from a pitch deck. OSR splits failover in two, and prices them differently:
Create-time failover: cheap, always on. No state exists yet, so a CapacityError or provider outage just means trying the next scored candidate. This is the regime that would have absorbed Daytona's burst-load bad day in the benchmark above, or any regional outage, with zero workload changes. It is also the regime that makes the Daytona closed-source episode survivable: new sandboxes route elsewhere by policy while you migrate old ones deliberately.
Mid-session migration: opt-in, best-effort, capability-gated. Moving a running sandbox requires snapshot on the source and restore on the target, plus a rebind of the durable binding. The spec documents the hard limit explicitly: filesystem state can move (with a portable tar export/import fallback), but live in-memory process state is not guaranteed across heterogeneous providers. A checkpoint taken inside a Firecracker microVM does not resume inside a gVisor sandbox, and no control plane can fix that — it is a property of the isolation technology, not the API. Automatic migration on provider degradation exists as a policy (onProviderDegraded: "migrate" | "fail"), but the default posture is fail loudly rather than migrate lossily.
That honesty extends to the roadmap: mid-session migration is scoped to filesystem-state portability first, MVP is create-time failover only, and the spec lists divergent capabilities as its number-one risk. Believe a roadmap that names its hardest problem first.
What the abstraction leaks
Three things leak through every provider-agnostic sandbox API, and OSR handles each differently:
1. Isolation strength. E2B and Vercel run Firecracker microVMs (separate kernel per sandbox), Modal runs gVisor on KVM, Daytona runs OCI containers with optional Kata/Sysbox hardening. These are genuinely different security postures for mutually untrusted tenant code, and the spec refuses to abstract them away: isolation is a policy floor (Firecracker > gVisor > container), and the router never places below it. If your threat model says microVM-or-nothing, half the candidate set is correctly excluded on every request. A router that hid this would be a vulnerability laundering machine.
2. Adapter maintenance. Provider APIs churn — the spec names Modal's filesystem API deprecation and Daytona's license change as the kind of drift adapters must absorb. The mitigation is thin adapters, per-provider contract tests, community-owned adapters, and capability probing to catch drift early. This is the project's real ongoing cost, and its open-core model (gateway, adapters, SDKs, and routing all OSS; hosted extras paywalled but no core capability gated) is structured to distribute it. Still: count the adapters before you bet the platform. MVP promises E2B, Modal, Vercel, plus the self-hosted Kubernetes reference.
3. The self-hosted default is the whole game. The Kubernetes adapter — each sandbox a Pod behind a RuntimeClass you choose (gVisor, Kata, Firecracker via something like Firekube, or plain runc for trusted workloads) — is what makes the router a portability story instead of a meta-vendor story. Keep it as the default backend and the commercial providers become burst capacity and failover targets rather than load-bearing dependencies. This is converging with real code elsewhere: AgentENV already runs Firecracker microVMs behind an E2B-compatible HTTP API so existing E2B SDK code works unmodified, kubernetes-sigs/agent-sandbox is standardizing SandboxClaim CRDs with warm pools, and e2bgateway already routes an E2B-shaped API across E2B Cloud, agent-sandbox, and Alibaba's OpenSandbox. The E2B-compatible contract is becoming the portable narrow waist whether or not OSR wins — the router just makes it routable.
Benchmark once, fail over, self-host the default
For a deploy-from-chat platform — or any agent product executing untrusted model output — the adoption shape the spec points to is concrete: declare capabilities per workload class, let the router benchmark and place, keep the self-hosted adapter as the default so your unit economics and your threat model both rest on infrastructure you control, and treat commercial backends as scored, policy-gated overflow. The durable binding plus audit log gives platform teams the per-tenant attribution that raw provider SDKs never provide.
Two caveats before you npm install anything. First, the spec is draft v0.1: the routing pipeline, capability model, and cost normalizer are designed, not shipped, so adopt the shape (capability declarations, create-time failover, self-hosted default) even if you hand-roll the first version. Second, E2B's repeated appearances across every agent framework show where the gravity is: build against the E2B-compatible contract and you inherit portability across AgentENV, agent-sandbox, e2bgateway, and OSR's future adapters alike.
The sandbox market spent 2026 fragmenting — new entrants, diverging licenses, three billing philosophies. A router that is honest about what it can and cannot abstract is exactly the right response: fail over what is stateless, declare what is required, and own the default.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



