Skip to main content

Anthropic Split the Agent's Brain From Its Hands: Why Infrastructure Is Now the Bottleneck

8 min readDora NodaDora Noda
Share
On this page

On April 8, 2026, Anthropic stopped being just a model company. The launch of Claude Managed Agents in public beta — a fully hosted runtime for autonomous agents, complete with sandboxed code execution, persistent sessions, and built-in tools — was interesting on its own. But the more important artifact was the engineering post behind it: "Scaling Managed Agents: Decoupling the Brain from the Hands." In it, Anthropic virtualized the agent into three independent primitives — a session, a harness, and a sandbox — and argued, in effect, that once model quality plateaus, the bottleneck to running agents in production is not intelligence. It is infrastructure: sandboxed execution, checkpointing, credential scoping, and recovery.

That framing is a gift to anyone deciding where each piece of the agent stack should run. Because once the brain is decoupled from the hands, you get to ask an uncomfortable question about the hands: why am I renting these by the second?

The three primitives, concretely​

Anthropic's decomposition is clean enough to state in one table. A session is the append-only event log of everything that happened. A harness is the loop that calls Claude and routes its tool calls to the relevant infrastructure. A sandbox is the isolated execution environment where code runs and files change. The brain (Claude plus harness) is stateless. The hands (sandboxes) are disposable. The session log lives outside both, so if a harness crashes, a new one calls wake(sessionId), replays events, and picks up where the last one stopped.

PrimitiveWhat it isManaged Agents (Anthropic runs it)Self-hosted equivalent (you run it)
SessionAppend-only log of every eventDurable session store, wake(sessionId) recoveryPostgres-backed event log, git-checkpointed progress files
HarnessLoop: call model, route tool callsModel selection, retries, scheduling, gradingYour orchestrator (Temporal, LangGraph, a cron job) + any model API
SandboxIsolated code execution + filesSandboxed Linux containers, built-in bash/fs/web toolsFirecracker microVMs, gVisor, Docker via Daytona or raw hosts

The June 2026 extension (announced at Code with Claude events in San Francisco and Tokyo) added three more managed capabilities on top: dreaming (a scheduled pass that curates the agent's memory between sessions), performance outcomes (rubric-based grading of every run), and multi-agent orchestration. Note what all three have in common: they are harness and session features. None of them change the sandbox story. The hands remain a separable, replaceable layer — which is exactly why the decoupling matters.

Why the split is the whole point​

Before this architecture, "running an agent" meant running one big stateful blob: the conversation history, the working directory, the credentials, and the retry logic all lived in the same process on the same machine. If it crashed at hour three of a migration job, you lost the thread — literally. Anthropic's own long-horizon agent work identified two dominant failure modes: agents trying to do too much at once, and agents prematurely declaring success after partial progress. Both are mitigated by externalizing state: checkpoint progress to git or a structured log at every meaningful unit of work, so any restart reconstructs state from artifacts instead of from in-context history.

Independent failure and replacement of each component is the payoff. The sandbox can fail and be re-provisioned without data loss because the session log is durable and external. The harness can be upgraded — new model, new retry policy — without touching execution. And crucially, each layer can live with a different owner: Anthropic's brain driving sandboxes you control, reaching your private services through outbound-only MCP tunnels. The Totalum production playbook for Managed Agents describes exactly this shape: Anthropic orchestrates, schedules, and recovers, while tool execution runs in a sandbox you control with vault-stored secrets injected per session and network egress gated to an allow-listed tunnel set.

That is the shape to internalize, because it turns a vendor decision into an infrastructure decision. The brain is a commodity you shop for by quality and price. The hands are infrastructure you own, tune, and keep your data inside.

The cost math of hands​

Here is where decoupling gets concrete. Managed sandboxes bill per second, and agent workloads are long: a coding agent that runs for ten minutes per task, a thousand tasks a day, is 10,000 sandbox-minutes daily. At E2B's published default rate of about $0.0000303 per second for a 2 vCPU / 512 MiB sandbox, that workload costs roughly $545 per month — and that is compute alone, before model tokens. A 2026 comparison of E2B alternatives puts the same workload on a self-hosted Daytona setup at roughly $6–12 per month of flat VPS cost.

Workload: 1,000 tasks/day × 10 minManaged (E2B default sandbox)Self-hosted (Daytona on your VPS)
Sandbox compute~$545/mo (per-second meter)~$6–12/mo (flat host)
IsolationFirecracker microVMsDocker (Kata/Sysbox optional)
Data residencyVendor cloudYour server

The variable that drives the result is volume. At ten tasks a day, the managed meter rounds to pocket change and zero ops wins. At a thousand tasks a day, the per-second meter is a ~50x markup over owned compute doing the same work. Sensitivity runs one direction: the more your agents run, the worse renting hands looks. That crossover — not any feature checklist — is the real reason teams graduate from managed sandboxes to owned execution infrastructure.

There is a second cost that never appears on the invoice: data leaving your boundary. Every file the agent reads, every credential it touches, every private repo it clones flows through the vendor's sandbox fleet under the fully managed shape. For side projects that is fine. For production agents operating on customer code and production credentials, residency is the feature, and only the self-hosted column has it.

What "hands" must actually provide​

Owning execution is not just running containers. Anthropic's framing is useful as a requirements checklist — five capabilities that separate a real hands layer from a Docker daemon with ambitions:

  1. Isolation tier matched to trust. Untrusted agent-generated code wants kernel-level boundaries: Firecracker microVMs (E2B's choice), gVisor (Modal's choice), or Kata containers. Trusted internal workloads can live in plain Docker. Pick per workload, not per platform — a hands layer should offer both.
  2. Checkpointing and durable sessions. Long-running agents must survive disconnection, eviction, and timeout. That means progress externalized continuously — structured logs, git commits per unit of work, or an append-only event store — so a fresh sandbox resumes instead of restarting.
  3. Credential scoping. Secrets must be injected per session without ever landing in prompt context or persisting in the sandbox image. Vault-backed, short-lived, least-privilege: the agent gets exactly the keys its task needs, and they expire with the session.
  4. Egress control and private access. The sandbox reaches internal services through outbound-only tunnels (the MCP-tunnel pattern), with an allow-listed destination set. No inbound ports, no ambient network access, no exfiltration path wider than the task requires.
  5. Scheduling, retry, and lifecycle. Sandboxes need provisioning in seconds, idle reaping, retry-with-backoff on failure, and cron or event triggers for recurring agents. This is the unglamorous half of the hands layer and the half most DIY setups skip until the 3 a.m. page.

Every item on this list is infrastructure work, not model work. None of it improves when the next checkpoint drops. That is precisely Anthropic's point: the intelligence layer keeps getting better on someone else's roadmap, while the execution layer is where your operational leverage compounds — if you own it.

Where a self-hosted deploy platform sits​

Stack the checklist against what a Cluster-API-based platform already does — declarative machine lifecycle, container scheduling on owned hardware, persistent volumes, secret management, network policy — and the overlap is nearly total. Provisioning a sandbox per agent session is a small step from provisioning an app per git push: same nodes, same isolation primitives, same secret store, with session-scoped credentials and idle reaping added on top. The platform becomes the hands layer underneath any vendor's brain — Claude today, whatever tops the evals tomorrow — without re-architecting when the model changes.

That is the durable position in the decoupled stack. Model vendors will keep competing the brain toward commodity pricing. The session log is a database. But the hands — isolated, checkpointed, credential-scoped execution on machines you control — is the layer where residency, cost at volume, and operational control all land. Anthropic's own architecture says so: it made the hands disposable and replaceable, which is another way of saying it made them yours to own.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. The same primitives that deploy your apps — declarative infrastructure, isolated execution, secret management — are the hands layer your agents need. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide