Point your coding agent at http://ai, hand it a fake key, and it just works — the real Anthropic key never leaves the gateway, and every token gets billed to the human or device that actually spent it. That is the whole pitch of Aperture, Tailscale's AI gateway, which went from an internal LLM proxy to general availability on August 26, 2026. The mechanism is worth understanding precisely, because a WireGuard-mesh vendor building a credential broker specifically for model API keys tells you where the next credential-sprawl problem actually lives — and what a self-hosted equivalent has to replicate.
Here is the request lifecycle, all five steps of it. The client names a model, not a provider: it sends an ordinary Anthropic- or OpenAI-shaped request to the gateway's tailnet address. The gateway resolves the caller's Tailscale identity — user email or device tag — from the connection itself via WhoIs, so there is no client secret to present or steal.
A deny-by-default grant check runs against that identity: without a matching tailscale.com/cap/aperture grant, nothing is allowed. On a match, the gateway injects the real upstream provider key, which lives only in the gateway's own configuration, and forwards the request. Telemetry — full request and response, token counts, duration, tool use, session grouping — is captured asynchronously after the response goes back.
The grant is the interesting object, because it is where "network identity" stops being a slogan and becomes policy:
{
"grants": [{
"src": ["group:engineering"],
"app": {
"tailscale.com/cap/aperture": [
{ "role": "user" },
{ "models": "anthropic/**" }
]
}
}]
}Roles are user and admin, model patterns are globs, and role is required — a matching model pattern with no role still gets HTTP 403. Grants can live in the Aperture configuration or the tailnet policy file; the two forms differ only in the dst field. A connectors entry extends the same grant to MCP tools and HTTP connector proxies, which matters more than it looks: the gateway stopped brokering just model keys months ago and now brokers agent tool access too.
The self-hosted parts list
That lifecycle decomposes into seven capabilities. Each one has a self-hosted building block — and each block has a gap where the hosted control plane was doing invisible work. Here is the whole replication surface in one table:
| Aperture capability | Self-hosted building block | The gap you must close yourself |
|---|---|---|
| Identity (WhoIs on every connection) | Headscale + WireGuard peer identity, or mTLS/SPIFFE IDs on your mesh | WhoIs-equivalent lookup: mapping a live connection to a user or device tag at request time, including off-mesh callers Aperture serves via its CLI bridge |
| Routing + upstream key injection | LiteLLM, Envoy AI Gateway, or any OpenAI-compatible proxy holding the real keys | Model-name routing across providers with per-provider auth shapes (headers, Bedrock signing, Vertex) |
| Grants (deny-by-default, glob models) | OPA, Cedar, or gateway-native policy | Same-schema policy in two places: Aperture grants live in gateway config or tailnet policy with one capability string |
| Token-bucket quotas per user/device | Gateway metering + a bucket per identity with refill | Per-person buckets keyed on mesh identity, not API key; on_exceed: reject semantics (HTTP 429) |
| Telemetry + session tracking | Log pipeline (OTel/Cribl-style export; Aperture itself exports to S3) | Session grouping: auto-detecting Claude Code/Codex session IDs so related requests stay together for forensics |
| MCP server proxying + connectors | Self-hosted MCP proxy with per-tool grants | Per-user OAuth connectors and the audit join between a model call and the tool calls it triggered |
| Sandboxed code execution | gVisor, Firecracker, or container sandboxes | Tying sandbox runs back to the same identity and budget as the chat session that spawned them |
Two rows do most of the work. Identity is the load-bearing one: everything downstream — grants, quotas, audit — keys off "who is on this connection" rather than "what secret did they present." A tunnel-bound or mTLS-bound identity cannot be pasted into an attacker's endpoint, which collapses the prompt-injection-to-credential-theft chain into nothing replayable. Session tracking is the underrated one: per-request logs tell you what happened, but only session grouping tells you which agent run did it, and that is the unit you actually investigate.
The honest summary of the table: routing, quotas, and policy are commodity parts, while mesh-native identity resolution and session-aware audit are the custom engineering. Anyone scoping a self-hosted Aperture should budget accordingly.
Spend attribution is a join, not a counter
Per-user audit logs and token-spend leaderboards sound like a dashboard feature. They are really a data join your gateway has to perform on every request: identity × model × tokens × price. Aperture centralizes OpenAI, Anthropic, Gemini, OpenRouter, and self-hosted endpoints behind one gateway, which means the join has to work across all five — and the fifth one breaks the pattern, because self-hosted tokens have no price.
Here is what one user's day looks like as a ledger (prices illustrative — the point is the shape, not the rates):
| User | Endpoint | Model | Tokens (in/out) | Unit cost source | Attributed cost |
|---|---|---|---|---|---|
| priya | Anthropic | claude-sonnet | 1.2M / 180K | provider price sheet | $5.94 |
| priya | OpenAI | gpt-5-mini | 4.0M / 600K | provider price sheet | $2.70 |
| priya | OpenRouter | open-weight routed | 900K / 120K | OpenRouter receipt | $0.41 |
| priya | Gemini | gemini-flash | 2.5M / 300K | provider price sheet | $1.05 |
| priya | self-hosted (vLLM) | llama-8b | 6.0M / 900K | GPU-hour allocation | $1.80 |
The first four rows are lookups: tokens counted at the gateway, price from a sheet, receipt reconcilable against the provider bill. The fifth row is an allocation your platform invents: tokens divided by throughput, times the GPU node's hourly cost, charged to whoever the mesh identity says ran them. Aperture's token-bucket quotas (capacity, rate, on_exceed: reject) then enforce against these attributed totals per user or per device — which only works if the self-hosted row exists. A gateway that meters cloud tokens precisely and treats the on-prem GPU pool as free will watch every cost-capped agent route itself onto the "free" models.
That is the row a self-hosted build most often skips, and the one that decides whether spend controls survive contact with a hybrid fleet. If you take one requirement from this post, take this: price the self-hosted endpoint, or your quotas are decorative.
Ten months from proxy to platform, compressed
Aperture's trajectory reads as a map of where sprawl kept moving. Three dates carry it: open alpha on February 17, 2026 as an LLM proxy — one-line client change, real keys centralized, per-identity usage tracking and full session histories from day one. Self-serve followed on March 23, and by mid-2026 the product had grown guardrails (pre-request hooks), an MCP server proxy, outbound connectors with per-user OAuth, and integrations with Oso, Cerbos, and Cribl. GA on August 26, 2026 added built-in purchasable tokens for open-weight and closed models, a chat UI with Projects, sandboxed code execution, and two MCP endpoints letting agents provision tailnet nodes over Tailscale SSH with per-action approval and logging.
Each expansion chased the same problem into a new hiding place. Keys first, then spend, then the tool calls agents make with the access they were given. By GA the product's own summary was "models, tools, and infrastructure without handing over API keys or unrestricted access" — notice that keys are only the first noun. The sprawl was never just the strings; it was every standing permission an agent held, and the gateway kept widening until it brokered all of them.
Where the sprawl actually lives
That widening is the signal worth taking seriously. Service-to-service secrets are a solved problem — Vault, OpenBao, and Infisical won that war, and nobody building a new platform hand-rolls env-var secret distribution for service auth anymore. The unsolved surface is everything an agent touches that is not service-to-service:
| Service-to-service secrets | Per-agent / per-tool model access | |
|---|---|---|
| Example | Database passwords, internal API tokens | Model keys per agent, MCP server credentials, sandbox permissions |
| Lifecycle | Long-lived, rotated on schedule | Minted per run, per tool, per session; often never revoked |
| Broker | Vault/OpenBao/Infisical (solved) | AI gateways like Aperture (being built now) |
| Failure mode | Leaked string in a repo | Replayable Bearer token the workload itself can be tricked into sending anywhere |
Classic virtual-key proxy gateways centralize billing but keep the core failure: the agent still holds a secret string that is valuable anywhere, so anywhere the agent can be tricked into sending it becomes a full compromise. Identity-based designs — Aperture on the tailnet, Pangolin over its own SSO-backed tunnels — try to shrink the number of replayable secrets toward zero instead of managing them better. I compared those two patterns directly in a previous post on Pangolin's AI gateway; the short version is that both authenticate the connection rather than the request, and differ mostly in which network identity you already have.
The reason a mesh vendor built this, rather than a secrets vendor, is that the mesh already owns the identity the gateway needs. Tailscale did not have to invent an agent identity layer — every device and user on a tailnet already has one, and WhoIs already resolves it per connection. The gateway is a small step once identity exists; without identity, it is a key warehouse with a dashboard. That ordering is the real lesson: identity first, broker second.
Build order for a PaaS team
For a team running its own machines, the parts list from the top of this post turns into a build order. Do it in this sequence, and defer everything else until the earlier layer holds:
- Mesh-identity gateway first. Put one proxy in front of every model endpoint — cloud and self-hosted — and authenticate callers by mesh identity (WireGuard peer, mTLS/SPIFFE ID), never by a distributed key. No other layer matters until the agent holds no replayable secret.
- Quotas second, with the self-hosted row priced. Per-user and per-device token buckets with reject-on-exceed, joined against provider sheets for cloud models and GPU-hour allocation for your own endpoints.
- Session-aware audit third. Request logs grouped by agent session, exportable to your existing pipeline, so one run's model calls and tool calls stay together.
- MCP and tool brokering fourth. Extend the same grants to agent tool access once model calls are governed — this is where Aperture itself went next, and for the same reason.
- Sandboxing last. Tie code execution to the same identity and budget only after the first four layers attribute correctly.
What to defer deliberately: built-in token resale (a billing business, not a gateway feature), a chat UI (your users already have clients — the one-line base-URL change is the integration that matters), and fine-grained guardrail hooks (valuable, but premature before identity and metering are right).
The industry direction is unmistakable either way: agents authenticate as themselves, not with borrowed secrets. Teams that adopt that posture early — on infrastructure they own, where the gateway and the models can live on the same machines — get the security win and the cost-visibility win together.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



