Skip to main content

Pangolin Puts SSO and WireGuard in Front of LLM Access Instead of API Keys: What Tunnel-Based Identity Means for Agent Credential Hygiene

10 min readDora NodaDora Noda
Share
On this page

Public GitHub commits introduced 28.65 million new hardcoded secrets in 2025, up 34% year over year — and the fastest-growing categories are now AI-service credentials, leaking 81% faster than the year before. AI-assisted commits leak secrets at roughly twice the rate of human-written ones. Your coding agents are burning through API keys, and the keys are ending up exactly where keys always end up: in repos, in config files, in chat threads. So when Pangolin — the self-hosted, WireGuard-based zero-trust access platform with over 22,000 GitHub stars — showed up on Hacker News in early September 2026 with an AI gateway whose pitch is "SSO and WireGuard instead of API keys for LLM access," it deserved more than a skim. The interesting move isn't another gateway. It's deleting the credential instead of managing it.

That is the core idea worth stealing: authenticate the network connection, not the request. Each user gets a WireGuard tunnel back to the gateway, established after they log in through their existing identity provider via a desktop app. The tunnel is the auth, so there is no key to generate, embed, rotate, or leak — the gateway already knows who is on the other end of the connection. This post walks through how that works mechanically, what changes when self-hosted models join the same gateway as public ones, how the pattern compares to Tailscale's Aperture and classic virtual-key gateways, and what it means for anyone running agents against models on infrastructure they own.

The tunnel is the credential​

Pangolin's architecture is a hub-and-spoke reverse proxy: Traefik does the HTTP proxying, a plugin called Badger authenticates every request, Gerbil manages the WireGuard server side, and Newt — a fully userspace WireGuard client — runs wherever the private resource lives, so nothing needs inbound ports or root-managed tunnels. The 1.22 release added an AI Gateway as a dedicated Pangolin resource: an identity-aware proxy between AI clients and model providers that administrators point coding agents at instead of configuring each agent against OpenAI or Anthropic directly.

The auth flow, as a Pangolin maintainer described it on the Show HN thread, runs like this. The desktop client opens a websocket control channel to the server, authenticated with the user's login token from SSO (Okta, Azure, Google, or whatever the org already uses). It shares a fresh WireGuard public key, the server provisions a peer, and the handshake completes — an outbound WireGuard tunnel from client to server. The client then overrides local DNS so agent traffic routes into the tunnel, and on the server side Pangolin trusts the WireGuard packet headers to associate the peer with an identity inside the gateway. No Bearer [REDACTED] ever crosses the wire from the agent, because there is nothing to send: identity arrives with the packets.

For the cases that cannot run the desktop client — headless CI machines, unattended agents, automation — Pangolin still issues traditional virtual keys, alongside the standard gateway features: session logging, budgets, usage analytics, and governance controls. That fallback matters and we'll come back to it, because every "no keys" story has a boundary where keys sneak back in, and the honest version of this pitch names it up front.

On-prem models join the same gateway, no inbound ports​

The second half of the announcement is arguably the bigger deal for self-hosters. The gateway supports OpenAI, Anthropic, Google Gemini, Amazon Bedrock, Google Vertex AI, Microsoft Foundry, OpenRouter, and the Vercel AI Gateway — and, through the same interface, self-hosted providers like Ollama, vLLM, and Bifrost. Multiple providers attach to one gateway, so a single Pangolin endpoint can serve Claude from Anthropic, GPT models from OpenAI, and a locally hosted Llama behind Ollama, with allow and block lists controlling which models are exposed through it.

Because the connector is just Newt — the same lightweight tunnel client Pangolin already uses for every other private resource — an on-prem model joins the gateway the way any private service does: drop the connector inside the cluster, even on something as small as a DGX Spark, and the model shows up next to the public cloud ones. Users switch between internal and external models without changing endpoints or juggling separate credentials for internal infrastructure.

There are no inbound firewall rules to open, no separate auth scheme for "the models behind the firewall" versus "the models with an API key." That symmetry is the part most AI gateways get wrong: self-hosted models are usually an afterthought bolted onto a proxy designed around forwarding static Bearer [REDACTED] upstream. Here the tunnel model makes them first-class by construction, since everything behind the gateway is reached over an outbound tunnel.

Three gateway patterns, compared​

It helps to place Pangolin's approach next to the two patterns teams actually run today: the classic virtual-key proxy gateway, and Tailscale's Aperture, which secures LLM access with tailnet identity instead of distributed keys.

Classic AI gateway (virtual-key proxy)Tailscale AperturePangolin AI Gateway
Auth primitiveStatic virtual key per client, proxy forwards a real provider key upstreamTailnet identity (WhoIs); one provider key stays inside the gatewayWireGuard tunnel + SSO identity; no client key at all
What the agent holdsA secret string in env/configA dummy placeholder key the gateway ignoresNothing (desktop tunnel) or a virtual key (headless)
Self-hosted modelsUsually a bolted-on custom endpointSupported alongside cloud providersFirst-class: Newt connector joins the same network
Network requirementOutbound HTTPS to the gatewayMembership in the tailnetPangolin client or connector with a tunnel
Fully self-hostableOften yes (open-source proxies exist)No — depends on Tailscale's control plane and identityYes — Community Edition is free and self-hosted
Audit storyPer-key logs and budgetsPer-identity logs, rate limits, session trailPer-identity logs, budgets, analytics, governance

Aperture and Pangolin rhyme: both move the real provider keys into the gateway and authenticate callers by who they are on the network rather than what secret they present. The difference is which network identity you already have. Aperture's answer is the tailnet — elegant if your org lives there, a non-starter if it doesn't. Pangolin's answer is a tunnel it provisions itself after SSO, which means it works for orgs with no tailnet and, critically, extends the same identity to models running on hardware the org owns. The classic proxy gateway keeps the key-sprawl problem and centralizes only the billing; both identity-based designs try to shrink the number of secrets that can leak in the first place.

What actually changes in your threat model​

This is where the pitch has to survive contact with the adversary, so let's be precise about what disappears and what doesn't.

What disappears is the largest, dumbest attack surface in agent operations: static provider keys sitting in agent configs, env files, and MCP server JSON. GitGuardian's 2026 State of Secrets Sprawl report found 24,008 unique secrets in public MCP configuration files alone, including over 2,100 still-valid credentials — and 1.27 million AI-service secrets overall. Keys in agent-accessible files aren't just leakable; they're stealable by the workload itself. The Claude Code case (CVE-2026-21852), where a malicious repo could redirect API requests — credentials included — to an attacker's endpoint, and the Amazon Q case (CVE-2026-12957), where a malicious MCP process inherited live AWS credentials, are both instances of one failure: the agent held a Bearer [REDACTED] that was valuable anywhere, so anywhere the agent could be tricked into sending it became a full compromise. A tunnel-bound identity can't be pasted into an attacker's endpoint. There is no string to exfiltrate, which collapses an entire class of prompt-injection-to-credential-theft chains into "the attacker gets nothing they can replay."

What doesn't disappear: a compromised endpoint still means a live identity. If the machine with the Pangolin client is owned, the tunnel is the attacker's tunnel until someone notices — revocation becomes peer removal rather than key rotation, which is arguably faster (kill the peer, the identity is gone everywhere at once) but still requires detection. The headless fallback is keys again: every CI runner and unattended agent on a virtual key is back in the rotation-and-leak business, just with a smaller blast radius than a raw provider key. And centralizing identity in the gateway makes the gateway itself the highest-value target — which is why the per-identity session logging and budgets aren't bonus features but load-bearing parts of the design.

Net: tunnel-based identity doesn't remove trust, it moves it from strings scattered across every agent's config to one gateway plus the endpoint's integrity. Given that 64% of secrets that were valid in 2022 were still unrevoked by January 2026, consolidating the thing you must revoke into one place you control is a genuine upgrade — as long as you count the virtual-key fallback honestly instead of pretending headless agents don't exist.

What a self-hosted PaaS should steal from this​

Strip away the product specifics and Pangolin's gateway is an argument about where agent credentials should live: not in per-provider env vars copied into every agent's environment, but in per-agent network identity issued at login and revoked in one place. Three pieces of that transfer directly to anyone running a deploy-from-git or deploy-from-chat platform on owned hardware.

First, per-agent identity beats per-provider keys as the unit of auth. A platform that provisions an identity per agent — tied to the agent's network attachment, not a pasted secret — gets revocation, scoping, and audit per agent for free. The moment your agents call models through env-var keys, you've recreated the exact sprawl GitGuardian keeps measuring, just inside your own fleet.

Second, the tunnel connector belongs inside the tenant's cluster. Pangolin's Newt-in-the-cluster pattern is the shape to copy: whatever serves models or tools inside tenant infrastructure should join the platform's network over an outbound-only connection, so tenant GPUs and internal endpoints appear in the same gateway as public providers with no inbound ports and no second credential scheme. For a platform with a roadmapped MCP server, this is the pairing that matters — the MCP server as the tool plane, the tunnel connector as the identity-bearing transport underneath it.

Third, treat the gateway as the audit boundary. Session logs, per-identity budgets, and allow/block lists per model are what make "the tunnel is the auth" operable instead of merely clever. An agent platform that can't answer "which agent spent what on which model this week" hasn't built credential hygiene; it's built a faster way to burn budget anonymously.

The industry direction is unmistakable: Aperture on the tailnet, Pangolin over its own tunnels, and every virtual-key proxy adding identity features are all converging on "agents authenticate as themselves, not with borrowed secrets." The teams that adopt that posture early — on infrastructure they own, where the gateway and the models can live on the same machines — get the security win and the cost-visibility win together.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide