An agent shipped to production at 2am. The deploy went out clean, the health checks passed, and nobody thought about it again until the incident review — when someone asked the question nobody could answer: whose credentials did it use?
This is not a hypothetical. In March 2026, a confused-deputy flaw in FastMCP's OAuth proxy (CVE-2026-27124) let an attacker perform actions on behalf of victims through a GitHub OAuth flow without explicit consent, because consent verification was missing from the callback. The tool call executed. The authorization behind it belonged to nobody in particular. For a platform that ships deploy-from-chat as MCP tools, that gap — between "the agent called the tool" and "a specific human authorized this specific action" — is the entire identity problem.
Three identities hide inside every agent deploy call
When a developer asks a chat agent to deploy, three distinct principals are actually involved, and a secure platform has to keep all three visible at the moment of the tool call:
| Identity | Who it is | What it contributes |
|---|---|---|
| The developer | The human who connected the agent and typed the request | Consent — this action is wanted |
| The agent | The service identity running the session (the host, the runner) | Execution context — this software is acting, with these limits |
| The tenant org | The team being billed and blast-radius-bounded | Boundary — this action lands here, and only here |
Watch what happens when these collapse into one token. A deploy tool call arrives at the MCP server carrying a bearer token. The server must answer three checks before it runs: did a human approve this action? (developer), is this caller the agent runtime we issued credentials to, operating inside its task scope? (agent), and does this deploy target a service owned by the tenant this session belongs to? (tenant). A single inherited user token answers the first question and waves through the other two. That is exactly the shape of the FastMCP failure: a token that looked like authorization, with no binding between the consent, the actor, and the target.
The fix is not "more OAuth." It is making each of the three identities explicit, independently verifiable, and recorded — at every tool call, not just at login.
What the protocol started requiring in July 2026
The MCP specification revision of July 28, 2026 — the largest since the protocol launched — dragged MCP authorization much closer to real OAuth deployments. The headline changes for anyone running a remote MCP server:
- OAuth 2.1 with PKCE is mandatory for remote servers. API keys pasted into an MCP config are not spec-compliant for remote servers in 2026; they survive only in local-server contexts like IDE assistants.
- Dynamic Client Registration is deprecated in favor of Client ID Metadata Documents (CIMD), tightening how a client proves what it is before it ever sees a token.
- Tokens are audience-bound. A token issued for one server cannot be replayed against another — the direct fix for the token-relay class of attacks.
- Token passthrough is explicitly forbidden. An MCP server must not forward a token it received as if it were its own credential. That pattern is the confused deputy, and the spec now says so.
- A client-credentials extension (
io.modelcontextprotocol/oauth-client-credentials) covers machine-to-machine authentication for unattended agents, CI jobs, and workers — supplementing, not replacing, the user authorization flow.
The stricter discovery and issuer-validation rules matter too: a 2026 flaw in the MCP Python SDK showed a client that failed to validate the authorization-server issuer or bind stored credentials to the server they belonged to — letting a malicious server redirect the client to an attacker-chosen token endpoint and harvest the client secret, the authorization code, and the PKCE verifier. Every one of those fields exists to bind this token to this server and this session; skip the binding and OAuth becomes theater.
Delegate with token exchange, not token inheritance
So how should the three identities actually travel to the tool call? The industry answer is converging on RFC 8693 OAuth 2.0 Token Exchange: the agent presents the human's token as the subject and its own credential as the actor, and the authorization server mints a derived token carrying both — sub for the human the agent acts for, act for the agent itself — with scopes narrower than the human's original grant. An IETF draft from February 2026 (draft-oauth-ai-agents-on-behalf-of-user) extends this specifically for AI agents, adding requested_actor to name the agent in the authorization request and actor_token to authenticate it during the code-for-token exchange.
A deploy-from-chat session built on this pattern looks like this:
- The developer connects the agent via the normal OAuth 2.1 + PKCE flow and consents to a bounded scope set.
- The agent authenticates itself with its own credential — preferably a signed JWT assertion tied to verified runtime identity, not a reusable secret.
- For the deploy task, the authorization server mints a delegation token recording "agent X acting on behalf of developer Y," scoped to exactly the tools this task needs.
- Each tool call is authorized against the intersection of the agent's role, the user's authority, and the task context — never against the user's full entitlement set.
That per-tool scoping is where multi-tenancy lives or dies. A deploy tool and a log-reading tool must not share a scope:
| Tool | Required scope | Why it is separate |
|---|---|---|
deploy | deploy:write on the tenant's services | Mutates production; needs fresh human consent |
rollback | deploy:write on the tenant's services | Same blast radius as deploy, different direction |
scale | deploy:write on the tenant's services | Changes spend; bounded by tenant quota |
logs:read | logs:read on the tenant's services | Read-only, but still tenant-bounded — no cross-tenant reads |
Enterprise guidance for MCP servers converges on the same shape: scopes structured per resource, per verb, and per sensitivity level, with bounded lifetimes and immediate revocation. The agent that only needed to read logs should hold a token that cannot deploy, even if a poisoned tool description tricks it into trying.
A leaked agent credential must not outlive the chat
Long-running agent sessions break the token-lifetime assumptions that human sessions were designed around. A developer's session lasts minutes; an agent's task can persist for hours or days — and a credential minted for that task stays valid exactly as long as its lifetime says, regardless of whether the chat that authorized it is still open.
The practical policy, repeated across 2026 MCP hardening guidance, has three parts. First, bound access-token lifetimes tightly — one hour or less, with transparent refresh and immediate revocation — so a leaked token is a small window, not a standing key. Second, issue credentials per agent task, not per agent, so revoking the task's grant kills exactly that task's access. Third, prefer sender-constrained credentials for unattended agents: signed JWT assertions tied to runtime identity over reusable secrets, so a copied bearer string alone buys an attacker nothing.
Consider the failure mode this closes. A chat-minted credential leaks — pasted into a log, echoed into a tool result, exfiltrated through a prompt-injection side channel. With a months-long, hard-to-revoke token (a pattern still documented in production MCP deployments), that leak is a months-long impersonation of the developer. With a one-hour, audience-bound, task-scoped delegation token, it is a one-hour window to call one tenant's tools from one runtime — still bad, but bounded, attributable, and revocable. The lifetime is the blast radius.
The audit log is the product
When the 2am deploy becomes an incident, the audit log is the only witness. "The agent did it" is not attribution. Every tool call should log the full delegation picture: who approved, who delegated, what resource was touched, and whether the action was autonomous or human-gated.
A useful entry records at least this:
{
"tool": "deploy",
"tenant": "acme-corp",
"principal": "developer:ada@acme.example",
"actor": "agent:deploy-chat-runner-7f3a",
"delegation": "oauth-token-exchange",
"scopes": ["deploy:write"],
"approval": "human-confirmed",
"target": "svc/api-prod",
"decision": "allow"
}Note what is present that a naive "agent used the user's token" log would lack: the tenant boundary, the agent's own identity distinct from the human's, the delegation mechanism, and the approval mode. Without those fields, you cannot distinguish "the developer approved this deploy" from "the agent acted on stale intent with a still-valid token" — and that distinction is the entire incident review.
The shortcut a multi-tenant platform cannot afford
The tempting shortcut is "the agent just uses my token." It removes a consent screen, a token exchange, and a scope-design meeting. It also collapses least privilege: the agent can suddenly do everything the user can, access stays continuous after the original intent goes stale, and tool chains amplify one grant into many.
January 2026 research from Obsidian Security, summarized in the OWASP MCP security guidance, showed how concretely this breaks: shared client IDs enabling consent-caching bypass, URL capture bypassing MCP-layer consent entirely, cookie injection forcing state validation under attacker control — with Square's MCP server exposing merchant data, transactions, and banking details in the blast radius. Separately, Manifold Security found a confused-deputy flaw in Microsoft's official Azure DevOps MCP server, where a tool returning pull-request descriptions verbatim opened a prompt-injection path into agent hijacking. None of these required exotic attacker capabilities. All of them required an identity model that assumed a token meant what it looked like.
The NSA's 2026 cybersecurity information sheet on MCP and the OWASP MCP Top 10 both land in the same place: traditional AppSec tooling cannot secure agentic workflows because the gap is runtime governance — multiple agents sharing servers with no native identity boundary. The boundary has to be built per request: validate the actual audience and client on every call, never cache an authorization decision in session state and trust it later.
Shipping deploy-from-chat without the identity hole
For a self-hosted PaaS, the checklist falls out of the three-identity model directly. Run the July 2026 authorization profile: OAuth 2.1 + PKCE, CIMD clients, audience-bound tokens, no passthrough. Mint delegation tokens per task via token exchange, with sub naming the developer and act naming the agent. Scope each tool per resource and verb, tenant-bounded. Keep lifetimes short and revocation instant. And log the human principal behind every agent session, because the incident review will ask.
Get those five right and the 2am deploy stops being a mystery: the log says which developer consented, which agent executed, which tenant paid, and which scope allowed it. That is what "the agent deployed on the tenant's behalf" has to mean before a multi-tenant platform can claim it.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with AI agents as first-class operators. Star the repo on GitHub or deploy your first app today.



