Skip to main content

Shadow AI in the Pipeline: What to Audit From Laptop to Kubernetes Before Unvetted AI Becomes a Supply-Chain Bridge

13 min readDora NodaDora Noda
Share
On this page

Nearly every organization now has employees using AI tools nobody approved, nobody owns, and nobody monitors — and the pipeline is where that habit turns into a breach path. Verizon's 2026 Data Breach Investigations Report found unsanctioned AI use tripled in twelve months, from 15% to 45% of the workforce, making shadow AI the third most common non-malicious insider action it detects. IBM's 2025 Cost of a Data Breach Report put a price on it: organizations with high shadow-AI exposure averaged $4.63 million per breach, $670,000 more than those without. And in December 2025, Aikido Security disclosed PromptPwnd, described as the first confirmed real-world case of prompt injection compromising a CI/CD pipeline — a payload smuggled through a GitHub issue that reached into pipeline secrets.

This post is the audit that follows from those facts. CNCF published a threat model of the full delivery path in August 2026 — Matteo Bisi's "Shadow AI in CI/CD: Threat-modeling the path from developer laptop to Kubernetes" — and its core reframe is the right starting point: for platform teams, shadow AI is not a "developers using a chatbot" problem. It is an access problem. Ungoverned AI can reach source code, secrets, customer data, cloud environments, and deployment workflows, and once an AI system can call tools and take actions, it stops being productivity software and becomes a new non-human identity with permissions, a blast radius, and a place in your threat model.

The table below is the whole audit on one page. Everything after it unpacks each row for a team that runs its own build fleet — where owning the machines is an advantage, because every clamp below is something you can enforce rather than something you must ask a vendor for.

The six-stage audit, on one table​

Delivery stageTypical shadow AI usePrimary riskControl that clamps it
Developer laptopUnapproved code assistant, local model plugin, public chatbotSource, secrets, or architecture leave approved boundariesApproved tools + pre-commit secret scanning (gitleaks) + signed commits (Gitsign); VM-isolated agent workspaces
Source controlAI bot reviews PRs, generates commits, summarizes reposExcessive repo permissions, unsafe changes, no clear ownershipNamed owner, dedicated identity, minimal repo scope, short-lived credentials; mandatory human review of agent PRs
CI pipelineAI generates pipeline logic, analyzes logs, "fixes" failing buildsBuild secrets and cloud credentials exposed; automated supply-chain changesEphemeral, isolated runners; secrets out of prompts/logs; policy gates (Kyverno, OPA/Gatekeeper) before infra changes
Artifact registryAI-assisted image or dependency selectionVulnerable, malicious, or untraceable dependencies reach productionScan (Trivy, Grype), SBOM (Syft), sign (Cosign), provenance (in-toto), verify (Notation)
CD platformAI agent approves, modifies, or rolls out releasesBypassed change controls, unauthorized deploys, poor traceabilityAgent proposes, human approves via PR; GitOps controller (Argo CD, Flux) reconciles only what merged
Kubernetes runtimeAgent queries clusters, remediates alerts, scales workloadsOver-privileged ServiceAccounts, destructive actions, lateral movementSPIFFE/SPIRE identity, namespace-scoped RBAC, runtime detection (Falco, Tetragon), network policies

Why an agent with a token is a different threat class​

The risk rises sharply when AI stops giving advice and starts taking action. An assistant that suggests a snippet creates one class of risk: data leaving an approved boundary, or a subtly wrong suggestion trusted without review. An agent holding a Git token, cloud credentials, or a Kubernetes ServiceAccount creates a different class entirely. It can create, alter, or delete resources at machine speed, and Kubernetes will not distinguish between a harmful action taken by an attacker and the same action taken by an over-privileged automation identity.

That is why the useful questions are operational, not philosophical: which AI tools and agents are in use? What data do they receive? Which systems can they reach, and with what permissions? Who owns each agent's behavior? And can access be revoked immediately if an agent misbehaves? A workable model gives every agent a human owner, registers it as an identifiable workload, constrains it with least privilege, and monitors what it actually does.

Prompt injection is the through-line across every stage. Agents routinely read untrusted content — issue descriptions, READMEs, dependency changelogs, build logs — and any of it can steer an agent into disclosing data or taking an unsafe action. Prompt filtering alone will never be enough, which is why every row in the table pairs detection with a structural clamp the agent cannot talk its way past.

Auditing the six stages on machines you own​

1. Developer laptop. The path starts with a developer installing an AI coding extension or pasting an error log into a public service — a log containing an API token, an internal hostname, or a customer identifier. LayerX's 2025 enterprise report, built on browser telemetry rather than surveys, found 77% of employees paste data into generative-AI prompts, and 82% of those pastes come from unmanaged accounts.

The fix is not a blanket ban, which pushes usage further into the shadows. Provide approved tools so usage stays visible, back them with pre-commit secret scanning wired in as a git hook, and enforce signed commits — Gitsign's short-lived identity-based certificates make "who produced this change" auditable even when the author is an agent.

Where an agent genuinely needs to run commands autonomously, isolate rather than trust it: a disposable, VM-isolated workspace with its own kernel and filesystem, where the host's SSH keys, cloud credentials, and other checkouts simply are not present. The property matters more than the product — Firecracker microVMs, gVisor, Kata Containers, Sysbox — but the rule is absolute: an agent with broad autonomy should not run directly on the machine holding your credentials.

2. Source control. An unapproved AI bot connected to your Git host with broad permissions can read every repository, comment on PRs, create branches, or push code. The threat is not only leakage: if the integration token is stolen, or the bot is manipulated through a malicious issue or PR description — exactly the PromptPwnd shape — it can introduce unsafe changes or disclose repository content.

Every AI integration needs a named owner, its own identity, minimal repository scope, and short-lived credentials. An agent reviewing code for one team should not hold org-wide access. Treat agent-authored PRs like any other untrusted contributor: mandatory human review, no self-approval.

3. CI pipeline. CI systems hold some of your most powerful credentials: source-control tokens, registry credentials, cloud keys, signing keys. A shadow-AI capability that inspects build logs or autonomously "fixes" a failing build becomes an ungoverned privileged operator — and a prompt-injection payload hidden in source, a README, or a build log can change how it behaves.

This is the stage where owning your build fleet pays off most directly. Run AI-connected jobs on ephemeral runners destroyed after each job so nothing persists between builds. Keep long-lived secrets out of prompts, logs, and build environments. And put policy checks in front of any pipeline that can alter infrastructure or release software — in Kubernetes-native CI, an admission policy is a hard gate an agent cannot talk its way past. A Kyverno or OPA/Gatekeeper rule rejecting unsigned images holds regardless of what a compromised job tries to push.

4. Artifact registry and supply chain. AI-generated code can pull in insecure packages, weak configurations, or dependencies nobody reviewed. If you cannot establish what went into an image, who approved it, and whether it passed a gate, fast delivery is just unmanaged risk.

The practical open-source chain is well settled: Trivy or Grype to scan images and infrastructure code, Syft for a per-artifact SBOM, Cosign to sign images and attach attestations, in-toto for signed per-step attestations as the foundation of SLSA-style provenance, and Notation to verify signatures at the registry boundary. Hold AI-generated code to the same review and release bar as human-written code — no fast lane for machine output.

5. Continuous delivery. An AI release agent able to edit Helm charts, update GitOps manifests, or trigger rollbacks can bypass exactly the change-management controls a mature process depends on. Draw a hard line between an agent that recommends a release action and one that executes it: production deploys, privilege escalation, data export, deletion, and network-policy changes sit behind explicit approval gates with strong audit trails.

In a GitOps model the pull request is the approval gate — the agent proposes a manifest change, a human approves the merge, and Argo CD or Flux reconciles. Neither controller takes instructions from the agent; it only ever applies what is already merged. Enforce that no path to production skips the gate.

6. Kubernetes runtime. Often the most consequential stage. A remediation agent gets cluster-admin "temporarily" to investigate an alert, and a convenience tool becomes a high-value target. A compromised agent with broad permissions can enumerate secrets, deploy a malicious workload, alter network policies, or exfiltrate data.

The clamps are namespace boundaries, workload identity, minimal RBAC, admission control, runtime detection, and network segmentation. Concretely: SPIFFE/SPIRE so each workload, agents included, gets a cryptographic identity instead of a shared long-lived token; least-privilege RBAC scoped to a namespace and verb set, never cluster-admin; Falco or Tetragon for runtime detection of post-compromise behavior; and network policies to contain lateral movement.

An agent whose only job is restarting a Deployment needs a namespaced Role granting only get, list, and patch on deployments — a rollout restart is a patch, so create and delete are never needed. Even fully compromised, that agent cannot read secrets, touch other namespaces, or delete workloads.

The bridge is already proven: four incidents​

This threat model is not hypothetical. The receipts arrived before the guidance did.

  • PromptPwnd (December 2025). Aikido Security's disclosure showed a prompt-injection payload delivered through a GitHub issue reaching an AI agent wired into CI, exposing pipeline secrets — the first confirmed real-world case of prompt injection compromising a CI/CD pipeline.
  • The Nx compromise (August 2025). Attackers breached the Nx build system through a vulnerable GitHub Actions workflow, in a campaign that explicitly targeted AI coding-tool credentials alongside traditional developer secrets. Your AI assistant's API key is now as lootable as your npm token.
  • SANDWORM_MODE (February 2026). A family of 19 typosquatted npm packages, documented in a Cloud Security Alliance research note, deployed a rogue MCP server into hidden directories after a 48-hour delay, registering deceptively named tools that instructed AI assistants to silently extract SSH keys, AWS credentials, npm tokens, and environment variables — and to hide the activity from the user.
  • The 80% AI-review bypass. Research published on arXiv tested a five-agent CI/CD pipeline built specifically to enforce pipeline security, running five distinct LLMs across three providers — and a single cleverly worded external request pushed malicious code all the way to deployment, bypassing every automated check. If your review gate is itself an LLM, assume it is bypassable and keep the structural gates behind it.

Add Google's September 2026 finding that the threat actor UNC6780 used prompt injection against AI coding assistants and LLM-based security scanners among half a dozen supply-chain techniques, and the pattern is unmistakable: attackers already treat your AI tooling as part of your attack surface. Your threat model should too.

Controls that scale beyond one pipeline​

None of this is about slowing AI adoption. The point is making AI use visible, accountable, and proportionate to its risk — and four controls scale that posture across the fleet.

Build an AI inventory. Keep a living register of approved and discovered tools, models, extensions, agents, API integrations, and MCP servers. Each entry needs an owner, its purpose and users, the data it may touch, the systems it reaches and with what permissions, the model behind it, and how to revoke its access. You cannot govern what you cannot see; every other control assumes this one exists.

Treat agents as identities. Every agent gets a unique identity and never borrows a developer's credentials or a shared admin account. Least privilege by default means separate credentials per environment, short-lived tokens, read-only access where possible, namespace-scoped Kubernetes permissions, explicit allowlists for tools and repos, and immediate revocation. In-cluster, SPIFFE/SPIRE and cert-manager already do this job; outside the cluster, Keycloak is the usual anchor for service identities.

Match the control to the blast radius. Not every AI function needs the same treatment. Code explanation needs approved tools and data-classification rules; code suggestion adds human review, testing, secret scanning, and dependency checks; PR creation needs restricted repo permissions and mandatory peer review; pipeline modification needs isolated execution, policy-as-code checks, and an approval gate; cluster remediation needs strict namespace RBAC, limited scope, a full audit trail, and human approval for high-impact actions; autonomous production changes should require exceptional approval, time-bound access, a kill switch, and continuous monitoring.

Apply defense in depth. Governance, identity and access management, source-control protections, secrets management, secure CI/CD, dependency and container scanning, Kubernetes admission and runtime policy, network controls, centralized logging. The test is simple: if any single control fails, the agent must still be unable to cause material harm.

The honest gap​

Bluntness is owed about what is not solved. The pipeline and runtime layers are well served by mature CNCF projects, but AI governance itself — prompt inspection, agent discovery, tool-call policy — is thinner. No graduated project owns that space end to end, and the gateway layer only improved noticeably during 2026: kagent for running agents under Kubernetes' declarative model, the Envoy AI Gateway reaching v1.0 in June 2026, agentgateway as a policy proxy for MCP and agent-to-agent traffic, and agentregistry targeting the inventory problem while still early-stage.

One layer has no foundation-governed answer at all: scanning MCP servers for tool poisoning and prompt-injection payloads. The maintained options are vendor-developed, so treat that as an evaluation exercise rather than a settled choice. Until it matures, pair OWASP's agent guidance with a policy enforcement point in front of model and tool endpoints, and lean on the identity and admission controls above to constrain what an agent can do, even when you cannot fully inspect what it is thinking.

Governed identities or an untracked pile​

Shadow AI is the next evolution of shadow IT, with one critical difference: modern AI agents can interpret information, call tools, and act across engineering systems. That makes them powerful accelerators and potential high-speed paths to source-code leakage, credential compromise, supply-chain abuse, and production disruption. The decision is not whether your developers will use AI — with nearly every organization reporting unsanctioned use, they already are. The decision is whether you manage AI as an untracked pile of productivity tools or as a governed set of identities, data flows, and production-capable workloads.

For a team that owns its build fleet, the path forward is concrete and almost entirely open source: discover AI use, assign ownership, scope permissions with real workload identity, protect the supply chain with signing and provenance, enforce Kubernetes boundaries with admission and runtime policy, segment the network, monitor behavior, and keep a human in the loop for consequential actions. Let agents help teams move faster — but never let them operate beyond your ability to see, control, and stop them.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own, with AI agents as first-class operators. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide