Every container image you have ever pulled from registry.k8s.io got there through a single tool: kpromo, the Kubernetes image promoter. It copies images from staging registries to production, signs them with cosign, replicates signatures across more than 20 regional mirrors, and generates SLSA provenance attestations. If it breaks, no Kubernetes release ships.
In early 2026, SIG Release rewrote its core from scratch — and the scoreboard, reported by Sascha Grunert (Red Hat), is startling: over 16,000 lines deleted against 10,000 added, a net reduction of about 5,000 lines, leaving a codebase 20% smaller that does strictly more. The plan phase dropped from about 20 minutes to about 2, and signature replication across the mirror fleet fell from about 17 hours to about 15 minutes. Nineteen long-standing issues closed, and nobody noticed — which was the whole point.
This post is a close read of that rewrite for a different audience than the one it was written for: teams running their own registry story. If you operate a self-hosted PaaS that promotes tenant buildpack images from a staging registry to production on owned hardware, you are running a smaller version of exactly this pipeline. The copyable lesson fits in one sentence: split promotion into sequential phases that each own the full rate-limit budget, verify before you copy, sign after you promote, and push replication out of the critical path. The rest of this post shows the working — the failure modes that forced the rewrite, the seven phases, the three performance tricks with their numbers, the rollout discipline — and then maps each piece onto what a git-push platform should copy and what it should skip.
What kpromo actually does
Strip away the scale and kpromo's job is the same job every image pipeline has: decide what moves, move it, and prove it. Concretely, it reads GitOps-style YAML manifests from the kubernetes/k8s.io repo that declare which staging images should be promoted, copies those images server-side from staging registries to production while preserving digests, signs the promoted images with keyless cosign, replicates the signatures to every regional mirror, and attaches provenance records. The workflow itself dates to late 2018, when the promoter began as an internal Google project to replace a manual, Googler-gated copy process into k8s.gcr.io with a community-owned flow: push to staging, open a PR with a manifest, get review, let automation handle the rest.
Seven years of that automation working is precisely what made the rewrite necessary. By 2025 the codebase carried contributions from multiple SIGs and subprojects — cip, gh2gcs, krel promote-images, and promobot-files consolidated into one CLI, plus cosign signing, SBOM support, and vulnerability scanning bolted on by different authors. Forty-two contributors, about 3,500 commits, more than 60 releases.
It worked, but the README said the quiet part out loud: expect duplicated code, multiple techniques for the same thing, and several TODOs. Any platform team that has watched its own deploy pipeline accrete flags and special cases for years will recognize the shape. The difference is only that kpromo's tech debt had a production symptom with a number attached: core-image promotion jobs regularly took over 30 minutes and frequently failed on rate limits.
Why the monolith had to die
Three failure modes did the persuading, and each one maps to a lesson.
The first was rate-limit contention. Signing and signature replication ran interleaved inside one large promotion function, both hammering registry APIs against a shared quota. The contention between those two workloads caused most production failures — two subsystems that each behaved fine in isolation starved each other when fused.
The second was latency without a floor: the plan phase read 1,350 registries serially, and one stalled connection was once observed blocking the entire pipeline for over 9 hours, because no individual network operation had its own timeout. The third was extensibility — adding provenance verification or vulnerability scanning to the monolith was painful enough that two SIG Release roadmap items sat open while the team answered eight research spikes before daring to start.
Note the ordering there. The rewrite was not motivated by elegance; it was motivated by a 30-minute job that failed on quotas, a 9-hour hang with no timeout to blame, and features that could not be added. That is the correct trigger for a pipeline rewrite, and it is worth stating because the temptation runs the other way: rewriting a working pipeline for cleanliness while its failure modes are still unmeasured. SIG Release measured first, then cut.
The seven phases
The new pipeline runs promotion as seven sequential phases — Setup, Plan, Provenance, Validate, Promote, Sign, Attest:
| Phase | What it does |
|---|---|
| Setup | Validate options, prewarm the TUF cache. |
| Plan | Parse manifests, read registries, compute which images need promotion. |
| Provenance | Verify SLSA attestations on staging images. |
| Validate | Check cosign signatures; dry runs exit here. |
| Promote | Copy images server-side, preserving digests. |
| Sign | Sign promoted images with keyless cosign. |
| Attest | Generate promotion records using a dedicated in-toto predicate type. |
Two design decisions in that table do most of the work. First, phases run sequentially so each one gets exclusive access to the full rate-limit budget — the contention that caused most production failures is gone by construction, not by tuning. Second, signature replication to the mirrors is no longer part of this pipeline at all; it runs as a dedicated periodic Prow job. Signing (one write per image) and replication (one fan-out per image per mirror) have completely different cost shapes, and the rewrite's single most important act was refusing to schedule them together.
Do not confuse these seven runtime phases with the nine rollout phases the team used to ship the rewrite itself — rate limiting, interfaces, pipeline engine, provenance, scanner/SBOMs, the sign/replicate split, then three phases of pure deletion. The nine were a shipping strategy; the seven are the architecture. Both are worth copying, but they answer different questions, and the rollout half gets its own section below.
Three performance tricks worth stealing
With the architecture in place, the team turned to speed. Each win below comes with its mechanism, its number, and the condition under which it applies — because a trick without its applicability condition is how a 68x headline becomes a cargo-culted refactor.
Parallel registry reads: 20 minutes to 2 minutes. The plan phase reads 1,350 registries; parallelizing those reads cut it by roughly 10x (promo-tools#1736). Applies when your plan step fans out across many independent registries or repositories. It does nothing for a PaaS whose staging-to-prod map fits in one registry with a handful of repositories — there the plan phase is already seconds, and parallelism just adds goroutines to debug.
Source check before fan-out: 17 hours to 15 minutes. Before iterating all mirrors for an image, the replicator now checks whether the signature exists on the primary registry first; in steady state, where most signatures are already replicated, that single check skips the entire fan-out (promo-tools#1727). This is the biggest number in the post — roughly 68x — and the most conditional: it pays off exactly in proportion to how often the work is already done. Any reconciliation loop with a cheap "is this already true?" probe and an expensive repair path wants this shape.
Two-phase tag listing: roughly half the API calls. Instead of checking all 46,000 image groups across 20-plus mirrors, the pipeline first checks only the source repositories — and about 57% of images have no signatures at all because they were promoted before signing existed, so they are skipped entirely (promo-tools#1761). The general form: enumerate cheaply, filter on the cheapest predicate first, and never pay mirror-fan-out prices for work the source already tells you to skip. Rounding out the set are per-request timeouts with retries (the 9-hour hang class, gone), HTTP connection reuse across operations (which closed a request open since 2023), and local-registry integration tests so the pipeline can be exercised without touching production infrastructure.
How to rewrite infrastructure nobody may see break
The rollout is the part most worth studying, because "delete the monolith" is easy to say and the k8s.io promoter is load-bearing for every Kubernetes release. The team shipped the rewrite as nine independently reviewable, mergeable, validatable phases behind a tracking issue, cut v4.2.0 with the new engine alongside the old code, and let it soak in production before deleting anything. Only then did v4.3.0 remove the legacy path entirely, and v4.4.0 shipped the follow-up improvements with provenance generation and verification on by default. Over 40 PRs, three releases, a clear rollback path at every step — never needed, but present.
The hard requirement throughout was zero user-facing change: same kpromo cip flags, same YAML manifests, same Prow job, no workflow updates for anyone downstream. And the soak caught real bugs: one regression made every image appear "lost" so nothing was promoted; another set the default thread count to zero and blocked all goroutines. Both were fixed within hours because the blast radius of each phase was small and the old code path was still there to compare against.
"Nobody noticed" was not luck. It was the output of a rollout designed so that noticing was never required.
Your PaaS's registry story: copy this, skip that
Here is the explicit mapping for a self-hosted git-push PaaS promoting tenant images from staging to production on owned machines.
Copy the phase split. Even with two registries and a dozen tenants, separating plan, verify, promote, sign, and attest into distinct stages — each independently runnable, dry-runnable, and retryable — is what turns "the deploy pipeline failed" into a diagnosable sentence. The Validate phase's contract is the detail to steal first: a dry run that exits after verification, before anything is copied or signed.
Copy the sequential rate budget. You may never hit a registry quota the way 1,350 registries do, but the principle survives at any scale: stages with different cost shapes should not share a concurrency pool implicitly. If your pipeline builds, pushes, signs, and notifies in one concurrent blur, the kpromo failure is your future at higher tenant counts. Give each stage its own budget on purpose.
Copy sign-after-promote and verify-before-copy. Verify SLSA provenance on what staging holds, copy by digest, then sign the production digest with keyless cosign and attach a promotion record. Keyless signing matters operationally: no long-lived private key to rotate, store, or leak from CI — the OIDC identity of the workflow is the identity on the signature, logged in Rekor. A tenant asking "prove this image came from my repo's CI run" gets a cryptographic answer instead of a dashboard screenshot.
Skip the 20-mirror replication mesh. This is the part of kpromo's story you should admire and not reproduce. Upstream itself wants out: issue #1762 proposes eliminating signature replication entirely by having archeio, the registry.k8s.io redirect service, serve signature requests from a single canonical upstream (the archeio codebase already carries a SIGNATURE_UPSTREAM_ENDPOINT for exactly this). For a PaaS on owned hardware, the equivalent is one canonical registry plus pull-through caches or a redirect at the edge — not N-way signature fan-out on every deploy.
Defer what your scale does not need. Parallel plan reads, two-phase tag listing, and the source-check shortcut are solutions to thousands-of-registries problems. With one staging registry and one production registry, your plan phase is a manifest diff that takes seconds. Implement the phase boundaries now — they are cheap — and let the optimizations arrive when a measured job, not an imagined one, demands them.
That, more than any single trick, is the discipline the rewrite actually demonstrates: SIG Release measured a 30-minute failing job and a 9-hour hang, then cut. Measure yours, then cut.
Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.



