Skip to main content

Stop Guessing When to Renew: cert-manager 1.21 Lets Let's Encrypt Tell Your Fleet When

9 min readDora NodaDora Noda
Share
On this page

On February 29, 2020, Let's Encrypt discovered a bug in its CAA rechecking code — and then had to tell the holders of 3,048,289 certificates, 2.6 percent of everything active on the web's largest CA, to renew within days or be revoked. About a million of those were duplicates covering the same domain names, which helped. What didn't help was that every ACME client on the planet decided when to renew by staring at its own certificates and guessing. There was no channel for the CA to say "renew this one now, that one can wait." The scramble lasted five days, and Let's Encrypt ultimately chose not to revoke roughly a million affected certificates past the deadline because doing so would have broken too much of the web.

That incident is the reason ACME Renewal Information exists. ARI — now RFC 9773 — flips renewal scheduling around: instead of each client guessing a renewal time from local state, the client asks the CA when it should renew, and the CA answers with a suggested window per certificate. Let's Encrypt has served ARI in both staging and production since March 2023. Tailscale adopted it in May 2024. And on July 8, 2026, cert-manager 1.21 brought experimental ARI support to the fleet manager most Kubernetes platforms actually run, behind the ACMEUseARI feature gate.

The timing is not a coincidence. The CA/Browser Forum's Ballot SC-081v3, approved in April 2025, phases maximum public TLS validity from 398 days down to 200 days in March 2026, 100 days in March 2027, and 47 days by March 2029. Every step of that staircase makes fixed client-side renewal heuristics worse: shorter lifetimes mean more renewals per year, tighter windows, and less slack when a CA-side event demands early renewal. A platform issuing automated TLS for hundreds of tenant custom domains is exactly where guessing breaks first. Here is what the CA-suggested window changes, and how to upgrade past 1.21's breaking changes without stranding your Issuers.

What ARI actually changes: from local guess to CA-suggested window​

The historical renewal rule is simple and everyone runs some version of it: renew at two-thirds of the certificate's lifetime (cert-manager's default is the earlier of renewBefore and renewBeforePercentage, which defaults to two-thirds). It is simple, stateless, and blind. It cannot hear a CA-side mass revocation coming, and when hundreds of certificates share an issuance date — a new tenant batch, a fleet rebuild — it renews them all in one spike. That spike lands on the CA at the worst possible moment and can trip per-account rate limits, including Let's Encrypt's New Orders limit, right when you need renewals to succeed.

ARI replaces the guess with a question. For each ACME-managed certificate, the client does an unauthenticated GET against the renewalInfo URL the CA advertises in its directory object, with a certID derived from the certificate's authority key identifier and serial number. The response is shaped like this:

json
{
  "suggestedWindow": {
    "start": "2026-10-02T00:00:00Z",
    "end": "2026-10-04T00:00:00Z"
  }
}

The client picks a random moment inside that window and renews then. A Retry-After header hints when to poll again, so the CA can move windows later — pulling renewals forward ahead of a mass revocation, or spreading them out to flatten a spike it sees forming. Two properties fall out of this design that no local heuristic can match. First, the CA can stagger suggested windows across certificates, so a fleet that issued everything on the same day does not renew everything in the same hour. Second, the CA gets a voice in an emergency: "renew before Friday" becomes a machine-readable signal instead of a forum post and a prayer.

The server side is not theoretical. Let's Encrypt enabled ARI in production in March 2023 and published a client-integration guide in April 2024; Tailscale wrote up its adoption in May 2024, and the spec has since been published as RFC 9773. cert-manager 1.21 is the client catching up: with the gate enabled, the controller queries the ACME server's renewalInfo endpoint for the recommended window per certificate, so a CA can proactively prompt renewal during mass revocations or key rollovers. Enabling it is one controller flag:

bash
--feature-gates=ACMEUseARI=true

(Helm chart users can pass it through the controller's extra arguments; canary it on one issuer before enabling it fleet-wide — more on that below.)

To make the herd behavior concrete, consider a representative fleet: 200 tenant custom-domain certificates, all issued the same day during a platform migration, each with a 90-day lifetime. Under renew-at-two-thirds, all 200 become due within the same roughly 24-hour band around day 60 — a single spike of 200 new orders against the CA, plus whatever retries a rate-limit rejection triggers. With ARI, the CA hands each certificate its own suggested window spread across days, and each client randomizes inside its window: the same 200 renewals arrive as a low plateau instead of a spike. Now run the sensitivity the SC-081v3 staircase demands. At 47-day lifetimes, the two-thirds rule fires around day 31, the renewal cadence roughly doubles to nearly eight rotations a year per certificate, and the spike that used to arrive every two months arrives every month — with half the slack to absorb a CA-side early-renewal event. ARI's spreading matters more, not less, as lifetimes shrink; fixed heuristics degrade on exactly the schedule the Forum voted for.

The rest of 1.21 a fleet operator actually needs​

ARI is the headline, but two neighboring changes matter for the same audience. First, 1.21 adds per-certificate renewal policies: a new renewal field on the Certificate API that complements renewBefore and renewBeforePercentage, including scheduling renewal into approved maintenance windows and disabling automatic renewal entirely. Maintenance-window renewal is the quiet win here — a platform team can finally say "renew tenant certificates only inside our Tuesday window" declaratively instead of scripting around the controller.

Second, the caveat that sets your minimum version: 1.21.0 panicked the controller on any Certificate with spec.renewal.policy: Disabled, and v1.21.1 fixed it three weeks later alongside Issuers stuck at Ready=False (InvalidSolver) when a DNS-01 solver Secret was created after the Issuer, plus a log-spam and dropped-informer-events regression. Do not deploy 1.21.0. The floor is v1.21.1, and v1.21.2 exists with further ACME and scheduler fixes if you are upgrading today.

Two more features deserve one line each. waitInsteadOfSelfCheck lets an ACME solver skip cert-manager's own self-check and wait a configured duration before asking the CA to validate — an escape hatch for split-horizon DNS and NAT hairpin environments where the self-check can never pass. And the Vault issuer now supports AWS IAM authentication (IRSA, EKS Pod Identity, ambient EC2/ECS credentials), removing long-lived AWS secrets from one more place.

The safe upgrade order​

1.21 ships three breaking changes, all in Helm chart RBAC and metrics values. The order matters because getting it wrong strands issuance in ways that look like CA problems but are really local RBAC. Work through them before you touch the ARI gate.

1. Restore or replace the removed tokenrequest RBAC. The chart no longer creates the default Role and RoleBinding granting the controller serviceaccounts/token: create on its own ServiceAccount. No documented workflow needs it — the Route53 docs section that motivated it was removed back in 2024 — but if you use serviceAccountRef.name pointing at the controller ServiceAccount, you must now create your own Role/RoleBinding or, better, migrate to a dedicated ServiceAccount following the Vault or Route53 documentation. Audit this first: a missing token-creation permission fails issuance silently at the solver step.

2. Check tooling that creates Challenge or Order resources directly. The cert-manager-edit aggregate ClusterRole no longer grants create on challenges.acme.cert-manager.io or create/patch/update on orders.acme.cert-manager.io (GHSA-8rvj-mm4h-c258; these are internal to the ACME workflow, and Challenge patch/update are retained so you can still clear stuck finalizers). Note the mitigation already in your fleet: this change shipped earlier in v1.20.3 and v1.19.6, so if you are current on either line it is not a new break. Only bespoke tooling that mints Challenge or Order objects needs new explicit grants.

3. Strip the removed metrics Helm keys. prometheus.servicemonitor.targetPort, prometheus.servicemonitor.path, and prometheus.podmonitor.path are gone, and the controller Service metrics port is renamed from tcp-prometheus-servicemonitor to http-metrics. Because the values schema sets additionalProperties: false, leaving any removed key in your overrides fails the upgrade with a schema validation error — remove them before you run the upgrade, and update any ServiceMonitor selectors chasing the old port name.

Then: upgrade to at least v1.21.1, confirm Issuers and Certificates return to Ready, and only then enable ACMEUseARI — on a canary issuer first, watching renewal timing and renewalInfo fetch behavior, before rolling it across every tenant issuer.

Why the feature gate is the right posture for now​

Experimental-on-the-renewal-hot-path deserves respect. Renewal timing is the one cert-manager behavior where a bug does not page you today — it pages you sixty days from now as an expiry, possibly across hundreds of certificates at once. ARI adds a network dependency (the renewalInfo poll) and CA-driven timing into that path, and the poll cadence and Retry-After handling are the parts of a new client implementation most likely to still be settling. The gate lets the ecosystem shake this down without making every 1.21 upgrader a beta tester by default.

The practical posture: enable the gate on a non-production issuer, and watch three things before going fleet-wide — renewal-time skew (are renewals actually spreading across suggested windows rather than clustering?), renewalInfo fetch errors and fallback behavior when the endpoint is unreachable, and interaction with any renewal maintenance-window policies you set, since CA-suggested windows and local maintenance windows now both constrain timing. What you are buying with patience is the end of renewal guessing on a CA that has served the signal since 2023. That is worth a canary.

The end of renewal guessing​

Step back and the arc is clear. The 2020 revocation scramble proved clients need a signal from the CA. Let's Encrypt built and production-hardened that signal starting in 2023, the IETF standardized it as RFC 9773, and cert-manager 1.21 finally wires it into the controller most Kubernetes fleets run — just as the SC-081v3 staircase starts punishing fixed renewal heuristics on a fixed calendar. Renew-at-two-thirds served a 398-day world adequately. It does not serve a 47-day world at all.

For a self-hosted platform issuing TLS per tenant app, this is one of those rare upgrades that is both obviously correct and easy to stage: fix the RBAC breaks in order, land on v1.21.1 or later, canary the gate, and let the CA pick the renewal times from then on. Your future mass-revocation self will thank you.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex