Skip to main content

Your TLS Renewals Have an Expiration Date Too: Auditing cert-manager for Let's Encrypt's 45-Day Countdown

9 min readDora NodaDora Noda
Share
On this page

Let's Encrypt has put your renewal timers on a published schedule — and the first deadline already passed. On May 13, 2026, the opt-in tlsserver ACME profile started issuing 45-day certificates, and on February 10, 2027 — roughly four months from now — the default classic profile drops from 90 days to 64, then to 45 days in February 2028. Let's Encrypt's own announcement names the failure mode explicitly: clients renewing at a hardcoded 60-day interval will break, and the fix is to renew at about two-thirds of the way through each certificate's actual lifetime, preferably by asking the CA when to renew via ACME Renewal Information (ARI).

If you run cert-manager on a self-hosted platform, this post gives you a four-step audit you can run today, the renewal math that explains why fixed intervals break, and the one feature gate that moves you from guessing to asking.

The 4-step cert-manager audit​

Run these in order. Each has a pass/fail bar, and the sections below explain the why behind each one.

Step 1 — Check your controller version. ARI support exists only in cert-manager v1.21.0 (released July 8, 2026) and later, and even there it sits behind an experimental feature gate.

bash
kubectl get deploy cert-manager -n cert-manager \
  -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'

Pass: image tag is v1.21.0 or newer. Fail: anything older means ARI is not even available to you yet, so every renewal decision your cluster makes is a local guess.

Step 2 — Find every hardcoded renewal window. cert-manager's safe default is to renew at two-thirds of a certificate's lifetime. That default only protects you if nobody overrode it. List every Certificate with an explicit duration or renewBefore:

bash
kubectl get certificates -A -o json | jq -r '
  .items[]
  | select(.spec.renewBefore != null or .spec.duration != null)
  | "\(.metadata.namespace)/\(.metadata.name) duration=\(.spec.duration // "default") renewBefore=\(.spec.renewBefore // "default")"'

Pass: empty output (everything on defaults), or every explicit renewBefore is small relative to a 45-day lifetime. Fail: any renewBefore of 720h (30 days) or more — on a 45-day certificate that fires renewals at day 15, tripling your ACME load — and anything at or above the issued lifetime, which cert-manager's own docs warn sends the Certificate into a permanent renewal loop.

Step 3 — Check whether ARI is actually on. Version 1.21 ships the gate; it does not flip it. Inspect the controller's arguments:

bash
kubectl get deploy cert-manager -n cert-manager \
  -o jsonpath='{.spec.template.spec.containers[0].args}' \
  | tr ' ' '\n' | grep -i feature-gates

Pass: the output includes ACMEUseARI=true. Fail: the gate is absent or false, meaning your cluster never consults the CA's suggested renewal window even when the CA is trying to tell it something urgent, like a mass revocation. To enable it, add --feature-gates=ACMEUseARI=true to the controller's arguments (through your Helm values' feature-gate mechanism) and roll the deployment — on a staging cluster first, since the support is still experimental.

Step 4 — Confirm expiry alerting exists outside cert-manager. Automation you cannot see failing is a hope, not a system. At minimum, alert on certificates approaching expiry and on stuck ACME challenges:

bash
kubectl get certificates -A | grep -v True
kubectl get challenges -A --field-selector status.phase!=valid 2>/dev/null | head

Pass: the first command shows nothing but healthy True rows, and you have a Prometheus alert resembling certmanager_certificate_expiration_timestamp_seconds - time() < 7 * 86400 paging someone. Fail: you are relying on tenants to report browser warnings. Let's Encrypt's announcement recommends exactly this monitoring step alongside ARI.

If all four pass, your cluster survives the countdown on defaults plus ARI. If any fail, the rest of this post tells you what to change and in what order.

The timeline in one table​

Two schedules matter: Let's Encrypt's profile rollout (what your ACME client actually receives) and the CA/Browser Forum backstop (the industry ceiling every public CA must meet).

DateLet's EncryptCA/B Forum maximum (ballot SC-081v3)
Mar 15, 2026—200 days
May 13, 2026tlsserver profile → 45-day certs (opt-in); shortlived (~6-day) already GA since Jan 15—
Feb 10, 2027Default classic → 64-day certs, 10-day authz reuse—
Mar 15, 2027—100 days
Feb 16, 2028Default classic → 45-day certs, 7-hour authz reuse—
Mar 15, 2029—47 days

Two columns deserve emphasis. First, the authorization-reuse shrink from 30 days to 10 days to 7 hours means domain control gets re-validated far more often — HTTP-01 and DNS-01 challenges your automation runs occasionally today become a weekly rhythm, with DNS propagation waits and rate limits to budget for. (Let's Encrypt is standardizing DNS-PERSIST-01, a validation method whose TXT record stays put across renewals, as relief; it was expected during 2026.) Second, the Forum's 47-day ceiling in 2029 means this is not a Let's Encrypt quirk you can dodge by switching CAs — every public CA lands in the same place.

Why fixed intervals break: the renewal math​

cert-manager computes renewalTime = notAfter − renewBefore, where the default renewBefore is one-third of the issued lifetime — the two-thirds rule Let's Encrypt recommends. That default tracks shrinking lifetimes automatically. Hardcoded values do not. Here is what four common configurations do as lifetimes shrink:

Config90-day cert64-day cert45-day cert6-day (shortlived)
Default (2/3 of lifetime)renews day 60renews day ~43renews day 30renews day 4
renewBefore: 720h (30d)renews day 60renews day 34renews day 15 (3× ACME load)renewal loop
60-day cron / hardcoded intervalOK (30d margin)renews day 60 of 64 (4d margin — one failed attempt from outage)expires before renewal (breaks)breaks immediately
renewBefore ≥ issued lifetimelooplooplooploop

The middle two rows are where real outages hide. A 60-day cron that has worked for years keeps working right up until the February 2027 switch to 64-day certificates, then survives on a four-day margin — a single DNS-01 propagation hiccup or a rate-limit pause away from serving an expired certificate. And renewBefore: 720h, a value copied from countless tutorials written for 90-day certificates, silently triples issuance traffic at 45 days instead of breaking loudly, which is arguably worse: it looks healthy while hammering the CA and your challenge solvers.

The honest summary: defaults survive every row of this table, hardcoded values survive only the rows they were tuned for, and the table keeps adding rows.

What ARI actually does​

ACME Renewal Information, standardized as RFC 9773, inverts the renewal decision: instead of the client guessing from the calendar, the client asks the CA. The flow is simple — the client sends the certificate's ID (issuer key hash plus serial) to the CA's renewalInfo endpoint, and the CA answers with a suggested window (start/end) plus a Retry-After telling the client when to ask again. The client picks a random moment inside the window and re-checks as instructed.

That inversion buys three things fixed math cannot. First, the CA can stagger millions of clients across time instead of absorbing a thundering herd the day a popular lifetime fraction lands. Second, the CA can move the window to "right now" during a mass revocation or key rollover, turning every ARI-aware client into a self-healing one without a single operator reading a security advisory first. Third, clients stop encoding lifetime assumptions entirely — when the profile switches from 64 to 45 days, nothing client-side needs a new number.

On the cert-manager side, that is precisely what the ACMEUseARI gate enables: the controller queries the renewalInfo endpoint and folds the CA's suggested window into its renewal scheduling. It shipped as experimental in v1.21.0, which is why Step 1 gates on version and Step 3 gates on the flag separately — having the capability and using it are two different audits. Test against Let's Encrypt's staging environment (which gets each lifetime change about a month before production) before trusting it with tenant traffic.

Don't forget the chain moved too​

The same May 13, 2026 date that flipped the tlsserver profile also activated Generation Y, Let's Encrypt's new issuance hierarchy: two new roots and six new intermediates, cross-signed from the Generation X roots so existing trust stores keep working. Most servers noticed nothing — but the chain is not identical, and three differences are worth a checklist of their own:

  • P-384 ECDSA intermediates. The new ECDSA chain uses the P-384 curve, which broke TLS clients compiled without it — most visibly ESP32 devices running Mbed TLS without secp384r1 support when their MQTT broker's certificate renewed onto the YE1 intermediate in late July 2026. If you serve embedded, IoT, or ancient clients, re-test their trust paths against a Gen Y chain deliberately.
  • No more client-auth certificates. The tlsclient profile was retired on July 8, 2026, ending Let's Encrypt issuance for the client-auth extended key usage; Generation X retired with it. If any mTLS setup in your fleet uses Let's Encrypt client certificates, those renewals are already failing — move them to a private CA or a provider that still offers the EKU.
  • Cross-sign expiry horizon. Cross-signs are a compatibility bridge, not a permanent floor. Track ISRG's published expiry dates for the X1/X2 cross-signs the way you track your own renewals, because the day a cross-sign lapses is the day un-updated trust stores find out.

None of this changes the renewal audit above — but "we automate TLS for you" quietly depends on both the timers and the chain, and both moved this year.

The claim gets harder to keep​

Every tenant-facing platform makes some version of the promise: push code, get HTTPS, never think about certificates. That promise has always rested on renewal automation tuned for 90-day certificates with 30-day authorization reuse. By February 2028 both numbers are roughly an order of magnitude tighter, and the configurations that break — 60-day crons, tutorial-copy renewBefore values, monitoring that assumes renewals rarely fail — break silently right up until the outage.

The good news is the fix is small and entirely on your side of the ACME connection: run the four-step audit, delete the hardcoded windows, enable ARI, and alert on expiry from outside the renewal path. Do it before February 2027, while the margin for a misconfigured retry is still measured in weeks rather than hours.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex