The Healthy Pod That Had Stopped Working
I sat down to do some cluster work and kubectl just… hung. No error, no timeout message worth reading, just a prompt that never came back. That’s how I found out my management cluster had stopped renewing TLS certificates three weeks earlier.
Nothing had crashed. Nothing had alerted. The pod responsible was Running, its readiness probe was passing, and it had done no useful work since August 7th.
Pulling the thread backwards
Here’s the chain, in the order I untangled it rather than the order it happened.
kubectl from my workstation doesn’t talk to the production cluster directly — it routes through the Rancher proxy on my management box, ubuntu-hp01. So a hang there means the Rancher path is broken, not the cluster. And the Rancher path was broken because the cattle-cluster-agent running in the production cluster had been failing TLS verification against Rancher for nine straight days. It couldn’t establish its tunnel, so the proxy had nothing to proxy.
It was failing TLS verification because three certificates had expired across August 21st and 22nd: Rancher’s own, and the wildcard for my internal domain twice over, since it exists in two namespaces. All of them things cert-manager was supposed to renew automatically and hadn’t.
And cert-manager hadn’t renewed them because it had been wedged since August 7th. As best I can reconstruct it, its informer watches — the long-lived connections it uses to notice that a Certificate resource needs attention — died after the API server throttled it, and cert-manager never re-established them. This is the part I keep turning over: a controller whose entire job is reacting to changes had lost its ability to see changes, and there is no health check in the world that catches that from the inside. The process was alive. The HTTP endpoint answered. The pod was green. It just wasn’t watching anything anymore.
Two weeks of silence, then certificates expired on schedule, then nine days of a degraded control plane, and the first symptom that reached me was a hanging shell command.
Why nobody told me
I do have an alert for this. CertificateExpiringSoon has been in my Prometheus rules for a long time, and it works — it scrapes cert-manager’s metrics and fires when a certificate is inside its renewal window.
It runs on the production cluster’s Prometheus. cert-manager was wedged on the management cluster. The alert had no visibility into the metrics it needed, so it did exactly what it was configured to do: nothing, quietly, correctly, for three weeks.
I wrote a post a few months ago about backups that don’t share fate — the idea that a backup living on the disk it’s meant to protect isn’t really a backup. This is the same mistake wearing a different hat. My certificate alert was derived from cert-manager’s own view of the world. When cert-manager’s view of the world stopped updating, the alert inherited the blindness. A monitor that gets its truth from the component it’s monitoring will always agree with that component, including when the component is wrong.
The other thing worth naming: for a while there I had no way to reach the production cluster at all, which made the outage feel much worse than it was. The escape hatch turned out to be simple — SSH to a control-plane node and use the local RKE2 kubeconfig directly. Obvious in hindsight. Worth knowing before you need it.
Unwedging it
Restarting cert-manager brought the watches back, and it immediately tried to reissue all three certificates. Then it got stuck on the DNS-01 challenges, which is where I hit a cert-manager bug I’ve now seen enough times to recognize on sight.
When a challenge fails, cert-manager tries to clean up the TXT record it created in Cloudflare — but it issues the delete with an empty zone ID, gets a 7003 error back, and can’t finish. The challenge keeps its finalizer, the finalizer keeps the challenge alive, and the whole thing sits there forever with stale _acme-challenge records still in DNS. The fix is manual and always the same: delete the leftover TXT records through the Cloudflare API, then null out the finalizers so the dead challenges can actually go away.
kubectl patch challenge <name> -n <ns> \
--type merge -p '{"metadata":{"finalizers":null}}'
Then one more delay that isn’t anyone’s bug. After clearing the records, _acme-challenge for my domain kept resolving as NXDOMAIN at Cloudflare’s edge for a good half hour, while sibling names under the same zone answered instantly. That asymmetry is exactly what makes you think you’ve broken something. It’s negative caching — the zone’s SOA minimum is 1800 seconds, so a “this doesn’t exist” answer is entitled to stick around for thirty minutes. The fix was to wait it out. We waited it out. The challenges completed. All three certificates are now valid to November 29th.
The probe that doesn’t trust cert-manager
The fix I actually care about isn’t the restart. CC built it the same night: a small CronJob that runs every six hours in the production cluster’s monitoring namespace, opens a TLS connection to each of my endpoints from the outside, reads the expiry date off the certificate it’s handed, and logs it. Loki picks up the log line, and a TLSCertExpiringSoon rule fires if anything is inside fourteen days.
It’s deliberately, almost rudely dumb. It doesn’t query cert-manager. It doesn’t read a Certificate resource or scrape a metrics endpoint or care whether a controller believes it’s healthy. It asks the same question a browser asks — what date is on this certificate? — and that answer is true whether or not the thing that issued it is functioning.
Three certificates had to expire and my control plane had to be degraded for nine days before I noticed I had never actually been measuring the outcome. I’d been measuring a controller’s opinion about the outcome. Those are not the same thing, and the gap between them is invisible right up until it isn’t.
The version of cert-manager that did this to me is also years out of date — it turned up as end-of-life in the first run of a new inventory job I’ll write about separately. I have a much stronger opinion about that upgrade than I did three weeks ago.