Are All Our Containers Up to Date?

I asked CC — my Claude Code agent — a question I thought had a short answer: is everything I run up to date?

It did not have a short answer.

What came back was an audit of every third-party thing I run across both clusters, and most of it was behind. Not alarmingly behind in a way I’d have noticed. Behind in the way a house settles.

The part I didn’t want to read

A few items were past end-of-life entirely, which is a different category from “behind.” Behind means you’re missing features. End-of-life means nobody is writing security fixes for the version you’re running.

The one that stung: my Git server was more than a year past end-of-life. It hosts every repository I have, including the ones that describe how to rebuild everything else. The thing I’d reach for during a disaster was the thing running on code nobody maintains anymore.

The rest of the list was in the same spirit — the file sync, the dashboards, the distributed storage layer, the certificate controller, and on the management cluster a second copy of that controller from a release vintage I will not defend. Both Kubernetes control planes had aged past their support windows, and neither could move until the management layer above them moved first. One ingress controller wasn’t just old; the project itself had been retired, which means there is no newer version to go to and the real fix is a migration. And one component was sitting on a version with a published advisory against it — not theoretical, not disputed, just a fix I hadn’t applied.

I’m being vague about versions on purpose, and that vagueness is itself part of the story: a list like this is a map of where to push, and some of these are still open as I write.

And then the one that actually changed how I think about this. One app was running on a :latest tag. Not pinned — floating. The image in the registry had moved on to a new major version. Nothing was wrong, because nothing had restarted. One pod eviction, one node reboot, one of the many ordinary things that happen in a cluster every week, and it would have come back up on a major version I never chose, at a moment I wasn’t watching.

That’s not an out-of-date container. That’s a loaded gun with a slow trigger.

I don’t want to be the scanner

My honest reaction was less “I should go patch things” and more “I feel like I’m always chasing updates.” That turned out to be the useful complaint. A one-time audit tells you where you are today and starts rotting immediately. I’d done manual sweeps before. They don’t hold, because the next one depends on me remembering to care on a particular Saturday.

So instead of spending the weekend upgrading, CC and I spent it building the thing that asks the question on a schedule. It’s called update-radar, it runs once a week, and it inventories everything I could think of:

  • container images on both clusters — pinned tags read from the manifests, version probes run inside the pods where the tag doesn’t tell you enough, and registry digest comparison for the floating tags, which is the only way to catch the :latest case above
  • Helm releases, and the Kubernetes distribution versions along with their actual upgrade ceilings, because “the newest Kubernetes” is meaningless if the management layer has to move first
  • OS patch state for the cluster nodes and the other machines here, collected over SSH — which I already want to replace, because it means the scanner holds SSH credentials, and a thing whose whole job is reading shouldn’t need keys that good
  • the network gear’s firmware and the NAS

Then it compares all of that against upstream: release feeds and registries for what’s current, published support windows for what’s expired, and security advisories matched against the version actually running — not against the newest version, which is the mistake that makes vulnerability scanners useless.

Outputs by design, not by inbox

Where the results go is the part I care most about, because this is where tools like this usually die. A weekly email becomes a weekly email you don’t read. So the report isn’t the product. The product is a queue.

Every item that needs work gets one issue, and the issue closes itself when the component comes current. That’s the whole trick — the backlog has real state, rather than being a snapshot I re-read each week to work out what’s new. Alongside that it pushes metrics with a small number of alerts, posts a short Telegram summary, commits a dated report with a week-over-week diff, and rewrites a memory file so every future CC session starts out already knowing what’s behind. The agent shouldn’t have to rediscover the state of my lab every time I open a terminal.

A recent week’s report, three weeks in, reads: critical 0 · high 16 · medium 10 · low 24. New 1, resolved 4, changed 34.

Those last three numbers are the ones I actually read. Fifty open items is a number I can’t do anything with. “One new thing, four closed” I can act on in a morning. The four that closed were all floating :latest tags I’d gone and pinned — the report noticed they’d stopped being a risk without me telling it so.

The severity rules ended up being the most opinionated part of the whole thing, and only two of them really matter:

  • End-of-life beats “behind.” Six minor versions behind on something still supported is housekeeping. One version behind on something nobody patches anymore is a decision I’m making whether I know it or not.
  • A floating tag that would jump a major on restart is its own category. Not “out of date” — it’s current, technically. It’s a scheduling risk, and it deserves a different word.

The meta moment

CC did all of it — wrote the collectors, the scoring, the issue integration, the metrics, the report format, and the memory file, and it ran the original audit that started this. I provided the complaint.

I’m genuinely unsure whether the severity rules are tuned right. Sixteen high items could mean the lab is in bad shape, or it could mean my thresholds are too eager; I won’t know until I’ve watched more weeks of diffs and seen whether the high list actually shrinks or just churns. So far it has mostly held steady, which is its own answer and not a flattering one.

I’ve also deliberately not given this thing permission to upgrade anything. It files issues. I still decide. The components on that list are the ones where a bad upgrade is a bad weekend, and I’d rather be the slow part of that loop.

It also isn’t finished. It runs as a timer on one laptop, which makes the thing that watches my cluster dependent on a machine that isn’t in it. The plan is to move it in-cluster and to stop collecting host patch state over SSH. Neither is built. I’m aware of the shape of this — I built a thing to find components I’d stopped maintaining, and it currently has a single point of failure I haven’t fixed.

The answer to my original question wasn’t a list of version numbers. It was how a lab that mostly runs itself drifts — floating tags that are fine until they restart, support windows that expire without an event, retired projects where there is no upgrade to do. None of that shows up as an outage. It shows up as an audit, if you ever run one.

Now something runs one every Monday, and I don’t have to be the one who remembers.