The Alerts Were Events. I Wanted Cases.
I built the alert pipeline in my homelab because I wanted to know when something broke. Alertmanager fires, a dispatcher picks the alert up, a headless Claude Code worker diagnoses it against the live cluster, and I get a Telegram card: here’s what’s wrong, here’s the proposed fix, Approve or Reject.
It worked. That was the problem.
Over about a month it produced roughly 450 of those cards. I stopped reading them. What I said out loud turned out to be the actual bug report: I feel flooded, I can’t tell if they’re repeats.
So before touching any code I had CC go count. Pull every dispatch out of the database and ask how many distinct problems they represented.
Twenty-eight.
Four hundred and fifty cards, twenty-eight problems. One failure accounted for 248 of them on its own, each card diagnosed from scratch by a worker burning tokens to re-derive a cause it had already found, for a Job object that could never succeed again.
Fingerprint-per-instance is not dedupe
I thought I had deduplication. Alertmanager computes a fingerprint over an alert’s label set, and the dispatcher keyed on it: same fingerprint, no new card. Technically true and completely useless, because the label set includes the name of the object the alert is about — and that name was new every time.
Here is what was actually behind the 248. A log-review CronJob runs on a half-hour schedule, and its Jobs are pinned by podAffinity to a long-running Deployment pod. That Deployment had been stuck since the third of September: one replica, RollingUpdate strategy, a ReadWriteOnce Longhorn volume. maxUnavailable of 25% of one replica rounds down to zero, so the old pod can’t be removed, while maxSurge rounds up to one, so a new pod gets created first — and then sits in ContainerCreating forever, because the volume it needs is still attached to the old pod on another node. A rollout that can neither finish nor fall back.
The CronJob kept firing anyway. Every half hour it stamped out a Job, the Job’s pod couldn’t be placed, the Job failed, the alert arrived carrying a brand-new Job name in its labels, and the dispatcher — looking at a fingerprint it had genuinely never seen before — woke a worker, paid for a diagnosis, and sent me a card. That ran from the third of September until the day I rebuilt the pipeline.
One deadlocked Deployment, 248 notifications. The same shape shows up anywhere the object name moves: pods carry a random hash, StatefulSet members an ordinal, CronJob children a timestamp. What changed every time was the instance; the thing I cared about never changed at all.
The pipeline was deduplicating events. Nothing in it had any concept of a problem that persists — which is the only unit I actually care about, because I don’t fix alerts, I fix causes.
The approvals I wasn’t giving
Underneath the noise was something I liked less, and it wasn’t about the machine.
An earlier count, over a different and much shorter window, looked at what I’d done with the cards rather than how many there were. Fifty-five percent of them had expired unanswered. Just over one percent ended in a resolution. I hadn’t rejected them and I hadn’t argued with them. I’d let them time out.
Some could never have worked anyway. The executor’s service account holds remediation rights in some namespaces and not others, and a stuck-Job cleanup in a namespace it can’t write to isn’t something I can authorize from a phone — approving it doesn’t grant anyone the permission to carry it out. Extending that RBAC is still open.
I wrote a post a few months ago about who gets to push the button in an agentic system, and I was pleased with myself for drawing a clean line between an agent that proposes and a human who approves. The line was in the right place. I just wasn’t standing on my end of it — and in a fraction of cases there was nothing on the other end either.
Cases, not events
The redesign went out the same day. The core of it is one change of unit: alerts no longer create notifications, they fold into cases.
A case is keyed on the alert name, the namespace, and a normalized target — CronJob time suffixes stripped, pod hashes stripped, StatefulSet ordinals stripped. Everything else follows from that:
- A repeat of an existing case bumps a counter instead of creating anything.
- A case whose plan is less than 24 hours old is never re-diagnosed. No worker, no tokens, no new card.
- Acknowledged and snoozed cases swallow their repeats silently.
- A verified auto-fix resolves the case — and if the alert comes back, the case reopens on its own.
Telegram now pages only for things I’d want to stop what I’m doing and go look at: anything critical, a node down or unreachable, etcd, clock drift, Longhorn, volumes filling, backups, the API server or a kubelet going away. Everything else waits in the queue without making a sound, and I get one digest message a day at seven in the morning. If I’m curious in between, a /queue command tells me where things stand.
For the actual work there’s a CLI:
forge queue list
forge queue show <case>
forge queue ack|snooze|resolve|approve|reject <case>
forge queue resolve-stale
Plus a Claude Code skill that turns “clear the queue” into a sit-down session. CC walks the open cases, re-verifies each one against the live cluster — because a diagnosis from Tuesday may describe a problem that no longer exists — proposes what to do, and I approve once per root cause instead of once per alert. That’s the ratio that was broken.
What the numbers did
CC tested the migration against a copy of the live dispatcher database before anything touched production, which I’d like to claim credit for and can’t — it was the right instinct and it wasn’t mine. The exact count was 455 dispatches, and they folded into 28 cases, 20 of them still open.
Then the first clearing pass. Fifteen cases were stale — alerts silent for more than two days, describing problems that had resolved themselves or been fixed by hand weeks earlier — and went away in one bulk resolve. Two more had cleared and were closed properly. Three were acknowledged and left open, waiting on manual deletes outside the executor’s reach, which is where that RBAC gap stops being theoretical.
Five days later the dispatcher had recorded 12 dispatches, 7 of them quiet queue entries I read in a digest. Five days and one cluster isn’t a trend — but it’s the first week in a month where I read everything the pipeline sent me.
The open cases that remain are honest ones. A Nextcloud cron pointing at a PVC name that doesn’t exist, which has had a Job stuck for most of a year. Stuck Jobs in a couple of namespaces the executor genuinely cannot write to, which now surface as cases marked as needing my hands, with no button attached implying otherwise. Both are real work I’d been unable to see through the noise.
The lesson I’d generalize
Monitoring systems are built to emit events, because events are what they observe. Humans do not work in events. We work in cases: a thing that is wrong, which has a state, which persists until someone changes something, and which should get louder only if it gets worse.
Every layer between Prometheus and a person is a place to do that conversion, and I had skipped it entirely. I’d gone straight from “an alert fired” to “ping a human,” added an AI diagnosis step in the middle, and then been surprised that making each notification smarter didn’t make 450 of them tolerable. Adding intelligence to a per-event pipeline just produces well-reasoned spam.
The other half, which I’ll be chewing on longer: I had a correct principle — the agent proposes, I approve — and I built a loop that asked me to approve the same fix 248 times to clear one broken object. The principle wasn’t wrong. But a system that asks for that much consent isn’t really asking, and I was always going to stop answering.