The Monitoring Died With the Node
A worker node in my production cluster hung at 4:38 on a Saturday afternoon. The NIC link stayed up at the switch, so the port looked perfectly healthy. The host itself was dead — no ping, no SSH, nothing. Kubernetes marked it NotReady, its pods went to Terminating, and there they sat.
Not a single alert fired. Not that day, not the next day, not for the rest of the week.
The reason is almost too neat. Prometheus ran on that node. So did Loki. So did the worker that dispatches alerts to me. There was no heartbeat from outside the cluster — no dead-man switch, nothing watching for silence. So when the node died it took the alerting path down with it, and an absence of alerts looks exactly like everything being fine.
What went down, and what didn’t
Eleven Longhorn volumes were stranded on the dead node. Forgejo, my git server. Grafana, Prometheus and Loki. Paperless. Sonarr. OpenClaw. FalkorDB, which holds the shared memory my agents read and write. Home Assistant’s config. The alert dispatcher’s own state.
Replacement pods got scheduled and then couldn’t attach anything, because Longhorn’s node-down policy was set to do nothing. That’s the default, and I’d never looked at it. When a node goes away, Longhorn is content to leave the volume marked as attached to a machine that no longer answers. The old pod stays Terminating forever, the new pod stays stuck, and the cluster sits in a stable, permanent, wrong state. Home Assistant had a second problem on top of that: it’s pinned to that specific node with a nodeSelector, so it wasn’t coming back no matter what Kubernetes did.
The part that didn’t matter in the end: the data. Every volume came back intact, and Longhorn’s replication across the other nodes is the reason. This was an availability outage, not a data incident, and that’s the one piece of the design that worked as intended.
CC found it during a routine cluster status check — not because anything told it to look. It worked out the standard recovery immediately: force-delete the stranded Terminating pods so the volumes release. Then its own permission guard refused the command, because a bulk destructive operation against live pods is precisely the category it isn’t allowed to run unsupervised. So it wrote out the exact commands and handed them to me instead.
I’m not going to pretend the guard was the problem. The guard was right — that is exactly the class of command I don’t want an agent running on its own against live pods. But it meant the recovery needed hands, and the node needed a physical power cycle regardless, which meant it needed me in the basement.
One power button
The fix was a hard power cycle at 12:25 on the 18th, and recovery after that was undramatic in a way I found slightly insulting. The node came back Ready inside a minute. Longhorn released all eleven volumes on its own, the stuck pods terminated, the replacements attached, and every service was back inside five minutes. The force-deletes CC had queued up turned out not to be needed at all. Six days of outage, one power button, five minutes of recovery.
The cause
At 06:01 UTC that day — just before midnight my time, the night before — unattended-upgrades installed a libc6 update. The post-upgrade pass that restarts services linked against the old library restarted two things that matter a great deal on a storage node: iscsid and systemd-journald. Longhorn volumes attach over iSCSI, and they were live and in use at the time.
journald never came back. Roughly seventeen hours later, the node hung.
That’s the sequence, and I’ll be honest about how much weight it carries: I’m confident in it, but it’s a reconstruction built from what the restart pass did and when the host went quiet, not from a crash dump. The host didn’t go down in a way that left much of a body to examine. A plausible mechanism — iSCSI being yanked out from under mounted volumes — plus a trigger that correlates exactly, and no autopsy to confirm it. I’d rather say that plainly than dress up a strong hypothesis as a proven root cause.
The one concrete bug that did fall out of the investigation: multipathd had no blacklist on five of my seven nodes, and on one of them it had grabbed a Longhorn device out from under the volume it belonged to — Sonarr’s config, as it happens. Longhorn documents the blacklist you’re supposed to apply, I’d never applied it, and CC rolled it out to every node via an Ansible playbook the same day. Unrelated to the hang as far as I can tell, but it was a live landmine sitting in the storage layer.
What changes
A dead-man switch outside the cluster, first and above everything else. I had a monitoring stack that could tell me about every failure except its own absence. Something off-cluster needs to expect a regular heartbeat and complain when it stops arriving, because a system that only reports problems it can see will always be silent about the one that blinded it.
Then: stop co-locating the entire observability stack on one worker. Prometheus, Loki and the alert dispatcher sharing a single node is a single point of failure dressed up as a monitoring system, and anti-affinity rules are not hard to write.
Set Longhorn’s node-down policy to actually delete the stranded pods. Exclude iscsid from automatic service restarts on storage nodes, or stop running unattended upgrades on them at all. Enable the hardware watchdog so a hung node reboots itself instead of waiting for me to walk downstairs. And label the boxes — they’re identical mini PCs on a shelf and I have never once labelled them, which is a problem you only have at the exact moment you need to find one.
There’s one more thing I noticed afterward and don’t much enjoy. A NodeClockNotSynchronising alert for that node had fired more than once in the hours before the hang — its clock had stopped correcting drift, the error growing unbounded from six seconds to sixteen while every other node stayed fine. I saw it. I didn’t act on it. It’s the only warning the system gave me, and I was the component that ignored it.
I wrote ten days ago about an expiry alert that lived on the one cluster it couldn’t see, and before that about backups sharing fate with the thing they protect. This is the third time I’ve written down the same lesson wearing a different costume. At some point the pattern stops being a series of interesting incidents and starts being a thing I apparently need to check for on purpose, every time, as a habit: if this component dies, what tells me? If the answer is “this component”, I haven’t built monitoring. I’ve built a component that agrees with itself.