The Uplink Was in Another Room
Power went out at 00:51 and came back at 03:16. My Kubernetes cluster did not come back for another four and a half hours.
Every machine in the rack had power back. The chain that kept the cluster down for the rest of the morning ran entirely outside it.
What the UPS did right, and what it couldn’t do
The rack UPS is healthy and its batteries are new. It carried the load for 25 to 28 minutes, and in that window NUT did exactly what it’s configured to do: shut the management host down cleanly before the battery ran out.
That’s the good half. The bad half is that the cluster nodes aren’t NUT clients. They were never told the power was going away, so they were hard-powered-off mid-write when the outlets dropped. The gateway, which I had assumed was on battery and was not, died within about a minute of the mains — so for the entire event the network had no edge.
So I had a UPS performing a careful, orderly shutdown of exactly one machine while the rest of the rack got its plugs pulled. I’d somehow never looked at it that way.
When power returned at 03:16, most of the rack booted itself. One worker node and one NAS stayed off, because their power-recovery settings differ from everything around them — the node’s firmware is set to stay off after a power loss, and the NAS wasn’t set to restart either. Going through the rest of the nodes afterwards turned up a second one with the same firmware setting, which simply hadn’t mattered this time. Easy to miss, because you only find out during the event you were trying to survive.
The gap after the power came back
Here’s the part I didn’t predict. The nodes that came back at 03:16 came back to nothing. Without a default route they had no path to the control plane or to anything outside, so the cluster sat there powered on and blind.
The rack’s only uplink is a wireless bridge, and the switch that powers the other end of that bridge is not in the rack. It took most of the morning to come back — I still don’t know whether it recovered on its own or was power-cycled — and the cluster got a default route within a minute of that switch coming up.
I built that uplink as a stopgap and then stopped thinking about it. The whole cluster’s connectivity turned out to rest on one device outside the rack that no part of my monitoring had an opinion about. Nothing inside the rack can tell you that. You can only see it by tracing power and link, device by device, which is exactly what nobody wants to do at four in the morning.
I got the diagnosis wrong first
CC’s first read — and mine, because I believed it — was that this was a repeat of the failure from two weeks earlier: a node hung by unattended-upgrades restarting services underneath live storage. That diagnosis came from one piece of evidence, the NotReady timestamp, and it was wrong.
What made it wrong was obvious once anything else got read. The node journals end abruptly with no shutdown record at all, which is a power cut, not a hang. The NUT log has the transfer-to-battery event and the clean shutdown that followed. The switch outside the rack has its own, much later boot time, and the rest of the network gear lines up with that switch rather than with the rack. Nothing we read after the first hour agreed with the story we’d opened with.
CC corrected the incident record and said plainly that its first read had been wrong. The root cause got revised twice before it settled. That matters more to me than getting it right the first time would have. The failure mode I actually worry about with an agent isn’t a bad hypothesis — I produce those too — it’s a confident narrative built on a single data point and never revisited. Pattern-matching to the most recent similar incident is a very easy mistake to make, and the fix is boring: before you explain an outage, go read the logs that would contradict you.
One concrete tell came out of this. A hung or powered-off mini PC does not show its switch port as Down. It shows Up at 10 Mb/s — the PHY sitting in standby for wake-on-LAN. A rack port negotiated at 10 Mb on a gigabit host now goes in my notes as the signature of a crashed or powered-off node, alongside the port map I should have written down years ago. I’d been reading “link up” as “machine alive” for far too long.
What followed
One thing changed during the recovery itself. Longhorn’s node-down pod deletion policy had been set to do nothing, which is why earlier incidents left me force-deleting stranded pods by hand. I switched it to delete, and the stranded volumes — Forgejo, OpenClaw, Prometheus — re-attached on surviving nodes in minutes instead of hours. That one is live on the cluster but still isn’t captured in any manifest, which means it would quietly vanish on a rebuild. Writing it down is on the list.
Everything else is still a plan, and I’d rather say that than imply otherwise. The rest of the work split cleanly in two, with two CC sessions taking a side each rather than racing to the same diagnosis. The cluster side is a graceful-shutdown pipeline, designed and not yet built: NUT on the management host notices the UPS going to battery and runs an Ansible playbook that drains and powers off the workers, then the control plane, then the NAS, then itself, inside the battery window. A graceful shutdown is worth more than runtime on a box that’s going down regardless.
Coming back up is the other half of that, and it’s where the power-recovery settings matter. Trying to fix the stuck firmware setting remotely failed outright — this particular BIOS and kernel combination rejects every attribute write, not just the one I wanted — so that’s a physical trip to the console with a keyboard. Wake-on-LAN is armed on both nodes in the meantime, which at least gives me a remote way to bring them up.
The network side is the uplink itself, and that one is a rebuild rather than a setting: a path from the rack to the gateway that doesn’t depend on a single device I’d forgotten I was depending on. I’m deliberately not writing the rest of that list down in public until it’s done.
The power cut was unavoidable. The gap after it was architecture. The wrong diagnosis was a process problem, and that’s the only one of the three I could fix for free.