The Uplink Was in Another Room
Mains power failed at 00:51 and came back at 03:16, but my cluster stayed dark most of the morning. The reason was a switch outside the rack, and my first explanation for the whole thing was wrong.
Tech, projects, opinions, and whatever else is on my mind.
Mains power failed at 00:51 and came back at 03:16, but my cluster stayed dark most of the morning. The reason was a switch outside the rack, and my first explanation for the whole thing was wrong.
My homelab sent me about 450 Telegram alert cards in a month. Reviewed properly, they were 28 actual problems, and one stuck Deployment accounted for 248 of them. Rebuilding the pipeline around cases instead of events fixed the flood — and showed me how little I'd been doing with the approvals it asked for.
Minecraft 26.3 broke gamepad support on our family server because the one mod that provides it hadn't been ported. CC finished the maintainer's unfinished port branch, verified it in the real client, and sent the generalizable part upstream — where the maintainer closed it, and his review found the mistake we'd made.
A worker node hung on a Saturday afternoon and took eleven Longhorn volumes down with it. No alert fired, because Prometheus, Loki and the alert dispatcher were all running on the node that died.
cert-manager on my management cluster quietly stopped renewing certificates three weeks before I noticed. The pod stayed green the whole time, my expiry alert lived on the one cluster that couldn't see it, and I only found out when kubectl started hanging.
I asked whether everything I run was up to date and got an answer I didn't want: most of the stack was behind, and several pieces were past end-of-life. The fix wasn't an upgrade weekend, it was turning the question into a weekly job.
A few weeks ago I decided my AI agent should orchestrate, not do everything itself. Tonight we built the thing it orchestrates with: Forge, a system that dispatches headless Claude Code workers into my Kubernetes cluster to build software. Here's what we made, the bug we hit, and where this is going.
I tried to finish a year-old broken Longhorn upgrade, it blew up exactly the way my AI had warned it would, and the rollback was clean. Here's the post-mortem — including why the failure was the software working correctly, and why I'm not filing a bug.
My production cluster's etcd snapshots were sitting on the same disks they were supposed to protect. Closing that gap turned into a nice little lesson in not fighting your tools.
I wanted my homelab backups to survive the NAS dying. That one requirement — no shared fate — quietly rewrote the whole design: 3-2-1 with fan-out instead of a chain, a backup key per app, and the boring discipline of actually testing a restore.
I asked whether Ansible is dead now that I have an AI agent managing my homelab. The answer was no — and figuring out why produced a cleaner architecture for cluster updates than I had before: the agent gets the eyes and the brain, Ansible keeps the hands.

My 2007 RAV4 threw a misfire code, so I ran the whole repair past two AI assistants in parallel — one with web search, one without. The one searching the web got the part numbers right. The one working from memory invented them. The lesson is older than AI: you can't reason your way to a fact you don't have.
I gave Claude Code a dedicated Linux Mint workstation with a GUI, browser control, and full dev tools. Here's why, and how we set it up together.
I built a custom MCP server that gives Claude Code vision and desktop control — screenshots, mouse clicks, keyboard input — and then it sent a Slack message on its own.
How I deployed Paperless-ngx on my homelab cluster for household document management — from planning to AI-powered document intake via Telegram.

Weekend Project: The Farmhouse Bathroom
How I used Claude and an AI agent to design, spec, and build this blog without writing most of the code myself.
GiSquared started as a computer repair shop in 2006. Twenty years later, it's back as something new.