Skip to content

Engineering notes from a working homelab

A homelab run like production — one hypervisor, ninety-nine pipelines, an agent that operates it under a written authority envelope, and a decision log that records what was rejected as well as what was chosen.

This site is written for a peer — a systems engineer or an architect who will want to know why a thing is shaped the way it is, and who will notice if the answer is thin. It is not a tour of what software is installed.

The estate behind it is a single Proxmox host and the services on it, run deliberately against production discipline: everything declared in a repository, every change applied by a pipeline, every service owing monitoring, observability, alerting, vulnerability coverage and documentation before it counts as delivered. An AI agent does the day-to-day engineering under a written authority envelope, and the ways it has fooled itself are a named taxonomy it carries into every session.

  • 471architecture decisions recorded
  • 97CI/CD pipelines, 76 of them deploying
  • 1,503pages in the documentation corpus
  • 402recorded agent lessons

Every number on this site is derived at build time from the estate itself and stamped with the moment it was taken. None of it is hand-typed into prose, and none of it is presented as live.

  • Platform and compute

    Three declaration paths, one privileged path to the hypervisor, and a snapshot helper that refuses to pretend a rollback point exists when it does not.

  • Delivery and GitOps

    Ninety-nine pipelines on a two-agent pool, trigger economy as capacity design, and why a green pre-push hook is the most convincing counterfeit in the estate.

  • Observability and SRE

    RED and USE with an honest fallback chain, log shipping proven rather than assumed, and alerting designed around a rule that nothing may ever wake the operator.

  • Security and compliance

    Zero-trust at the edge with an approval register behind it, fleet SBOMs, a secret-scan conformance gate, and two prudential standards run properly rather than claimed.

  • AI-assisted operations

    An authority envelope that has narrowed as well as widened, knowledge retrieved rather than remembered, and a failure taxonomy loaded before the work instead of written up after it.

Six things that went wrong, and what came out of them

Section titled “Six things that went wrong, and what came out of them”

Each of these is a real incident with a real remedy, written as problem, design, trade-off and evidence. They are the shortest route into how this estate actually thinks.

  • Estate topology

    The derived service-dependency graph, published with every identifier stripped and the stripping asserted.

  • Decision log

    Every architecture decision in the corpus, filterable by domain, status and year — including the gaps.

  • Delivery performance

    The four DORA keys over twenty-eight days of real pipeline history, each with its derivation and sample count.

  • Agent operations

    Decision cadence, the failure-class taxonomy, and how well each recorded lesson is evidenced.

The standards and designs sections are the estate’s own internal documents, published verbatim where they can be. They are written for an operator, not for a visitor, so they are dense and they assume context. That is deliberate: a showcase that only contains material written to be read by a stranger is a brochure, and the point of this one is that you can walk into the source material and cross-examine it.

Everything on this site is a snapshot of a running system, and it says so on every panel. If you want to talk about any of it, the colophon has the contact details and explains how the site is built.