Status archive

Dated snapshots of the Status page. That page shows only the present; this one shows the trajectory, which for a project whose whole subject is the gap between designed and running is the more honest view.

Newest first. Nothing here is edited after the fact — each entry is the status page exactly as it read on that date.

How the numbers moved

DateCompleteIn progressNot startedTotal
2026-09-1839131466
2026-09-1438111665
2026-09-0436111764
2026-08-3136101864
2026-08-2736101864
2026-08-2433122065
2026-08-2232112265
2026-08-202882965
2026-08-182583265
2026-08-162383465
2026-08-131664365
2026-08-121564465
2026-08-091384465
2026-08-071084765
2026-07-29954862
2026-07-28954862

Snapshots

2026-09-18 — 39 complete, 13 in progress, 14 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 18 September 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The application tiers are filling in fast. A bookmark archive, webmail, news reader, workflow automation, file sync and two independent personal finance tools on the productivity side; continuous integration, a container registry and now software supply-chain scanning on the development side; and a local language model runtime with a chat interface and a routing gateway in front of it, now also reachable through single sign-on rather than only from inside the mesh. Several are complete. Several are running without the monitoring or backups that this project counts as finished, and are listed as in progress for that reason rather than as done. A photo library cleared the manual step that had it blocked and took its first real import. A dedicated GPU host has joined the AI tier and is still being tested, not yet counted as finished.

An internal phone system has made its first real, encrypted call in and out, though nothing yet watches it stay up. An SMS-to-email bridge sends real messages through a carrier line; the reverse direction is still being built. Power draw for the physical gear now has its own dashboard.

Two separate tracks now run alongside the main build: an education lab of blade servers checked out through an API, and an application built by another agent that uses the lab as a tenant rather than changing it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, encrypted repositories
Personal finance, document archiveLive, restore drill passed
Continuous integration and container registryLive, both alert-proven
Local language model, chat interface, routing gatewayLive, no cloud providers, monitoring incomplete
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, backed up, login proven
Supply-chain scanningLive, images and deployed hosts both scanned

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • The photo library’s storage block cleared, and it is now actively serving. A first real import ran end to end — hundreds of gigabytes, tens of thousands of photos, one upload error out of all of them — and the background processing it kicked off is still draining. It is counted as in progress rather than complete because monitoring and backup have not caught up to it yet, not because it isn’t working.
  • File sync went from not started to reachable. The setup wizard has run and the service answers requests. Single sign-on, the shared-storage connection to the rest of the file library, and backup are all still outstanding — this is a service that exists, not one that is finished.
  • The bulk storage tier is the other part-built piece on the productivity side, started rather than designed and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • Two live services have no liveness monitor. The language model runtime and its chat interface both answer requests and neither would raise an alert if it stopped. They are counted as in progress rather than complete for exactly this reason: by this project’s own definition, a service that runs and is not watched is not finished.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharing, external storage, single sign-on, offsite backup, restore rehearsedComplete
OutlineNotes and internal wikiComplete
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveComplete
Firefly IIIPersonal financeComplete
Actual BudgetPersonal finance, second instance, single sign-on onlyComplete
n8nWorkflow automationComplete
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLIn progress
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone system, internal extensions, voicemailIn progress
SMS bridgeEmail-to-SMS two-way messaging, outbound provenIn progress
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hosting and container registryComplete
Woodpecker CIBuild automation on an isolated VLANComplete
Dependency-TrackSBOM and vulnerability trackingComplete
Syft + GrypeSBOM generation and scanningComplete
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtime, CPU, with one GPU-backed host under testIn progress
Open WebUIChat interfaceIn progress
LLM gatewayRouting and per-service keys, no cloud providersComplete
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierOne card selected and under test on a new dedicated hostIn progress

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete
UPS / power telemetryLoad and runtime, alert-proven by induced failureComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Thirty-nine components complete. Thirteen in progress. Fourteen not started.

Every one of those sixty-six has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-09-14 — 38 complete, 11 in progress, 16 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 14 September 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The application tiers are filling in fast. A bookmark archive, webmail, news reader, workflow automation, file sync and two independent personal finance tools on the productivity side; continuous integration, a container registry and now software supply-chain scanning on the development side; and a local language model runtime with a chat interface and a routing gateway in front of it, now also reachable through single sign-on rather than only from inside the mesh. Several are complete. Several are running without the monitoring or backups that this project counts as finished, and are listed as in progress for that reason rather than as done. A photo library cleared the manual step that had it blocked and took its first real import. A dedicated GPU host has joined the AI tier and is still being tested, not yet counted as finished.

Two separate tracks now run alongside the main build: an education lab of blade servers checked out through an API, and an application built by another agent that uses the lab as a tenant rather than changing it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, encrypted repositories
Personal finance, document archiveLive, restore drill passed
Continuous integration and container registryLive, both alert-proven
Local language model, chat interface, routing gatewayLive, no cloud providers, monitoring incomplete
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, backed up, login proven
Supply-chain scanningLive, images and deployed hosts both scanned

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • The photo library’s storage block cleared, and it is now actively serving. A first real import ran end to end — hundreds of gigabytes, tens of thousands of photos, one upload error out of all of them — and the background processing it kicked off is still draining. It is counted as in progress rather than complete because monitoring and backup have not caught up to it yet, not because it isn’t working.
  • File sync went from not started to reachable. The setup wizard has run and the service answers requests. Single sign-on, the shared-storage connection to the rest of the file library, and backup are all still outstanding — this is a service that exists, not one that is finished.
  • The bulk storage tier is the other part-built piece on the productivity side, started rather than designed and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • Two live services have no liveness monitor. The language model runtime and its chat interface both answer requests and neither would raise an alert if it stopped. They are counted as in progress rather than complete for exactly this reason: by this project’s own definition, a service that runs and is not watched is not finished.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharing, external storage, single sign-on, offsite backup, restore rehearsedComplete
OutlineNotes and internal wikiComplete
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveComplete
Firefly IIIPersonal financeComplete
Actual BudgetPersonal finance, second instance, single sign-on onlyComplete
n8nWorkflow automationComplete
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLIn progress
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hosting and container registryComplete
Woodpecker CIBuild automation on an isolated VLANComplete
Dependency-TrackSBOM and vulnerability trackingComplete
Syft + GrypeSBOM generation and scanningComplete
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtime, CPU, with one GPU-backed host under testIn progress
Open WebUIChat interfaceIn progress
LLM gatewayRouting and per-service keys, no cloud providersComplete
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierOne card selected and under test on a new dedicated hostIn progress

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Thirty-eight components complete. Eleven in progress. Sixteen not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-09-04 — 36 complete, 11 in progress, 17 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 4 September 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The application tiers are filling in fast. A bookmark archive, webmail, news reader, workflow automation, personal finance and document archive on the productivity side; continuous integration, a container registry and now software supply-chain scanning on the development side; and a local language model runtime with a chat interface and a routing gateway in front of it, now also reachable through single sign-on rather than only from inside the mesh. Several are complete. Several are running without the monitoring or backups that this project counts as finished, and are listed as in progress for that reason rather than as done. A photo library cleared the manual step that had it blocked and took its first real import.

Two separate tracks now run alongside the main build: an education lab of blade servers checked out through an API, and an application built by another agent that uses the lab as a tenant rather than changing it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, encrypted repositories
Personal finance, document archiveLive, restore drill passed
Continuous integration and container registryLive, both alert-proven
Local language model, chat interface, routing gatewayLive, no cloud providers, monitoring incomplete
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, backed up, login proven
Supply-chain scanningLive, images and deployed hosts both scanned

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • The photo library’s storage block cleared, and it is now actively serving. A first real import ran end to end — hundreds of gigabytes, tens of thousands of photos, one upload error out of all of them — and the background processing it kicked off is still draining. It is counted as in progress rather than complete because monitoring and backup have not caught up to it yet, not because it isn’t working.
  • File sync went from not started to reachable. The setup wizard has run and the service answers requests. Single sign-on, the shared-storage connection to the rest of the file library, and backup are all still outstanding — this is a service that exists, not one that is finished.
  • The bulk storage tier is the other part-built piece on the productivity side, started rather than designed and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • Two live services have no liveness monitor. The language model runtime and its chat interface both answer requests and neither would raise an alert if it stopped. They are counted as in progress rather than complete for exactly this reason: by this project’s own definition, a service that runs and is not watched is not finished.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingIn progress
OutlineNotes and internal wikiComplete
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveComplete
Firefly IIIPersonal financeComplete
n8nWorkflow automationComplete
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLIn progress
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hosting and container registryComplete
Woodpecker CIBuild automation on an isolated VLANComplete
Dependency-TrackSBOM and vulnerability trackingComplete
Syft + GrypeSBOM generation and scanningComplete
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtime, CPU onlyIn progress
Open WebUIChat interfaceIn progress
LLM gatewayRouting and per-service keys, no cloud providersComplete
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Thirty-six components complete. Eleven in progress. Seventeen not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-31 — 36 complete, 10 in progress, 18 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 31 August 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The application tiers are filling in fast. A bookmark archive, webmail, news reader, workflow automation, personal finance and document archive on the productivity side; continuous integration, a container registry and now software supply-chain scanning on the development side; and a local language model runtime with a chat interface and a routing gateway in front of it, now also reachable through single sign-on rather than only from inside the mesh. Several are complete. Several are running without the monitoring or backups that this project counts as finished, and are listed as in progress for that reason rather than as done. A photo library cleared the manual step that had it blocked and took its first real import.

Two separate tracks now run alongside the main build: an education lab of blade servers checked out through an API, and an application built by another agent that uses the lab as a tenant rather than changing it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, encrypted repositories
Personal finance, document archiveLive, restore drill passed
Continuous integration and container registryLive, both alert-proven
Local language model, chat interface, routing gatewayLive, no cloud providers, monitoring incomplete
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, backed up, login proven
Supply-chain scanningLive, images and deployed hosts both scanned

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • The photo library’s storage block cleared, and it is now actively serving. A first real import ran end to end — hundreds of gigabytes, tens of thousands of photos, one upload error out of all of them — and the background processing it kicked off is still draining. It is counted as in progress rather than complete because monitoring and backup have not caught up to it yet, not because it isn’t working.
  • The bulk storage tier is the other part-built piece on the productivity side, started rather than designed and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • Two live services have no liveness monitor. The language model runtime and its chat interface both answer requests and neither would raise an alert if it stopped. They are counted as in progress rather than complete for exactly this reason: by this project’s own definition, a service that runs and is not watched is not finished.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiComplete
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveComplete
Firefly IIIPersonal financeComplete
n8nWorkflow automationComplete
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLIn progress
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hosting and container registryComplete
Woodpecker CIBuild automation on an isolated VLANComplete
Dependency-TrackSBOM and vulnerability trackingComplete
Syft + GrypeSBOM generation and scanningComplete
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtime, CPU onlyIn progress
Open WebUIChat interfaceIn progress
LLM gatewayRouting and per-service keys, no cloud providersComplete
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Thirty-six components complete. Ten in progress. Eighteen not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-27 — 36 complete, 10 in progress, 18 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 27 August 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The application tiers are filling in fast. A bookmark archive, webmail, news reader, workflow automation, personal finance and document archive on the productivity side; continuous integration, a container registry and now software supply-chain scanning on the development side; and a local language model runtime with a chat interface and a routing gateway in front of it, now also reachable through single sign-on rather than only from inside the mesh. Several are complete. Several are running without the monitoring or backups that this project counts as finished, and are listed as in progress for that reason rather than as done. A photo library has started and is blocked on a manual step, not a decision.

Two separate tracks now run alongside the main build: an education lab of blade servers checked out through an API, and an application built by another agent that uses the lab as a tenant rather than changing it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, encrypted repositories
Personal finance, document archiveLive, restore drill passed
Continuous integration and container registryLive, both alert-proven
Local language model, chat interface, routing gatewayLive, no cloud providers, monitoring incomplete
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, backed up, login proven
Supply-chain scanningLive, images and deployed hosts both scanned

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • A photo library was started and then blocked, not on a decision but on a file share the automation account cannot create for itself on the storage box — a four-minute manual step, waiting on a person rather than a part.
  • The bulk storage tier is the other part-built piece on the productivity side, started rather than designed and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • Two live services have no liveness monitor. The language model runtime and its chat interface both answer requests and neither would raise an alert if it stopped. They are counted as in progress rather than complete for exactly this reason: by this project’s own definition, a service that runs and is not watched is not finished.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiComplete
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveComplete
Firefly IIIPersonal financeComplete
n8nWorkflow automationComplete
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLIn progress
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hosting and container registryComplete
Woodpecker CIBuild automation on an isolated VLANComplete
Dependency-TrackSBOM and vulnerability trackingComplete
Syft + GrypeSBOM generation and scanningComplete
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtime, CPU onlyIn progress
Open WebUIChat interfaceIn progress
LLM gatewayRouting and per-service keys, no cloud providersComplete
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Thirty-six components complete. Ten in progress. Eighteen not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-24 — 33 complete, 12 in progress, 20 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 24 August 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The application tiers are filling in fast. A bookmark archive, webmail, news reader, workflow automation, personal finance and document archive on the productivity side; continuous integration and a container registry on the development side; and a local language model runtime with a chat interface and a routing gateway in front of it, now also reachable through single sign-on rather than only from inside the mesh. Several are complete. Several are running without the monitoring or backups that this project counts as finished, and are listed as in progress for that reason rather than as done. A photo library has started and is blocked on a manual step, not a decision.

Two separate tracks now run alongside the main build: an education lab of blade servers checked out through an API, and an application built by another agent that uses the lab as a tenant rather than changing it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, encrypted repositories
Personal finance, document archiveLive, restore drill passed
Continuous integration and container registryLive, both alert-proven
Local language model, chat interface, routing gatewayLive, no cloud providers, monitoring incomplete
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, no backup yet, login never tested

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • A photo library was started and then blocked, not on a decision but on a file share the automation account cannot create for itself on the storage box — a four-minute manual step, waiting on a person rather than a part.
  • The bulk storage tier and the notes and wiki service are the other two part-built pieces on the productivity side, both started rather than designed and neither running anywhere permanent yet.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • One live service still has no backup and no proven login, unchanged since 20 August. The notes and wiki service answers requests and is monitored, and its database sits on local disk in no backup repository of any kind. Its single sign-on client has never had a login attempted, so whether a user can get in, and with what permissions, is unknown. Every other service on that machine reached this state during its build and then had it closed.
  • Three live services have no liveness monitor. The document archive, the language model runtime and its chat interface all answer requests and none of them would raise an alert if they stopped. They are counted as in progress rather than complete for exactly this reason: by this project’s own definition, a service that runs and is not watched is not finished.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiIn progress
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveIn progress
Firefly IIIPersonal financeComplete
n8nWorkflow automationComplete
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLIn progress
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hostingComplete
Woodpecker CIBuild automation on an isolated VLANComplete
ZotContainer registryComplete
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtime, CPU onlyIn progress
Open WebUIChat interfaceIn progress
LLM gatewayRouting and per-service keys, no cloud providersComplete
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Thirty-three components complete. Twelve in progress. Twenty not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-22 — 32 complete, 11 in progress, 22 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 22 August 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The application tiers are filling in fast. A bookmark archive, webmail, news reader, personal finance and document archive on the productivity side; continuous integration and a container registry on the development side; and a local language model runtime with a chat interface and a routing gateway in front of it. Several are complete. Several are running without the monitoring or backups that this project counts as finished, and are listed as in progress for that reason rather than as done.

Two separate tracks now run alongside the main build: an education lab of blade servers checked out through an API, and an application built by another agent that uses the lab as a tenant rather than changing it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, encrypted repositories
Personal finance, document archiveLive, restore drill passed
Continuous integration and container registryLive, both alert-proven
Local language model, chat interface, routing gatewayLive, no cloud providers, monitoring incomplete
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, no backup yet, login never tested

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • Four services are part-built: the bulk storage tier, git hosting, the feed reader, and the mail server above. Started rather than designed, and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • One live service still has no backup and no proven login, unchanged since 20 August. The notes and wiki service answers requests and is monitored, and its database sits on local disk in no backup repository of any kind. Its single sign-on client has never had a login attempted, so whether a user can get in, and with what permissions, is unknown. Every other service on that machine reached this state during its build and then had it closed.
  • Three live services have no liveness monitor. The document archive, the language model runtime and its chat interface all answer requests and none of them would raise an alert if they stopped. They are counted as in progress rather than complete for exactly this reason: by this project’s own definition, a service that runs and is not watched is not finished.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiIn progress
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveIn progress
Firefly IIIPersonal financeComplete
n8nWorkflow automationNot started
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hostingComplete
Woodpecker CIBuild automation on an isolated VLANComplete
ZotContainer registryComplete
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtime, CPU onlyIn progress
Open WebUIChat interfaceIn progress
LLM gatewayRouting and per-service keys, no cloud providersComplete
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Thirty-two components complete. Eleven in progress. Twenty-two not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-20 — 28 complete, 8 in progress, 29 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 20 August 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier is running: metrics, dashboards, alerting, log aggregation and security monitoring, with alerts proven by inducing the failures rather than assuming delivery. Offsite backups exist as of 18 August.

The first services meant for people rather than for the lab itself are now live — a bookmark archive, a webmail client and a news reader, all behind single sign-on. A notes and wiki service is running alongside them and is not finished: it has no backup, and nobody has ever logged into it.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, eight encrypted repositories
Bookmark archive, webmail client, news readerLive behind single sign-on, backed up
Notes and wikiLive, no backup yet, login never tested

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • Four services are part-built: the bulk storage tier, git hosting, the feed reader, and the mail server above. Started rather than designed, and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • One live service has no backup and no proven login. The notes and wiki service answers requests and is monitored, and its database is on local disk in no backup repository of any kind. Its single sign-on client has never had a login attempted against it, so whether a user can actually get in, and with what permissions, is unknown. Every other service on that machine reached this same state during its build and then had it closed. This one has not yet.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiIn progress
LinkwardenBookmark archiveComplete
SnappymailWebmail client, gated by single sign-onComplete
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerComplete

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hostingComplete
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Twenty-eight components complete. Eight in progress. Twenty-nine not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-18 — 25 complete, 8 in progress, 32 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 18 August 2026, and audited against the build record on the same day.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier landed in the last three days: metrics, dashboards, alerting, log aggregation and security monitoring all moved from designed to running, with alerts proven by inducing the failures rather than assuming delivery. The storage tier remains unbuilt and blocks the offsite half of every backup in the lab.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisorOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Seven nodes on the meshLive
Twenty-three uptime monitorsLive, alert paths proven individually
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven
Git hostingLive, backed up, restore drill passed
Offsite backup chainLive, eight encrypted repositories

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • Four services are part-built: the bulk storage tier, git hosting, the feed reader, and the mail server above. Started rather than designed, and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Offsite backups now exist, and each one’s key lives on the machine it protects. Eight encrypted repositories were built on 18 August, in four legs that fail independently, with staleness alerts on each. The archives are encrypted at the source, so the copies hold only ciphertext.

The password for each repository currently exists in exactly one place: the host whose data it protects. So losing a host loses the key to that host’s backup. The archive survives, replicates offsite, and is unreadable. The chain was built to make host loss survivable and does not yet do that. It is the outstanding item on the whole backup effort, and it is deliberately a person’s job rather than an automated one, because it is credential handling.

  • Nothing measures the alert channel as a whole. Every alert in the lab is proven individually by inducing its failure and watching the notification arrive. None of that could detect the state found on 18 August, where a single misconfigured monitor had produced 97% of every alert sent over a week. The delivery path was healthy; it was carrying 38 false alarms a day. Alert quality turns out to be a property of the channel rather than of any one check, and it was not being measured.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage no longer holds up offsite backups, which were built on 18 August against the bulk storage box instead. It still holds up the integrity tier itself, which is where the small, precious, checksummed data is meant to live.

  • Ten of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host, the security monitoring host, and hosts for the productivity and development tiers whose services are still being built.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Get every backup repository’s key into a second place, so that losing a host does not lose the means to read its own backup.
  2. The quorum witness, which is what turns two servers into a cluster.
  3. The integrity storage tier, once the boot adapter arrives.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerIn progress

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
GiteaGit hostingComplete
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted at source, four independent legsComplete
Restore proceduresIdentity and git rehearsed; the vault is notIn progress
Maintenance playbookPatching and routine operationsNot started

Twenty-five components complete. Eight in progress. Thirty-two not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-16 — 23 complete, 8 in progress, 34 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 16 August 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it, and the first services on top of it are now running too. The second physical server is online, so the lab is no longer a single machine, though the two are not yet clustered. Eight service virtual machines are built.

The observability tier landed in the last three days: metrics, dashboards, alerting, log aggregation and security monitoring all moved from designed to running, with alerts proven by inducing the failures rather than assuming delivery. The storage tier remains unbuilt and blocks the offsite half of every backup in the lab.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisor, single nodeOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Four hosts on the meshLive
Eight uptime monitors, alert-provenLive
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly, restore rehearsed
Second physical hypervisorOnline, monitored, standalone
Shared secrets vaultLive behind single sign-on, local backups only
Metrics, dashboards and alertingLive, host-death alert proven by induced failure
Log aggregationLive, eight hosts shipping
Security monitoringLive, eight agents reporting
Certificate expiry probingLive, six targets
RAID and disk health alertingLive, both alert-proven

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.
  • Four services are part-built: the bulk storage tier, git hosting, the feed reader, and the mail server above. Started rather than designed, and not yet running anywhere permanent.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.
  • Every backup is on-host only. Nightly backups run for the identity provider and the secrets vault, with guards proven to fire, and the identity restore has been rehearsed against a real dump. None of it leaves the machine it protects. That covers corruption and does not cover losing the host, and it stays that way until the storage tier is built.
  • The metrics host cannot report its own host’s death. Monitoring runs as a virtual machine on the server it watches, so if that server dies the watcher dies with it. The second physical node now probes the first every minute precisely because of this, but the general shape is worth naming: a monitor inside the thing it monitors can only report degradation, never death.

Not started

  • The hypervisor cluster. Both nodes are now online, monitored, and running independently. They are not clustered, because a two-node cluster cannot break a quorum tie on its own and the third vote does not exist yet. Until the witness is built, there is no high availability, only two servers.
  • The storage tiers, and the reason is worth stating because it is the kind of thing that does not appear in a design document. The integrity-tier box is recycled enterprise hardware that does not support UEFI boot, so the plan to boot it from an NVMe drive in a PCIe slot does not work: the firmware cannot boot from that card. The current approach is to source an adapter that lets a standard SATA disk take the place of the onboard optical drive and serve as the boot volume. Waiting on a part, not on a decision.

That blockage reaches further than one machine. The storage tier is what makes offsite backups possible, so every backup in the lab is currently on-host only — they survive corruption, not the loss of the machine.

  • Eight of the sixteen service VMs. Built so far: the home-side gateway, two DNS resolvers, the control node, the identity server, the secrets vault, the observability host and the security monitoring host.

Verified how, not just verified

Log shipping and the security monitoring stack were both proven on throwaway virtual machines first, and both are now running for real. That rehearsal earned its keep: executing the automation found defects that four separate document reviews had missed entirely, including a case where the stated mitigation for a known vulnerability did not hold, and several where a verification step was structurally incapable of failing.

The standard applied since is that an alert is not counted until the failure has been induced and the notification observed arriving. A host was deliberately taken down to prove the host-death alert. A drive failure was simulated to prove the RAID alert. A service was stopped to prove each liveness monitor. Every alert described as proven on this page was proven that way, and the ones that were not are named as unproven rather than quietly counted.

The gap, stated plainly

Roughly forty runbooks exist. Around a dozen have now been substantially executed on real hosts, up from four a week ago.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. The quorum witness, which is what turns two servers into a cluster.
  2. The storage tiers, once the boot adapter arrives.
  3. Offsite backups, which the storage tier unblocks.
  4. The mail server, which gives this site a corrections route.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisorsBoth nodes online and monitored, not clusteredIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsIn progress
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsComplete

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerIn progress

Mail

ComponentPurposeStatus
StalwartMail serverIn progress
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
ForgejoGit hostingIn progress
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation, 8 hosts shippingComplete
WazuhSecurity monitoring, 8 agents reportingComplete
PrometheusMetrics collectionComplete
GrafanaDashboards and alertingComplete
Blackbox exporterCertificate expiry and endpoint checksComplete
Disk and hardware health monitoringRAID and SMART, both alert-provenComplete

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted, offsite copiesNot started
Restore proceduresTested restores, not just backupsNot started
Maintenance playbookPatching and routine operationsNot started

Twenty-three components complete. Eight in progress. Thirty-four not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-13 — 16 complete, 6 in progress, 43 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 13 August 2026.

Two rented gateways are live, both serving TLS, with the tunnel to the house up and carrying traffic. Single sign-on is running, which unblocks every service above it. Five virtual machines run on the single physical server. The hypervisor cluster and the storage tier remain unbuilt.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisor, single nodeOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Four hosts on the meshLive
Eight uptime monitors, alert-provenLive
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive
Identity provider (single sign-on)Live, current stable, backed up nightly

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.
  • Failover has never been drilled. The scripts exist and their alerting was just repaired — both referenced a token that does not exist, so a real failover would have changed every public DNS record and then crashed before telling anyone. Fixed, and still never exercised. A capability that has not been performed is a plan.

Not started

  • The hypervisor cluster. One node is online. High availability needs a second node plus the quorum witness; both machines exist and are racked, but neither has been powered on or had an operating system installed. Blocked on commissioning, not on procurement.
  • The storage tiers. Hardware on hand, not yet built out.
  • Eleven of the sixteen service VMs. Five are built: the home-side gateway, two DNS resolvers, the control node and the identity server.

Proven on test hardware, not yet in production

Not deployed, but no longer theoretical. The automation for the VM baseline, log shipping, and the security monitoring stack has been executed against throwaway virtual machines and verified on the running result rather than from the automation’s own self-report.

That distinction earned its keep. Running the code found defects that four separate document reviews had missed entirely — including one case where the stated mitigation for a known vulnerability did not actually hold, and several where a verification step was structurally incapable of failing.

The gap, stated plainly

Roughly forty runbooks exist. Four of them have been substantially executed on real hosts.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Commission the second hypervisor node and the witness.
  2. Rack and configure the storage tiers.
  3. The mail server, which gives this site a corrections route.
  4. Services on top of identity, now that identity exists.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisor, single nodeFirst server online, carrying unrelated productionIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsNot started
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceComplete
VaultwardenShared and machine secretsNot started

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerNot started

Mail

ComponentPurposeStatus
StalwartMail serverNot started
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
ForgejoGit hostingNot started
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation — proven on test hardware onlyIn progress
WazuhSecurity monitoring — proven on test hardware onlyIn progress
PrometheusMetrics collectionNot started
GrafanaDashboards and alertingNot started
Blackbox exporterCertificate expiry and endpoint checksNot started
Disk and hardware health monitoringPhysical-layer alertingNot started

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted, offsite copiesNot started
Restore proceduresTested restores, not just backupsNot started
Maintenance playbookPatching and routine operationsNot started

Sixteen components complete. Six in progress. Forty-three not started.

Every one of those forty-three has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-12 — 15 complete, 6 in progress, 44 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 12 August 2026.

Two rented gateways are live, both serving TLS. The tunnel from the public edge to the house is up and carrying traffic. Four virtual machines run on the single physical server. The hypervisor cluster and the storage tier remain unbuilt.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU)Live, hardened, serving TLS
First physical hypervisor, single nodeOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Four hosts on the meshLive
Eight uptime monitors, alert-provenLive
Rathole tunnel, both gatewaysLive, carrying traffic, proven end to end
Caddy on the standby gatewayLive, ACME proven on staging then production
External heartbeat watchdog, off-siteLive, negatives tested
Configuration-management control nodeLive

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The mail server. Design settled and mailbox automation working; the records and the service itself are not live yet. This is what the site’s corrections route is waiting on.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.

Known gaps, stated rather than counted as coverage

  • The external watchdog is half proven. An off-site heartbeat now exists on unrelated hosting, outside this infrastructure, and its failure modes were tested until they failed — stale, future-dated and missing state all return an error rather than a reassuring answer. The monitor watching it was proven by disabling the heartbeat outright rather than faking a response. The other direction is not proven. The off-site host also probes the gateway and alerts by email, and email is the path that matters when the gateway is down. Its arrival has never been confirmed. Counted as unproven — why that distinction matters.

Not started

  • The hypervisor cluster. One node is online. High availability needs a second node plus the quorum witness; both machines exist and are racked, but neither has been powered on or had an operating system installed. Blocked on commissioning, not on procurement.
  • The storage tiers. Hardware on hand, not yet built out.
  • Twelve of the sixteen service VMs. Four are built: the home-side gateway, two DNS resolvers and the control node.

Proven on test hardware, not yet in production

Not deployed, but no longer theoretical. The automation for the VM baseline, log shipping, and the security monitoring stack has been executed against throwaway virtual machines and verified on the running result rather than from the automation’s own self-report.

That distinction earned its keep. Running the code found defects that four separate document reviews had missed entirely — including one case where the stated mitigation for a known vulnerability did not actually hold, and several where a verification step was structurally incapable of failing.

The gap, stated plainly

Roughly forty runbooks exist. Four of them have been substantially executed on real hosts.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Commission the second hypervisor node and the witness.
  2. Rack and configure the storage tiers.
  3. The mail server, which gives this site a corrections route.
  4. Identity tier — because nothing above it can be built first.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay, both gateways, proven end to endComplete
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisor, single nodeFirst server online, carrying unrelated productionIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsNot started
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceNot started
VaultwardenShared and machine secretsNot started

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerNot started

Mail

ComponentPurposeStatus
StalwartMail serverNot started
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
ForgejoGit hostingNot started
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation — proven on test hardware onlyIn progress
WazuhSecurity monitoring — proven on test hardware onlyIn progress
PrometheusMetrics collectionNot started
GrafanaDashboards and alertingNot started
Blackbox exporterCertificate expiry and endpoint checksNot started
Disk and hardware health monitoringPhysical-layer alertingNot started

Operations

ComponentPurposeStatus
Ansible automationControl node live; baseline role proven against a real hostComplete
Offsite backup chainEncrypted, offsite copiesNot started
Restore proceduresTested restores, not just backupsNot started
Maintenance playbookPatching and routine operationsNot started

Fifteen components complete. Six in progress. Forty-four not started.

Every one of those forty-four has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-09 — 13 complete, 8 in progress, 44 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 9 August 2026.

Two rented gateways are live, the first physical server is online, and three virtual machines are running on it: a home-side gateway and two DNS resolvers. The hypervisor cluster and the storage tier remain unbuilt.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU), hardened and assertedLive, no traffic yet
First physical hypervisor, single nodeOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service
Home-side gateway VMLive, 8/8 acceptance checks
Two internal DNS resolversLive, split-horizon + DNSSEC
Four hosts on the meshLive
Five uptime monitors, alert-provenLive

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The tunnel back to home. No longer blocked — the home-side gateway VM exists. Remaining dependency is a shared secret, which is an operator handover rather than work. Three defects are already known in that runbook section, found by reading rather than by running.
  • A real fully-qualified name for the mail server. Now resolved on both gateways, but the mail domain and its forward records still need creating before the mail server runbook starts.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.

Known gaps, stated rather than counted as coverage

  • There is no external watchdog. Uptime monitoring runs on the same host as the services it watches, so if that host dies the monitor dies with it and no alert is sent. Worse than it sounds: when the monitoring container was stopped for eighteen minutes, every monitor reported UP throughout, because a heartbeat table records what was last seen — why that happens. The standby gateway now exists and could watch the primary, so this is no longer blocked on hardware. It is simply not built, and a restart policy is deliberately not standing in for it.

Not started

  • The hypervisor cluster. One node is online. High availability needs a second node plus the quorum witness; both machines exist and are racked, but neither has been powered on or had an operating system installed. Blocked on commissioning, not on procurement.
  • The storage tiers. Hardware on hand, not yet built out.
  • Thirteen of the sixteen service VMs. Three are built: the home-side gateway and two DNS resolvers.

Proven on test hardware, not yet in production

Not deployed, but no longer theoretical. The automation for the VM baseline, log shipping, and the security monitoring stack has been executed against throwaway virtual machines and verified on the running result rather than from the automation’s own self-report.

That distinction earned its keep. Running the code found defects that four separate document reviews had missed entirely — including one case where the stated mitigation for a known vulnerability did not actually hold, and several where a verification step was structurally incapable of failing.

The gap, stated plainly

Roughly forty runbooks exist. Eight sections of one of them have been fully executed on a real host.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Finish the remaining sections of the primary gateway runbook.
  2. Rack and configure the hypervisor cluster and storage.
  3. Complete the tunnel home, pending an operator-supplied secret.
  4. Identity tier — because nothing above it can be built first.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay — buildable, pending an operator-supplied secretIn progress
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationTemplate validated; three VMs built from itComplete
Physical hypervisor, single nodeFirst server online, carrying unrelated productionIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsNot started
UnboundInternal DNS, both resolvers liveComplete
Home-side gateway VMTerminates the tunnel inside the networkComplete

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceNot started
VaultwardenShared and machine secretsNot started

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerNot started

Mail

ComponentPurposeStatus
StalwartMail serverNot started
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
ForgejoGit hostingNot started
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation — proven on test hardware onlyIn progress
WazuhSecurity monitoring — proven on test hardware onlyIn progress
PrometheusMetrics collectionNot started
GrafanaDashboards and alertingNot started
Blackbox exporterCertificate expiry and endpoint checksNot started
Disk and hardware health monitoringPhysical-layer alertingNot started

Operations

ComponentPurposeStatus
Ansible automationConfiguration managementIn progress
Offsite backup chainEncrypted, offsite copiesNot started
Restore proceduresTested restores, not just backupsNot started
Maintenance playbookPatching and routine operationsNot started

Thirteen components complete. Eight in progress. Forty-four not started.

Every one of those forty-four has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-08-07 — 10 complete, 8 in progress, 47 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 7 August 2026.

Two rented gateways are live and the first physical server is online, carrying unrelated production alongside this build. The hypervisor cluster, the storage tier and every service VM remain unbuilt.

The design is still far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Past versions of this page are kept in the status archive, so the trajectory is visible rather than only the present state.

Live

Running and verified. The US gateway carries real traffic; the rest is built and asserted but not yet in service:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive
Standby gateway (EU), hardened and assertedLive, no traffic yet
First physical hypervisor, single nodeOnline, carrying unrelated production
10 Gb bonded uplink, failover provenBuilt, not yet in service

The mesh coordinator matters more than it looks: it is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

The edge is now a pair rather than a single host, which is what the failover design has always assumed. Failover itself is still not built.

In progress

  • The tunnel back to home. Blocked on the home-side gateway VM, which cannot exist until there is a hypervisor to run it on. Three defects are already known in that runbook section, found by reading rather than by running.
  • A real fully-qualified name for the mail server. Now resolved on both gateways, but the mail domain and its forward records still need creating before the mail server runbook starts.
  • VLAN plan applied to one host only. IDs were assigned on 7 August and a trunk proven on the new bonded link. Nothing else has been migrated, and adopting the plan fully means renumbering the live management network.
  • Blocklist status of the gateway addresses. Two lists report both addresses clean. One is unusable through a public resolver, and the rest need re-running from a host with working DNS — why the first answer was wrong.

Known gaps, stated rather than counted as coverage

  • There is no external watchdog. Uptime monitoring runs on the same host as the services it watches, so if that host dies the monitor dies with it and no alert is sent. The dashboard reads all-green until it reads nothing at all. The fix needs the second gateway, which does not exist yet. Recorded here as a gap rather than counted as monitoring — why that distinction matters.

Not started

  • The hypervisor cluster. One node is online. High availability needs a second node plus the quorum witness; both machines exist and are racked, but neither has been powered on or had an operating system installed. Blocked on commissioning, not on procurement.
  • The storage tiers. Hardware on hand, not yet built out.
  • Every service VM. Sixteen of them, zero built.

Proven on test hardware, not yet in production

Not deployed, but no longer theoretical. The automation for the VM baseline, log shipping, and the security monitoring stack has been executed against throwaway virtual machines and verified on the running result rather than from the automation’s own self-report.

That distinction earned its keep. Running the code found defects that four separate document reviews had missed entirely — including one case where the stated mitigation for a known vulnerability did not actually hold, and several where a verification step was structurally incapable of failing.

The gap, stated plainly

Roughly forty runbooks exist. Eight sections of one of them have been fully executed on a real host.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Finish the remaining sections of the primary gateway runbook.
  2. Rack and configure the hypervisor cluster and storage.
  3. Build the home-side gateway VM and complete the tunnel.
  4. Identity tier — because nothing above it can be built first.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay to the home networkNot started
Standby gateway (EU)Failover targetComplete
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationCommon build for every VM — proven on test hardware onlyIn progress
Physical hypervisor, single nodeFirst server online, carrying unrelated productionIn progress
10 Gb bonded uplinkBuilt and failover-proven, not yet in serviceIn progress
VLAN plan (802.1Q tags)Assigned and trunked on one hostIn progress
Proxmox clusterTwo-node with HA — blocked on a second nodeNot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsNot started
UnboundInternal DNS, redundant pairNot started
Home-side gateway VMTerminates the tunnel inside the networkNot started

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceNot started
VaultwardenShared and machine secretsNot started

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerNot started

Mail

ComponentPurposeStatus
StalwartMail serverNot started
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
ForgejoGit hostingNot started
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation — proven on test hardware onlyIn progress
WazuhSecurity monitoring — proven on test hardware onlyIn progress
PrometheusMetrics collectionNot started
GrafanaDashboards and alertingNot started
Blackbox exporterCertificate expiry and endpoint checksNot started
Disk and hardware health monitoringPhysical-layer alertingNot started

Operations

ComponentPurposeStatus
Ansible automationConfiguration managementIn progress
Offsite backup chainEncrypted, offsite copiesNot started
Restore proceduresTested restores, not just backupsNot started
Maintenance playbookPatching and routine operationsNot started

Ten components complete. Eight in progress. Forty-seven not started.

Every one of those forty-seven has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-07-29 — 9 complete, 5 in progress, 48 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 29 July 2026.

One machine out of roughly two dozen is doing real work. The design is far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Live

The US gateway VPS — tier 1, the public edge. Running and carrying real traffic:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive

That last group went in on 28 July and matters more than it looks: the mesh coordinator is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

In progress

  • The tunnel back to home. Blocked on the home-side gateway VM, which cannot exist until there is a hypervisor to run it on. Three defects are already known in that runbook section, found by reading rather than by running.
  • A real fully-qualified name for the mail server. The gateway’s hostname was set correctly, but the change left it without a proper domain-qualified name. Nothing currently cares. The mail server will, so it needs deciding before that runbook starts.

Known gaps, stated rather than counted as coverage

  • There is no external watchdog. Uptime monitoring runs on the same host as the services it watches, so if that host dies the monitor dies with it and no alert is sent. The dashboard reads all-green until it reads nothing at all. The fix needs the second gateway, which does not exist yet. Recorded here as a gap rather than counted as monitoring — why that distinction matters.

Not started

  • The standby gateway in Europe. Its runbook is written and carries every correction learned on the primary. Nothing has been provisioned.
  • All physical hardware. The hypervisors, the storage tier, and the witness node are specified in detail and not yet racked. The architecture describes an intended end state, not a room.
  • Every service VM. Sixteen of them, zero built.

Proven on test hardware, not yet in production

Not deployed, but no longer theoretical. The automation for the VM baseline, log shipping, and the security monitoring stack has been executed against throwaway virtual machines and verified on the running result rather than from the automation’s own self-report.

That distinction earned its keep. Running the code found defects that four separate document reviews had missed entirely — including one case where the stated mitigation for a known vulnerability did not actually hold, and several where a verification step was structurally incapable of failing.

The gap, stated plainly

Roughly forty runbooks exist. Eight sections of one of them have been fully executed on a real host.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Finish the remaining sections of the primary gateway runbook.
  2. Rack and configure the hypervisor cluster and storage.
  3. Build the home-side gateway VM and complete the tunnel.
  4. Identity tier — because nothing above it can be built first.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay to the home networkNot started
Standby gateway (EU)Failover targetNot started
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationCommon build for every VM — proven on test hardware onlyIn progress
Proxmox clusterTwo-node hypervisor with HANot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsNot started
UnboundInternal DNS, redundant pairNot started
Home-side gateway VMTerminates the tunnel inside the networkNot started

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceNot started
VaultwardenShared and machine secretsNot started

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerNot started

Mail

ComponentPurposeStatus
StalwartMail serverNot started
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
ForgejoGit hostingNot started
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation — proven on test hardware onlyIn progress
WazuhSecurity monitoring — proven on test hardware onlyIn progress
PrometheusMetrics collectionNot started
GrafanaDashboards and alertingNot started
Blackbox exporterCertificate expiry and endpoint checksNot started
Disk and hardware health monitoringPhysical-layer alertingNot started

Operations

ComponentPurposeStatus
Ansible automationConfiguration managementIn progress
Offsite backup chainEncrypted, offsite copiesNot started
Restore proceduresTested restores, not just backupsNot started
Maintenance playbookPatching and routine operationsNot started

Nine components complete. Five in progress. Forty-eight not started.

Every one of those forty-eight has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.

2026-07-28 — 9 complete, 5 in progress, 48 not started

Where the build actually is, as opposed to where the documentation says it should be. Updated as things change; last revised 28 July 2026.

The honest summary: one machine out of roughly two dozen is doing real work. The design is far ahead of the deployment, which is a problem this project is aware of and writing about rather than hiding.

Live

The US gateway VPS — tier 1, the public edge. Running and carrying real traffic:

ComponentState
Hardened SSH (key-only, non-default port)Live, verified
Firewall rulesetLive, verified across reboots
Caddy with real Let’s Encrypt certificatesLive, via DNS challenge
Edge intrusion detection and blockingLive, verified
Dynamic home-address tracking for tunnel rulesLive, verified
WireGuard (server side)Live, no peers enrolled yet
Mesh coordination serviceLive
Notification serviceLive
Uptime monitoringLive

That last group went in on 28 July and matters more than it looks: the mesh coordinator is a prerequisite for any VM joining the private network, so its being live unblocks the entire VM tier.

In progress

  • The tunnel back to home. Blocked on the home-side gateway VM, which cannot exist until there is a hypervisor to run it on. Three defects are already known in that runbook section, found by reading rather than by running.
  • A real fully-qualified name for the mail server. The gateway’s hostname was set correctly, but the change left it without a proper domain-qualified name. Nothing currently cares. The mail server will, so it needs deciding before that runbook starts.

Not started

  • The standby gateway in Europe. Its runbook is written and carries every correction learned on the primary. Nothing has been provisioned.
  • All physical hardware. The hypervisors, the storage tier, and the witness node are specified in detail and not yet racked. The architecture describes an intended end state, not a room.
  • Every service VM. Sixteen of them, zero built.

Proven on test hardware, not yet in production

Not deployed, but no longer theoretical. The automation for the VM baseline, log shipping, and the security monitoring stack has been executed against throwaway virtual machines and verified on the running result rather than from the automation’s own self-report.

That distinction earned its keep. Running the code found defects that four separate document reviews had missed entirely — including one case where the stated mitigation for a known vulnerability did not actually hold, and several where a verification step was structurally incapable of failing.

The gap, stated plainly

Roughly forty runbooks exist. Eight sections of one of them have been fully executed on a real host.

Every review pass, every new design document, and every good idea added to the backlog widens that gap. The gap — not any missing feature — is the actual risk to this project.

There is now a rule about it: new ideas get one paragraph in a backlog file and nothing else until the current tier is running. An idea earns its way out only if deferring it would make some decision expensive to reverse later. Two have qualified so far, and neither was a feature — both were addressing decisions that become costly once hardware is physically installed.

What comes next

  1. Finish the remaining sections of the primary gateway runbook.
  2. Rack and configure the hypervisor cluster and storage.
  3. Build the home-side gateway VM and complete the tunnel.
  4. Identity tier — because nothing above it can be built first.

Progress against this list is what the status updates cover.


Full component inventory

Everything in the design, with an honest status against each. The three values mean specific things:

  • Complete — deployed and verified on the running system, not inferred from a config file. This site has written at length about why that distinction matters.
  • In progress — partially deployed, or proven on throwaway test hardware but not yet running anywhere permanent.
  • Not started — designed and documented, nothing built.

Note how much sits in the third column. That gap is the point.

Public edge

ComponentPurposeStatus
Hardened SSHKey-only remote accessComplete
nftables firewallPacket filtering, survives rebootComplete
CaddyReverse proxy, TLS terminationComplete
Let’s Encrypt via DNS challengeReal certificates, including internal-only namesComplete
CrowdSecEdge intrusion detection and blockingComplete
Dynamic egress-IP trackingKeeps tunnel rules following a changing home IPComplete
HeadscaleMesh VPN coordinationComplete
ntfyPush notifications for alertsComplete
Uptime KumaExternal availability monitoringComplete
WireGuardVPN — server deployed, no peers enrolledIn progress
Rathole tunnelInbound relay to the home networkNot started
Standby gateway (EU)Failover targetNot started
DNS failover automationHealth-checked cutover, ~90s recoveryNot started

Core infrastructure

ComponentPurposeStatus
VM baseline automationCommon build for every VM — proven on test hardware onlyIn progress
Proxmox clusterTwo-node hypervisor with HANot started
Quorum witnessThird vote, breaks cluster tiesNot started
ZFS integrity tierChecksummed storage for small, precious dataNot started
BTRFS bulk tierCapacity storage for media and modelsNot started
UnboundInternal DNS, redundant pairNot started
Home-side gateway VMTerminates the tunnel inside the networkNot started

Identity and secrets

ComponentPurposeStatus
AuthentikSingle sign-on for every serviceNot started
VaultwardenShared and machine secretsNot started

Files and productivity

ComponentPurposeStatus
NextcloudFile sync and sharingNot started
OutlineNotes and internal wikiNot started
LinkwardenBookmark archiveNot started
SnappymailWebmail clientNot started
Paperless-ngxDocument scanning, OCR, archiveNot started
Firefly IIIPersonal financeNot started
n8nWorkflow automationNot started
FreshRSSNews readerNot started

Mail

ComponentPurposeStatus
StalwartMail serverNot started
rspamd + ClamAVSpam and malware filteringNot started
Mail backup and archive pipelineLong-term retentionNot started

Media

ComponentPurposeStatus
ImmichPhoto library with local MLNot started
JellyfinMedia streamingNot started
AudiobookshelfAudiobooks and podcastsNot started
Curated photo sharingSecond, public-facing photo instanceNot started

Communications

ComponentPurposeStatus
FreePBX / AsteriskPhone systemNot started
SMS bridgeMessaging integrationNot started
LAN-only SMS alert gatewayCritical alerts that survive an internet outageNot started
SMS query gatewayConversational status queriesNot started

Development

ComponentPurposeStatus
ForgejoGit hostingNot started
Woodpecker CIBuild automation on an isolated VLANNot started
ZotContainer registryNot started
Dependency-TrackSBOM and vulnerability trackingNot started
Syft + GrypeSBOM generation and scanningNot started
devpi / VerdaccioPackage proxiesNot started

AI

ComponentPurposeStatus
OllamaLocal language model runtimeNot started
Open WebUIChat interfaceNot started
LLM gatewayComplexity-tiered routing, per-service keysNot started
Voice pipelineLocal speech-to-text and text-to-speechNot started
GPU tierCard not yet selectedNot started

Observability and security

ComponentPurposeStatus
Loki + AlloyLog aggregation — proven on test hardware onlyIn progress
WazuhSecurity monitoring — proven on test hardware onlyIn progress
PrometheusMetrics collectionNot started
GrafanaDashboards and alertingNot started
Blackbox exporterCertificate expiry and endpoint checksNot started
Disk and hardware health monitoringPhysical-layer alertingNot started

Operations

ComponentPurposeStatus
Ansible automationConfiguration managementIn progress
Offsite backup chainEncrypted, offsite copiesNot started
Restore proceduresTested restores, not just backupsNot started
Maintenance playbookPatching and routine operationsNot started

Nine components complete. Five in progress. Forty-eight not started.

Every one of those forty-eight has a written runbook behind it. None of that counts for anything until it runs, which is the recurring theme here and the subject of the first post.