Compare commits

...
Author SHA1 Message Date
beatzaplentyandClaude Sonnet 4.6 cded77919d refactor: full repo sweep — variables, docs, and comment cleanup
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m31s
- variables.nix: switch to rec {}, extract giteaDomain/giteaRepoPath,
  extraAdminSshKeys, haLanNfsFqdn, tailscaleResolverIp, ports.dhcp,
  ports.dns; ipaServer now derives from homeDomain ref; section headers
- modules: use new vars throughout (pxe-boot, ts-dns-forwarder,
  cluster-config, configuration.nix, mount-pxe-images) — eval unchanged
- docs: delete ephemeral planning docs (AUDIT_REPORT, ha-network-audit,
  network-cutover); add docs/ha.md; drop migration reference table from
  ip-addressing.md; remove stale server example from beszel.md
- CLAUDE.md/README.md/AGENTS.md: fix build types (tailscale-router,
  ha-server, drop server); document scripts/ha/, scripts/ipa/, and
  all previously undocumented top-level and lib scripts

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-30 06:28:52 +10:00
15 changed files with 502 additions and 1162 deletions
+1 -1
View File
@@ -7,7 +7,7 @@ servers and workstation.
The flake exposes NixOS configurations named `<platform>-<buildtype>` The flake exposes NixOS configurations named `<platform>-<buildtype>`
(platforms: `linode`, `proxmox`, `lxc`, `baremetal`; build types: `minimal`, (platforms: `linode`, `proxmox`, `lxc`, `baremetal`; build types: `minimal`,
`nix-cache`, `server`, `docker`, `gui`, `pxe-boot`, `tailscale-router`, `nix-cache`, `docker`, `gui`, `pxe-boot`, `tailscale-router`,
`tor-relay`, `ha-server`), generated from `modules/platforms/*` and `tor-relay`, `ha-server`), generated from `modules/platforms/*` and
`modules/build-types/*` by the `mkTarget` function in `flake.nix`. Not every `modules/build-types/*` by the `mkTarget` function in `flake.nix`. Not every
combination is built — `pxe-boot` has no `linode` variant, `ha-server` only combination is built — `pxe-boot` has no `linode` variant, `ha-server` only
-150
View File
@@ -1,150 +0,0 @@
# Flake End-to-End Audit Report
**Date:** 2026-07-21
**Scope:** Full static lint/eval sweep + live build/deploy/interrogate/destroy testing of every `lxc-*` and `proxmox-*` flake target against `pve.sweet.home`, plus an audit of the operator's ability to manage the flake/secrets tooling.
**Branch:** `worktree-flake-e2e-audit` (this session's isolated worktree)
## Executive Summary
The flake itself is in good shape: `nixpkgs-fmt`, `statix`, and a full eval + dry-run build of every host and package are all clean. Every `lxc-*`/`proxmox-*` target's NixOS configuration builds successfully — no target has a broken derivation graph.
The issues found are **operational, not code-level**:
1. **pve.sweet.home is critically low on disk space** (91-95% full during this session) and cannot currently build the two largest closures (`gui`, `pxe-boot`) to completion — this actively blocks deploying/redeploying those hosts via the documented workflow.
2. **A real, reproducible secrets-decryption failure** was caught live: a stale cached container image (built before a same-day sops-key fix) boots with sshd never starting and every secret failing to decrypt. This is a **general hazard in `create-proxmox-resource.sh`'s "reuse the cached image if present" default**, not a one-off.
3. **sops key/anchor drift**: `proxmox-minimal` has a `.sops.yaml` recipient anchor with no corresponding private key anywhere in this environment; several `lxc-*`/`proxmox-*` targets have no sops registration at all yet.
4. One concrete script bug was found and **fixed in this session**: `create-proxmox-resource.sh` never enabled the QEMU guest agent channel on VMs it creates, despite the guest OS already running it.
5. A management-surface audit (of the operator's ability to run this repo day to day) found 5 process gaps, detailed below.
Nothing here required or received a `nixos-rebuild switch/boot/test`, `nixos-install`, or any disk-formatting command — all validation was `nix build`/`nix eval`, plus disposable `pct`/`qm` create-then-destroy cycles via the repo's own `create-proxmox-resource.sh`.
---
## 1. Static Analysis Results — all clean
`bash scripts/codex-maintenance.sh --full-check --dry-run` (whole-tree sweep, not just changed files):
| Check | Result |
|---|---|
| Secret grep | Clean — only the documented exceptions (installer's own hashed passwords, `access-tokens` comment references) |
| `nixpkgs-fmt --check` | 0/53 files would be reformatted |
| `statix` | No lint warnings |
| nix-cache host key drift check | Up to date |
| Full eval of every host's `system.build.toplevel` | All 19 `nixosConfigurations` targets evaluate cleanly |
| Dry-run build of every host + package | All succeed, no derivation errors |
No drift, no formatting issues, no lint findings anywhere in the tree.
---
## 2. Per-Target Test Results
Legend: **LIVE** = built on pve, `pct`/`qm` create → interrogated → destroyed. **BUILD-ONLY** = `nix build` validated the config (mostly `.config.system.build.toplevel`, occasionally `.tarball`), no resource created on pve.
| Target | Test type | Result | Notes |
|---|---|---|---|
| `lxc-docker` | BUILD-ONLY | ✅ PASS | Live redeploy skipped — CT105 is already running this identity in production; `--allow-duplicate-host` would have destroyed it. |
| `lxc-minimal` | **LIVE** | ✅ PASS (after retry) | First attempt reused a stale cached tarball predating a same-day sops-key commit → activation failed, sshd never started (see Finding #2). Redeployed with `--force-rebuild`: clean boot, `systemctl is-system-running` = `running`, secrets decrypted, sshd listening, users correct. |
| `lxc-nix-cache` | BUILD-ONLY | ✅ PASS (after retry) | Live redeploy skipped — CT101 is already running this identity. First local build attempt appeared to hang on a remote-builder handoff to nix-cache; killed and retried with `--builders ""` (local-only), succeeded. |
| `lxc-gui` | **LIVE (attempted)** | ⚠️ BLOCKED by pve disk space | Registered a fresh sops key (no prior registration existed), built successfully through the full NixOS system closure, then **failed packaging the tarball**: `No space left on device` on pve's root filesystem. Not a flake defect. |
| `lxc-pxe-boot` | **LIVE (attempted)** | ⚠️ BLOCKED by pve disk space | Same failure as `lxc-gui` — this target additionally builds a full nested installer/netboot image (`stage-installer-artifacts.nix`), making it similarly large. Failed with the same `No space left on device` error, immediately after the gui attempt had already consumed pve's remaining headroom. |
| `lxc-server` | BUILD-ONLY | ✅ PASS | No sops key registered yet; live deploy also would have hit `boot.zfs.extraPools` trying to import a real ZFS pool that doesn't exist in an isolated test container — an expected limitation of testing this build type outside its real hardware, not a bug. |
| `lxc-tailscale-exit-node` | BUILD-ONLY | ✅ PASS | No sops key registered yet. |
| `lxc-tor-relay` | BUILD-ONLY | ✅ PASS | Live redeploy skipped — CT106 already holds this identity in production. |
| `proxmox-docker` | BUILD-ONLY | ✅ PASS (after retry) | Live redeploy skipped — both CT105 *and* VM103 already hold `docker` identities. Combined `toplevel` + `diskoImagesScript` build crashed with a **Nix-internal assertion failure** (`worker.cc:360`) under this session's memory pressure (see Finding #6) — not a flake bug. Retried with `toplevel` alone: clean. |
| `proxmox-minimal` | **LIVE (attempted)** | ⚠️ BLOCKED by key drift → BUILD-ONLY | `.sops.yaml` has a registered `&proxmox-minimal` anchor but **no corresponding private key exists anywhere in this environment** — the script correctly refused to generate a mismatched replacement. Fell back to `toplevel` build: ✅ PASS. |
| `proxmox-nix-cache` | BUILD-ONLY | ✅ PASS | No sops key registered yet. |
| `proxmox-gui` | BUILD-ONLY | ⚠️ Killed after ~40min (resource-limited) | This session's local build machine has only 2GB RAM; swap filled completely (2.0/2.0GB) and the build stalled, so it was killed rather than risk destabilizing the session further. **Not a flake defect** — the equivalent `gui` NixOS configuration already proved fully buildable during the `lxc-gui` live attempt above (it built the entire system closure successfully and only failed at the pve-side tarball-packaging step due to disk space, not the config). |
| `proxmox-pxe-boot` | BUILD-ONLY | ⚠️ Killed after ~35min (resource-limited) | Was deep into building the nested installer's kernel initrd (this build type bundles a full netboot installer image via `stage-installer-artifacts.nix`) when killed to keep the audit moving. **Not a flake defect** — this target's own module logic was already effectively validated via the earlier *live* pve deploy attempt (`lxc-pxe-boot` above), which built the complete image and only failed at the final tarball-packaging step due to pve's disk space (Finding 1). |
| `proxmox-server` | BUILD-ONLY | ✅ PASS | No sops key registered yet; same ZFS-pool caveat as `lxc-server` would apply to a live deploy. |
| `proxmox-tailscale-exit-node` | BUILD-ONLY | ✅ PASS | No sops key registered yet. |
**Not tested at all:** `linode-*` targets (not deployable to Proxmox) and `installer` (not a normal host) — both were still covered by the static eval/dry-run-build sweep above.
---
## 3. Findings, Ranked by Severity
### Finding 1 — pve.sweet.home is critically low on disk space (blocks real deployments)
At session start: `/dev/mapper/pve-root` was **95% full, 5.3GB free** (of 94GB). After two failed large builds it recovered slightly to **91% full, 8.2GB free** (nix cleans up its own failed-build scratch space). `/nix/store` alone is 26GB; `nix-store --gc --print-dead` reports **zero** reclaimable garbage — everything currently in the store is a live GC root, so `nix-collect-garbage` won't help without first removing old roots.
**Why it matters:** `create-proxmox-resource.sh` builds every VM/CT image **directly on pve**, not on a build machine and transferred over. With <10GB headroom, any closure approaching a few GB (the `gui` build type: full Cinnamon desktop + Firefox + LibreOffice + GIMP + VS Code + xrdp; the `pxe-boot` build type: nginx/atftpd *plus* an entire nested installer/netboot image) cannot currently be built there at all. Both `lxc-gui` and `lxc-pxe-boot` failed live with `No space left on device` during this audit.
**Recommended action:** Expand `pve-root`'s LV, or free space by pruning old container templates in `/var/lib/vz/template/cache` (1.5GB) / old backups in `/var/lib/vz/dump` (306MB) / auditing what's pinning 26GB of `/nix/store` as live GC roots (likely `result-*` symlinks — see below). This is real production disk state; **not something this session touched or fixed** — it needs the operator's judgment on what's safe to remove.
**Secondary, smaller finding:** every `create-proxmox-resource.sh` run leaves a `result-<target>` symlink in the node's repo checkout as a permanent GC root (`ls /root/nixos/result-*` on pve showed 3 from this session alone: `lxc-docker`, `lxc-minimal`, `lxc-nix-cache`). These accumulate forever and pin their entire closures in the store. Consider having the script clean up its own `result-*` link after staging the built artifact (or use a temp `--out-link` under `/tmp`), so `nix-collect-garbage` can actually reclaim old build outputs.
### Finding 2 — Stale cached images can silently ship broken secrets (reproduced live)
`create-proxmox-resource.sh`'s default behavior is: if the node already has `<target>.tar.xz`/`.raw` staged, **reuse it** — only `--force-rebuild` forces a fresh build. This session hit exactly the failure mode `docs/auto-installer.md` already warns about: `lxc-minimal`'s cached tarball (built 2026-07-20T15:57Z) predated a same-day sops-key fix commit (2026-07-20T17:49Z, "clean up in ailse 3"). The deployed container booted with:
```
sops-install-secrets: failed to decrypt '.../common.yaml': Error getting data key: 0 successful groups required, got 0
Activation script snippet 'setupSecrets' failed (1)
```
— every secret permanently failed to decrypt, `sshd` never started (though the container otherwise looked "running"). This was **not a code bug**: the currently-committed `secrets/common.yaml` decrypts fine for that host's key when checked independently; the *cached artifact on pve* simply reflected an older commit's ciphertext. Redeploying with `--force-rebuild` fixed it immediately.
**Why it matters:** this is silent and easy to trigger by accident — any operator who redeploys a host without remembering `--force-rebuild` after a secrets change gets a container that looks like it started (`pct start` succeeds, `pct status` = running) but is completely inaccessible.
**Recommended action:** Have `create-proxmox-resource.sh` compare the cached image's build timestamp (or embed the source commit hash in the staged filename) against current HEAD, and warn (or refuse without `--force-rebuild`) if they differ — rather than silently trusting presence alone.
### Finding 3 — sops key/anchor drift
Two concrete instances hit live during this session:
- **`proxmox-minimal`**: `.sops.yaml` already has a registered `&proxmox-minimal` age recipient, but this environment's `host-keys/` directory has no corresponding private key file. `sync-host-keys.sh` correctly refused to generate a replacement (it would silently mismatch whatever's already registered/deployed) — but this means **no environment currently has this host's private key**, unless it exists on some other machine that was never backed up here.
- **`lxc-gui`**, and by the same logic `lxc-server`/`lxc-tailscale-exit-node`/most `proxmox-*` targets, have **no sops registration at all yet** — expected for undeployed hosts per `docs/auto-installer.md`, but this session's live-testing needed to register `lxc-gui`'s key on the fly, which immediately hit **Finding 3b**: registering a key locally does nothing for pve's build until it's pushed to `origin/main` (pve builds via `git pull`, not from this uncommitted worktree). This is exactly gap #4 the management-surface audit (below) already flagged in the abstract — this session hit it concretely.
**Recommended action:** for `proxmox-minimal`, decide whether to regenerate its key (destroying old-key decrypt access, if anything still holds it) or track down wherever the original private key lives and back it up here. For the general pattern, see the management-surface audit's recommendation to pre-flight-check key registration before building.
### Finding 4 — QEMU guest agent never wired up (found and fixed this session)
`modules/common/configuration.nix:44` sets `services.qemuGuest.enable = true` on every host — the guest-side agent daemon is correctly enabled everywhere. But `scripts/proxmox/create-proxmox-resource.sh`'s `qm create` call never passed `--agent 1`, so **Proxmox never created the virtio-serial channel** the agent needs. Every `proxmox-*` VM this script ever created was silently missing `qm guest exec`/IP-address reporting in the Proxmox UI, despite the guest daemon actually running.
**Status: fixed in this session's worktree** (`scripts/proxmox/create-proxmox-resource.sh`, `qm create` now includes `--agent enabled=1`) — see the diff, included in the PR from this session.
### Finding 5 — Orphaned container on pve (CT102)
`pve.sweet.home` has a stopped LXC container, **VMID 102**, with an essentially empty config (`lock: create` and nothing else — no hostname, no rootfs, no network) — the leftover of a `pct create` that started and never finished. It predates this session (not created by any of this audit's activity) and wasn't touched. **Recommend the operator confirm it's abandoned and remove it** (`pct destroy 102 --purge 1`) — left as-is it may be someone's genuine in-progress work, so it wasn't assumed safe to delete autonomously.
### Finding 6 — Nix-internal crash under memory pressure (tooling, not flake)
Building `proxmox-docker`'s `toplevel` and `diskoImagesScript` together crashed with a Nix-internal assertion failure (`Assertion '!awake.empty()' failed ... worker.cc:360`, a known class of bug in Nix's multi-goal build scheduler) while this session's 2GB-RAM build container was under heavy swap pressure (1.8-2.0/2GB swap in use) from a separate concurrent build. Retrying the same target alone (no concurrency) succeeded cleanly. **Not a flake defect** — purely an artifact of this session's constrained build environment; noted for completeness since it looked alarming in isolation.
### Finding 7 — Management-surface audit: 5 operability gaps
A focused audit of "can the operator actually run this repo day to day" (flake, home-manager, sops, related scripts) found:
1. **No documented recovery path if the `&admin` sops age key is lost without a backup.** `scripts/secrets/backup-admin-key.sh` exists and works but is referenced nowhere in `README.md`/`docs/` — no forcing function ensures a backup was ever taken. `rotate-admin-key.sh` requires the *old* key to re-key; there's no bootstrap-from-nothing path documented (the real fallback — deriving an age identity from any still-live host's own SSH key — isn't written down anywhere).
2. **home-manager has no standalone iteration path.** It's wired only inside `nixosConfigurations` (`flake.nix`) — no `homeConfigurations` output. The fastest real shortcut (`nix build .#nixosConfigurations.<target>.config.home-manager.users.nixos.home.activationPackage`) isn't documented anywhere, so the practical workflow is a full host rebuild to test one HM tweak.
3. **Gitea's flake-lock-update workflow pushes straight to `main` with no pre-merge validation.** `.gitea/workflows/update-flake-lock.yml` commits and pushes `nix flake update`'s result directly; `codex-maintenance.sh` only runs *after*, on the resulting push — a genuinely broken lockfile bump lands on `main` before anything catches it. (The GitHub-side workflow is safer — PR-based — but has the opposite gap: nothing alerts if the PR sits unmerged.)
4. **No pre-flight check that a build target has a registered sops key before building it.** `docs/auto-installer.md` documents the failure mode (silent, total secrets-decrypt failure) but nothing in `create-proxmox-resource.sh` refuses to proceed when it's about to build a target with no `.sops.yaml` anchor — it's on the operator to remember. This session's `lxc-gui` test hit close to this exact gap (needed the key added on the fly, mid-session).
5. **`vars.remoteBuilderAuthorizedKeys` has the same drift risk as `vars.nixCacheHostKey`, but no checker script.** `sync-nix-cache-host-key.sh --check` guards the latter; the former (and `vars.pxeServerIp`/`vars.pbsIp`) has no equivalent — a rotated/revoked client key just silently stops working with no diagnostic pointing back here.
---
## 4. Action Plan (priority order)
1. **Free up disk space on pve.sweet.home** (or expand `pve-root`). Blocking: `lxc-gui`, `proxmox-gui`, `lxc-pxe-boot`, `proxmox-pxe-boot` cannot currently be built/redeployed on this node at all.
2. **Decide on `proxmox-minimal`'s orphaned sops key**: locate the original private key and back it up here, or accept regenerating it (breaks decrypt access for whoever/whatever currently holds the old one).
3. **Merge this session's PR** (see below) to get the `--agent 1` fix and `lxc-gui`'s new sops registration onto `main` — required before `lxc-gui` can be live-redeployed with working secrets.
4. **Add a staleness guard to `create-proxmox-resource.sh`'s cache-reuse path** (Finding 2) — highest-leverage fix, since it silently produces a broken-but-"running" host.
5. **Add a pre-flight sops-anchor check to `create-proxmox-resource.sh`** (management-surface gap #4) — same root cause class as #4 above, catch it before building instead of at first boot.
6. Investigate/clean up **CT102** on pve (Finding 5) — confirm abandoned, then remove.
7. Document `backup-admin-key.sh` in `README.md`'s Security Notes and add the live-host-key bootstrap-recovery procedure to `docs/` (management-surface gap #1).
8. Add pre-push validation to the Gitea flake-lock-update workflow (management-surface gap #3).
9. Lower-priority: document the home-manager `activationPackage` shortcut (gap #2); extend `sync-nix-cache-host-key.sh`'s drift-check pattern to `remoteBuilderAuthorizedKeys` (gap #5).
10. Follow-up session: finish build-validating `proxmox-gui` and `proxmox-pxe-boot` (both killed here after 35-40min on this session's 2GB-RAM machine — not failures, just unfinished) once pve has headroom (item 1) — ideally from a machine with more RAM. `proxmox-server` and `proxmox-tailscale-exit-node` already passed build-only validation in this session, no follow-up needed.
---
## 5. Uncommitted Changes From This Session
This worktree (`worktree-flake-e2e-audit`) currently has:
- `scripts/proxmox/create-proxmox-resource.sh` — the `--agent enabled=1` fix (Finding 4).
- `.sops.yaml` / `secrets/common.yaml``lxc-gui`'s new age key registered as a recipient (generated live during this session's testing).
Per this session's standard workflow, these will be committed, pushed, and opened as a draft PR rather than pushed to `main` directly — merging it is the operator's call, and is also **prerequisite to live-redeploying `lxc-gui` successfully** (its build will keep hitting the sops-staleness failure from Finding 2 on pve until this registration is on `origin/main`).
+101 -9
View File
@@ -252,6 +252,14 @@ instead of copying it.
silent skip rather than a failure) only reports drift; the no-flags form silent skip rather than a failure) only reports drift; the no-flags form
updates both files in place. Declarative clients still need a rebuild to updates both files in place. Declarative clients still need a rebuild to
pick up the fix. pick up the fix.
- `scripts/secrets/push-host-keys.sh [--all | <target>] [--dry-run]
[--skip-git-check]` — pushes newly-generated SSH host keys from
`host-keys/` to already-running NixOS hosts, so they can decrypt sops
secrets after a rebuild following `sync-host-keys.sh
--regenerate-all-keys`. Verifies that `.sops.yaml` and `secrets/*.yaml`
are committed and pushed to the remote first (hosts rebuild from the
remote Gitea flake, so recipient changes must land there before any key
push).
### `scripts/proxmox/` ### `scripts/proxmox/`
@@ -273,6 +281,16 @@ instead of copying it.
failure just falls back to building from source / `cache.nixos.org`) so failure just falls back to building from source / `cache.nixos.org`) so
the node substitutes from and can offload builds to nix-cache on every the node substitutes from and can offload builds to nix-cache on every
subsequent run, not just this one. subsequent run, not just this one.
- `scripts/proxmox/clone-pve1-to-pve-test.sh <vmid> [--new-vmid <id>]
[--mode snapshot|suspend|stop] [--dry-run]` — ad-hoc clone of a single
VM or CT from pve1 (production) to pve-test (sandbox) via vzdump +
qmrestore/pct restore. Streams the archive directly between nodes (no
local staging copy). Always restores with `--unique 1` (fresh MAC
addresses) since the original is still running on the LAN. Cleans up
the vzdump archive from both nodes after a successful restore. The
script's own default is pve1 → pve-test, matching CLAUDE.md's policy
(unlike `create-proxmox-resource.sh`, which defaults to production for
the operator's own unqualified use).
- `scripts/proxmox/configure-nix-cache-client.sh [--dry-run] - `scripts/proxmox/configure-nix-cache-client.sh [--dry-run]
[--no-remote-builder] [--no-restart]` — the non-NixOS equivalent of [--no-remote-builder] [--no-restart]` — the non-NixOS equivalent of
`modules/nix-cache/client.nix`/`remote-builder-client.nix`, for a plain `modules/nix-cache/client.nix`/`remote-builder-client.nix`, for a plain
@@ -288,6 +306,50 @@ instead of copying it.
marked block rather than duplicating it); restarts `nix-daemon` by marked block rather than duplicating it); restarts `nix-daemon` by
default so the change takes effect immediately. default so the change takes effect immediately.
### `scripts/ha/`
HA cluster lifecycle and operational scripts. All mutate real cluster state
when run for real — always run against pve-test first unless the operator
explicitly targets pve1.
- `scripts/ha/deploy.sh [--skip-*] [--destroy] [--dry-run]` — full
lifecycle manager: phases through bridge creation, key sync, VM creation
(via `create-proxmox-resource.sh`), NIC/disk attachment, and cluster
initialisation. `--destroy` tears it back down. Safe to rerun
idempotently; each phase can be individually skipped.
- `scripts/ha/cluster-init.sh` — one-time cluster bootstrap run **as root
on ha-server-1** after both VMs are booted. Generates/distributes the
Corosync authkey, initialises DRBD metadata, creates XFS on `/dev/drbd0`,
configures LIO iSCSI, and registers all Pacemaker resources (DRBD → XFS
→ iSCSI → NFS → VIPs).
- `scripts/ha/health.sh` — read-only cluster health snapshot: SSH
reachability, quorum, DRBD state, Pacemaker resources, and VIP port
reachability. Safe to run from the workstation at any time.
- `scripts/ha/failover.sh [--to node1|node2] [--force] [--timeout <s>]
[--dry-run]` — graceful failover by putting the active node into
Pacemaker standby and waiting for resources to appear on the target.
- `scripts/ha/acceptance-tests.sh` — T1T7 acceptance tests (failover,
NFS/iSCSI connectivity, DRBD sync, etc.) that must all pass before the
cluster is considered production-ready.
- `scripts/ha/resize-data-disk.sh --size +NNg [--force] [--dry-run]` —
online data-disk resize: `qm resize` on both VMs, guest block-device
rescan, `drbdadm resize`, `xfs_growfs`. No downtime required.
- `scripts/ha/cluster-enable-stonith.sh` — enables the `fence_pve_ssh`
STONITH resource after the fence SSH key is deployed to both nodes and
authorised on the Proxmox host. Run once after `cluster-init.sh`.
- `scripts/ha/fence-pve-ssh.py` — Python STONITH fence agent for Pacemaker.
Deploy to `/etc/pacemaker/fence_pve_ssh` on both HA nodes (`chmod +x`).
SSHes to the Proxmox host and runs `qm stop/start <vmid>`.
### `scripts/ipa/`
- `scripts/ipa/create-nixos-ipa-host-account.sh [options] <hostname>` —
adds a NixOS host to the FreeIPA domain and produces a sops-encrypted
keytab at `secrets/<hostname>.keytab`, ready for `modules/ipa/client.nix`.
Replaces three error-prone manual steps: `ipa host-add`, `ipa-getkeytab`
(run on the DC, SCP'd back), and `sops encrypt` in the correct location
(must be at `secrets/<hostname>.keytab` for the creation rule to match).
### `scripts/lib/` ### `scripts/lib/`
Sourced by the scripts above, never run directly: Sourced by the scripts above, never run directly:
@@ -297,6 +359,15 @@ Sourced by the scripts above, never run directly:
`create-proxmox-resource.sh` runs over SSH. `create-proxmox-resource.sh` runs over SSH.
- `nix-eval.sh` — `NIX_EVAL_FLAGS` plus `list_flake_targets`/ - `nix-eval.sh` — `NIX_EVAL_FLAGS` plus `list_flake_targets`/
`flake_target_hostname` flake-introspection helpers. `flake_target_hostname` flake-introspection helpers.
- `nix-parallel.sh` — `run_nix_parallel`: fans out independent `nix eval`/
`nix build --dry-run` calls across up to `NIX_PARALLEL_JOBS` processes,
capped by available memory (~1 GB/job) rather than raw `nproc` to avoid
OOM on constrained CI runners. Used by `codex-maintenance.sh`.
- `clan-vars.sh` — helpers for reading/writing SSH host keys stored as clan
vars (`vars/per-machine/<target>/openssh/`, sops-encrypted) instead of
the gitignored `host-keys/` directory. Sourced by
`create-proxmox-resource.sh` and `sync-host-keys.sh`; depends on
`sops-age.sh` and `ssh-host-keys.sh` being sourced first.
- `ssh-host-keys.sh` — `generate_host_ed25519_key`/`ssh_pubkey_to_age`, - `ssh-host-keys.sh` — `generate_host_ed25519_key`/`ssh_pubkey_to_age`,
shared by `sync-host-keys.sh` and `prepare-host-key.sh`. shared by `sync-host-keys.sh` and `prepare-host-key.sh`.
- `sops-age.sh` — `age_pubkey_from_identity_file`/`sops_yaml_admin_pubkey`/ - `sops-age.sh` — `age_pubkey_from_identity_file`/`sops_yaml_admin_pubkey`/
@@ -316,6 +387,17 @@ Sourced by the scripts above, never run directly:
default cores/memory, `NIX_CACHE_HOST`, `LAN_DOMAIN`) sourced by default cores/memory, `NIX_CACHE_HOST`, `LAN_DOMAIN`) sourced by
`create-proxmox-resource.sh` and `scripts/installer/auto-install.sh`. Add `create-proxmox-resource.sh` and `scripts/installer/auto-install.sh`. Add
new cross-script config here instead of duplicating it per-script. new cross-script config here instead of duplicating it per-script.
- `scripts/recover-hosts.sh [<hostname> ...]` — fixes sops/SSH-key/GitHub-token
issues on deployed NixOS hosts and triggers a `Switch-nix` rebuild on each.
With no args discovers every known hostname; with args checks only those.
Fixes applied automatically (prompts before rebuilding): SSH host key drift
(restores the registered key) and stale GitHub access tokens (empties the
rendered `nix-github-token.conf` so Nix falls back to unauthenticated requests
until sops-nix re-renders the correct token after the next successful rebuild).
- `scripts/gc-hosts.sh [--dry-run]` — runs `nix-collect-garbage -d` on all live
NixOS hosts (workstation first, then pve1, then all Proxmox guests). Excludes
`nix-cache` (gc-ing the shared binary cache evicts store paths other hosts
depend on). Uses passwordless sudo where available; falls back to user-level gc.
- `scripts/bump-nixpkgs-release.sh` — bumps `flake.nix`'s `nixpkgs.url`/ - `scripts/bump-nixpkgs-release.sh` — bumps `flake.nix`'s `nixpkgs.url`/
`home-manager.url` in place. Exists because flake input URLs can't `home-manager.url` in place. Exists because flake input URLs can't
reference `variables.nix` (confirmed empirically — `nix flake metadata` reference `variables.nix` (confirmed empirically — `nix flake metadata`
@@ -354,12 +436,13 @@ nixosSystem {
``` ```
Platforms: `linode`, `proxmox`, `lxc`, `baremetal`. Build types: `minimal`, Platforms: `linode`, `proxmox`, `lxc`, `baremetal`. Build types: `minimal`,
`nix-cache`, `server`, `docker`, `gui`, `pxe-boot`, `tailscale-exit-node`, `nix-cache`, `docker`, `gui`, `pxe-boot`, `tailscale-router`, `tor-relay`,
`tor-relay`. Not every combination is built — e.g. `pxe-boot` has no `linode` `ha-server`. Not every combination is built — e.g. `pxe-boot` has no `linode`
variant (PXE/DHCP/TFTP need LAN L2 adjacency a Linode VPS doesn't have), variant (PXE/DHCP/TFTP need LAN L2 adjacency a Linode VPS doesn't have),
`tor-relay` currently only exists as `lxc-tor-relay`, and `baremetal` `tor-relay` only exists as `lxc-tor-relay`, `ha-server` only exists as
currently only exists as `baremetal-gui` (the real gui-host hardware — `proxmox-ha-server-{1,2}`, and `baremetal` only exists as `baremetal-gui`
see `hosts/nixos/host.nix` and `modules/platforms/baremetal.nix`). Treat (the real gui-host hardware — see `hosts/nixos/host.nix` and
`modules/platforms/baremetal.nix`). Treat
`flake.nix`'s `flake.nix`'s
`generatedTargets` as the source `generatedTargets` as the source
of truth for which hosts exist — `README.md`, `AGENTS.md`, of truth for which hosts exist — `README.md`, `AGENTS.md`,
@@ -391,7 +474,7 @@ removing a host.
`vzdump` backup-archive metadata this doesn't have), no install step — `vzdump` backup-archive metadata this doesn't have), no install step —
see `docs/auto-installer.md`. see `docs/auto-installer.md`.
- `modules/build-types/*.nix` — what a system is for: - `modules/build-types/*.nix` — what a system is for:
minimal/server/docker/gui/pxe-boot/nix-cache/tailscale-exit-node/tor-relay. minimal/docker/gui/pxe-boot/nix-cache/tailscale-router/tor-relay/ha-server.
- `modules/common/configuration.nix` — base NixOS config imported by every - `modules/common/configuration.nix` — base NixOS config imported by every
host: locale, users, nix settings, git. host: locale, users, nix settings, git.
- `modules/common/home.nix` / `hosts/nixos/home.nix` — Home Manager config for - `modules/common/home.nix` / `hosts/nixos/home.nix` — Home Manager config for
@@ -417,7 +500,7 @@ removing a host.
`modules/platforms/baremetal.nix` also imports `modules/platforms/baremetal.nix` also imports
`modules/services/zfs/enable-service.nix` for this (the `zfs_unstable` `modules/services/zfs/enable-service.nix` for this (the `zfs_unstable`
package, autoScrub/autoSnapshot/trim) — the only other importer today is package, autoScrub/autoSnapshot/trim) — the only other importer today is
`server`'s NFS data pool, an unrelated non-root ZFS use. `ha-server`'s NFS data pool, an unrelated non-root ZFS use.
- `modules/boot/efi.nix` — systemd-boot + EFI vars, paired with the disko module. - `modules/boot/efi.nix` — systemd-boot + EFI vars, paired with the disko module.
- `modules/installer/` — the auto-installer environment (ISO, also served as - `modules/installer/` — the auto-installer environment (ISO, also served as
PXE netboot): `common.nix` (shared config + the generated PXE netboot): `common.nix` (shared config + the generated
@@ -431,14 +514,21 @@ removing a host.
substituter + SSH remote-builder wiring; see `docs/nix-cache.md` for the substituter + SSH remote-builder wiring; see `docs/nix-cache.md` for the
full design (per-host local stores, no shared `/nix/store`, and how the full design (per-host local stores, no shared `/nix/store`, and how the
`nixremote` signing/SSH keys fit together). `nixremote` signing/SSH keys fit together).
- `modules/ha/` — HA cluster NixOS modules: `cluster-config.nix` (DRBD,
Corosync, Pacemaker, firewall rules, cluster-wide NFS/iSCSI port
authorisation — shared by both ha-server nodes), `pacemaker-stack.nix`
(Pacemaker + Corosync service enablement), and supporting modules. See
`docs/ha.md` for the cluster operational guide.
- `modules/ipa/client.nix` — FreeIPA client enrollment: sssd, Kerberos keytab,
and IPA host registration; imported by every real host via
`modules/common/configuration.nix`.
- `modules/beszel/enable-agent.nix` — enables beszel-agent, sets `HUB_URL`, - `modules/beszel/enable-agent.nix` — enables beszel-agent, sets `HUB_URL`,
fixes the upstream `StateDirectory` bug, and wires the universal fixes the upstream `StateDirectory` bug, and wires the universal
`beszel-token` sops secret (from `secrets/common.yaml`) into the agent's `beszel-token` sops secret (from `secrets/common.yaml`) into the agent's
`environmentFile`; see `docs/beszel.md` for the full setup guide. `environmentFile`; see `docs/beszel.md` for the full setup guide.
- `modules/tailscale/`, `modules/docker/`, `modules/networking/`, - `modules/tailscale/`, `modules/docker/`, `modules/networking/`,
`modules/traefik/`, `modules/tor/`, `modules/services/*` — single-purpose, `modules/traefik/`, `modules/tor/`, `modules/services/*` — single-purpose,
single-host single-host feature modules (e.g. `docker/enable-service.nix`,
feature modules (e.g. `docker/enable-service.nix`,
`services/zfs/enable-service.nix`). Grep `modules/build-types/*.nix` for `services/zfs/enable-service.nix`). Grep `modules/build-types/*.nix` for
each build type's `imports` list to see which modules apply where. each build type's `imports` list to see which modules apply where.
@@ -462,3 +552,5 @@ duplicating config.
- `docs/flake-lock-automation.md` — how `flake.lock` updates flow through CI - `docs/flake-lock-automation.md` — how `flake.lock` updates flow through CI
(scheduled `nix flake update` PR + host-eval-on-PR workflow) and why hosts (scheduled `nix flake update` PR + host-eval-on-PR workflow) and why hosts
should track the committed lock file rather than `nixos-rebuild --upgrade-all`. should track the committed lock file rather than `nixos-rebuild --upgrade-all`.
- `docs/ha.md` — HA file-server cluster: DRBD + XFS + LIO iSCSI + NFS managed
by Corosync + Pacemaker; network topology; lifecycle scripts in `scripts/ha/`.
+2 -3
View File
@@ -9,8 +9,8 @@ Targets are named `<platform>-<buildtype>`, generated from two orthogonal
pieces composed in `flake.nix`: pieces composed in `flake.nix`:
- **Platforms** (what it runs on): `linode`, `proxmox`, `lxc`, `baremetal` - **Platforms** (what it runs on): `linode`, `proxmox`, `lxc`, `baremetal`
- **Build types** (what it's for): `minimal`, `nix-cache`, `server`, `docker`, - **Build types** (what it's for): `minimal`, `nix-cache`, `docker`, `gui`,
`gui`, `pxe-boot`, `tailscale-router`, `tor-relay`, `ha-server` `pxe-boot`, `tailscale-router`, `tor-relay`, `ha-server`
Not every combination exists — `pxe-boot` has no `linode` variant, since Not every combination exists — `pxe-boot` has no `linode` variant, since
PXE/DHCP/TFTP need LAN L2 adjacency that a Linode VPS doesn't have, PXE/DHCP/TFTP need LAN L2 adjacency that a Linode VPS doesn't have,
@@ -24,7 +24,6 @@ hardware). The full list:
| `proxmox-minimal` | Minimal NixOS host profile on Proxmox — previously the flat `nix-minimal` target | | `proxmox-minimal` | Minimal NixOS host profile on Proxmox — previously the flat `nix-minimal` target |
| `lxc-minimal` | Minimal NixOS host profile in a Proxmox LXC container | | `lxc-minimal` | Minimal NixOS host profile in a Proxmox LXC container |
| `linode-nix-cache` / `proxmox-nix-cache` / `lxc-nix-cache` | Local Nix binary cache and remote builder — previously the flat `nix-cache` target | | `linode-nix-cache` / `proxmox-nix-cache` / `lxc-nix-cache` | Local Nix binary cache and remote builder — previously the flat `nix-cache` target |
| `linode-server` / `proxmox-server` / `lxc-server` | Storage, NFS, backup, and monitoring exporter host — previously the flat `server` target |
| `linode-docker` / `proxmox-docker` / `lxc-docker` | Docker host for the main container stack — previously the flat `docker` target | | `linode-docker` / `proxmox-docker` / `lxc-docker` | Docker host for the main container stack — previously the flat `docker` target |
| `linode-gui` / `proxmox-gui` / `lxc-gui` | Cinnamon desktop workstation — previously the flat `nixos` target | | `linode-gui` / `proxmox-gui` / `lxc-gui` | Cinnamon desktop workstation — previously the flat `nixos` target |
| `baremetal-gui` | Same Cinnamon desktop workstation, on the real gui-host hardware — ZFS RAID0 root, systemd-boot | | `baremetal-gui` | Same Cinnamon desktop workstation, on the real gui-host hardware — ZFS RAID0 root, systemd-boot |
-9
View File
@@ -82,15 +82,6 @@ services.beszel.agent.environment = {
}; };
``` ```
The `server` host uses this to expose its ZFS data pool:
```nix
services.beszel.agent.environment = {
EXTRA_FILESYSTEMS = "${vars.storageRoot}/${vars.nfsShares.dockerVolumes.subpath}";
LOG_LEVEL = "debug";
};
```
--- ---
## Optional: monitoring Docker containers ## Optional: monitoring Docker containers
-261
View File
@@ -1,261 +0,0 @@
# Storage/Cluster Network Segmentation Audit — pve1.sweet.home
**Date:** 2026-07-29
**Scope:** Read-only discovery of pve1.sweet.home host networking, HA cluster VMs (200/201), and Docker CT (105). No changes made.
> **Implementation status — 2026-07-29:** All recommendations from this audit have been
> implemented in the same session. See `docs/ip-addressing.md` for the current state.
> Key decisions that diverged from the original recommendations:
> - VLAN IDs renumbered: cluster → VLAN 10 (192.168.10.x), storage-client → VLAN 20 (192.168.20.x)
> - Two Pacemaker VIPs: `vip-lan` (192.168.2.229, NFS for LAN) and `vip-storage` (192.168.20.229, NFS + iSCSI for VLAN 20)
> - NFS served on **both** VIPs (each firewalled to its own subnet); iSCSI available on VLAN 20 but NFS is preferred for docker to support future Docker Swarm multi-host access
> - `corosync.conf` ring1 added using LAN IPs (RF-1 resolved)
> - `vmbr2` created and NICs added to HA VMs and docker CT (RF-6/RF-7 resolved)
> - iSCSI portal remains on `[::0]`; firewall enforces VLAN 20 restriction (RF-4 mitigated)
> - STONITH still disabled (RF-3 deferred — accepted risk during development phase)
> - iSCSI ACLs not configured (RF-5 deferred — iSCSI not in active use)
---
## 1. Current State Summary
### pve1.sweet.home Host — Physical NICs
| Interface | Speed/Duplex | Notes |
|-----------|-------------|-------|
| `nic0` | 2500 Mb/s / Full (2.5GbE) | Only active physical NIC; sole bridge port for vmbr0 |
| `nic1` | (not connected / no data) | Present in config, not UP |
| `wlp4s0` | DOWN | WiFi, unused |
No bonding configured. Every guest's traffic ultimately funnels through the single 2.5GbE `nic0`.
### Proxmox Bridges
| Bridge | Physical NIC | Host IP | Subnet | VLAN-aware | Purpose (current) |
|--------|-------------|---------|--------|------------|-------------------|
| `vmbr0` | `nic0` (2.5GbE) | 192.168.2.245/24 | 192.168.2.0/24 | No | General LAN, management, **iSCSI/NFS VIP** |
| `vmbr1` | **none** (internal-only) | — | 192.168.4.224/29 | No | Corosync heartbeat + DRBD replication |
`vmbr1` has `bridge-ports none` in `/etc/network/interfaces.d/vmbr1.conf` — it is a purely software bridge with zero physical uplink. All traffic on it stays inside the hypervisor's memory.
### Guest NIC Assignments
| Guest | VMID | Role | NIC | Bridge | IP | Traffic type |
|-------|------|------|-----|--------|----|--------------|
| nix-cache | CT 102 | Build cache | eth0 | vmbr0 | DHCP/LAN | LAN |
| pxe-boot | CT 103 | PXE/TFTP | eth0 | vmbr0 | DHCP/LAN | LAN |
| tor-relay | CT 104 | Tor | eth0 | vmbr0 | DHCP/LAN | LAN |
| **docker** | **CT 105** | **Docker host** | **eth0** | **vmbr0** | **192.168.2.225/24** | **LAN only** |
| pdm | CT 106 | Proxmox mgmt | eth0 | vmbr0 | 192.168.2.220/24 | LAN |
| server | VM 101 | General server | net0 | vmbr0 | DHCP/LAN | LAN |
| tailscale-router | VM 107 | Tailscale exit | net0 | vmbr0 | DHCP/LAN | LAN |
| domain-controller | VM 108 | FreeIPA | net0 | vmbr0 | 192.168.2.253/24 | LAN |
| **ha-server-1** | **VM 200** | **HA primary** | **net0** | **vmbr0** | **192.168.2.228/24 + VIP 192.168.2.229/24** | **LAN + VIP** |
| **ha-server-1** | **VM 200** | **HA primary** | **net1** | **vmbr1** | **192.168.4.228/29** | **Corosync + DRBD** |
| **ha-server-2** | **VM 201** | **HA secondary** | **net0** | **vmbr0** | **192.168.2.227/24** | **LAN** |
| **ha-server-2** | **VM 201** | **HA secondary** | **net1** | **vmbr1** | **192.168.4.227/29** | **Corosync + DRBD** |
### Corosync/Pacemaker State
- **Transport:** knet/UDP
- **Rings:** 1 only — ring0 on `192.168.4.228` / `192.168.4.227` (vmbr1/ens19)
- **Cluster status:** Both nodes online, DC = ha-server-1, quorum achieved
- **STONITH:** `stonith-enabled: false`
- **no-quorum-policy:** `ignore`
- **Resources (all active on ha-server-1):**
- `ms-drbd0` — promotable DRBD clone (Primary: ha-server-1, Secondary: ha-server-2)
- `xfs-data` — XFS on `/dev/drbd0``/srv/ha-data`
- `iscsi-target` — targetctl service
- `nfs-server` — nfs-server service
- `vip` — IPaddr2 at **192.168.2.229/24** (no `nic=` parameter specified; floats to ens18/vmbr0 automatically based on subnet match)
- **Resource ordering:** ha-group starts only after DRBD is promoted; collocated with Promoted DRBD clone.
### DRBD State
- **Resource:** `ha-data` (DRBD 8.4.11 kernel module, config in `/etc/drbd.conf`)
- **Protocol:** C (synchronous)
- **Replication endpoints:**
- ha-server-1: `192.168.4.228:7789` (ens19 / vmbr1)
- ha-server-2: `192.168.4.227:7789` (ens19 / vmbr1)
- **State at audit time:** Initial sync in progress — ~20% complete, ~40 MB/s, ~34 min remaining (100 GB disk)
- **Fencing config:** `fencing resource-only` + fence-peer/unfence-peer handlers
### iSCSI Target
- **IQN:** `iqn.2026-01.home.sweet:ha-storage`
- **Portal:** `[::0]:3260` — confirmed listening on all interfaces (`ss -tnlp` shows `*:3260 *:*`)
- **LUN 0:** fileio backstore — `/srv/ha-data/iscsi-lun.img` (10 GiB, write-thru)
- **ACLs:** **None** (`no-gen-acls`, `no-auth`)
- **VIP (intended portal):** 192.168.2.229 — on vmbr0/LAN, no storage-NIC-specific binding
### NFS Exports
Served from the same `ha-group` as iSCSI (starts/stops together):
| Export path | Client subnet |
|------------|---------------|
| `/srv/ha-data/docker/{config,databases,volumes,nextcloud-data}` | 192.168.2.0/24 |
| `/srv/ha-data/proxmox/{iso,lxc}` | 192.168.2.0/24 |
| `/srv/ha-data/pxe-boot/images` | 192.168.2.0/24 |
| `/srv/ha-data/raspi/volumes` | 192.168.2.0/24 |
All NFS exports restrict to 192.168.2.0/24 and are served via the VIP at 192.168.2.229. `rw,sync,no_subtree_check,no_root_squash`.
### Docker Host Current State
- **Single NIC:** eth0 on vmbr0, 192.168.2.225/24, gateway 192.168.2.254
- **Route to storage network (192.168.4.x):** none — no NIC and no route
- **iSCSI sessions:** none
- **iSCSI nodes discovered:** none
- **Docker networks:** several active compose-project networks (core_traefik, core_nextcloud, core_passbolt, core_gramps, core_docker-socket-proxy, plus CI isolation networks)
---
## 2. Risk Flags
### RF-1: Corosync has only one ring (no heartbeat path redundancy)
`corosync.conf` defines only `ring0_addr` for each node, using 192.168.4.x on vmbr1. No `ring1_addr` / second knet link is configured. On a single Proxmox host, vmbr1 is a software bridge (no physical NIC), so physical link failure is not the concern — but a kernel network stack hiccup, a `pveproxy` restart dropping bridge state, or vmbr1 getting disrupted during heavy DRBD sync all leave corosync with zero fallback path. Missed heartbeats on a two-node cluster with `no-quorum-policy: ignore` do not cause a clean shutdown; they cause a false failover or split-brain.
Adding the LAN addresses (192.168.2.228 / 192.168.2.227 via ens18/vmbr0) as a second knet link would provide a backup path with no infrastructure changes needed.
### RF-2: Corosync heartbeat and DRBD replication share vmbr1 — no isolation between them
Both corosync (knet/UDP, ~1 kB heartbeat packets every ~100 ms) and DRBD replication (protocol C, synchronous, syncing at ~40 MB/s on a 100 GB initial fill at audit time) traverse the same `vmbr1` virtual bridge and terminate on the same ens19 NIC pair inside each HA VM. Under heavy DRBD write load, the guest-kernel scheduler's NIC transmit queue processes both flows together. While corosync's heartbeat is tiny, the absence of QoS/priority marking on vmbr1 means a DRBD burst can delay a heartbeat enough to trigger a ring fault warning. This is a latent risk that grows under high-write workloads.
### RF-3: STONITH disabled — split-brain protection relies solely on DRBD's resource-only fencing
`stonith-enabled: false` in the CIB. With `no-quorum-policy: ignore`, both nodes will continue running if corosync loses communication. DRBD's `fencing resource-only` does call `fence-peer` before allowing a Primary promotion, which provides some protection, but there is no hard external power fence to guarantee the other node actually stops. In a real split-brain (both nodes believe they are Primary), data corruption on the shared XFS filesystem is possible. **This is the highest-severity risk in the current setup.**
Getting STONITH to work on Proxmox-hosted VMs requires either a `fence_pve` agent (Proxmox API fencing) or `fence_virtd` (QEMU guest agent fencing). Neither is configured.
### RF-4: iSCSI portal bound to `[::0]:3260` — listens on every interface, not just the VIP
The targetcli portal is `[::0]:3260` (confirmed: `*:3260 *:*` in ss). This means the target is reachable on:
- 192.168.2.229 (VIP — correct, failover-safe)
- 192.168.2.228 (ha-server-1 LAN IP — does **not** move during failover; an initiator session connecting here would break on failover)
- 192.168.4.228 (storage NIC — not reachable by the Docker host today, but unintentionally exposed)
Binding the portal explicitly to the VIP IP instead of wildcard eliminates the non-VIP reachability risks.
### RF-5: iSCSI has zero ACLs and no authentication
`targetcli ls` shows `acls: 0`, `no-gen-acls`, `no-auth`. Any host that can reach port 3260 on any of the above IPs can log into the LUN with no credentials. The Docker host is not yet configured as an initiator — but neither is it blocked.
### RF-6: Docker host has no path to the storage network — iSCSI would traverse vmbr0/nic0
CT 105 (docker, 192.168.2.225) has one NIC, on vmbr0. To reach the VIP at 192.168.2.229, iSCSI traffic would travel:
```
docker (eth0/vmbr0) → nic0 (2.5GbE) → vmbr0 → tap200i0 (VM 200 net0/ens18)
```
All of the following share this same path over vmbr0 → nic0:
- Docker container traffic (outbound and inter-container)
- CI/CD runner traffic (Gitea Actions jobs visible in `docker network ls`)
- NFS mounts from pxe-boot, proxmox host itself, and other LAN clients
- iSCSI block traffic (protocol-sensitive to latency and retransmit)
A Nextcloud upload or a CI `nix build` job can saturate nic0 and starve the iSCSI session, causing command timeouts and filesystem errors on the Docker host.
### RF-7: VIP is on the LAN interface with no storage-specific binding
The Pacemaker `vip` resource specifies `ip=192.168.2.229, cidr_netmask=24` with no `nic=` parameter. Pacemaker's IPaddr2 agent selects the interface by longest-prefix match, landing it on ens18 (vmbr0/LAN). There is no way to keep this VIP from competing with general LAN traffic on nic0 without moving the VIP to a separate subnet on a different virtual bridge.
---
## 3. Recommended Target Layout
### Design constraints
- Single Proxmox host: all traffic ultimately shares nic0's bandwidth. The goal is QoS partitioning via separate bridges and subnets, not true physical isolation.
- Future physical split: bridge/VLAN IDs chosen here should map cleanly to physical uplink VLAN tags when the HA nodes move to separate hardware.
### Proposed bridge layout
| Bridge | Physical port | VLAN tag (future) | Subnet | Purpose |
|--------|-------------|-------------------|--------|---------|
| `vmbr0` | nic0 | untagged / VLAN 1 | 192.168.2.0/24 | **LAN/management only** — no storage traffic |
| `vmbr1` | (none / VLAN 10 on future trunk) | VLAN 10 | 192.168.4.224/29 | **Corosync heartbeat + DRBD replication** (current, keep) |
| `vmbr2` *(new)* | (none / VLAN 20 on future trunk) | VLAN 20 | 192.168.5.0/24 | **Storage: iSCSI + NFS client access** |
This is the minimum-disruption path: vmbr1 stays as-is (no DRBD reconfiguration needed), and the new vmbr2 gives the Docker host a direct path to the storage VIP without crossing vmbr0.
If stricter isolation is later desired, DRBD can be migrated from vmbr1 to vmbr2 in a separate maintenance window (see §5), leaving vmbr1 as corosync-only.
### Per-guest NIC assignments in target layout
| Guest | VMID | NIC | Bridge | Proposed IP | Purpose |
|-------|------|-----|--------|-------------|---------|
| ha-server-1 | VM 200 | net0 | vmbr0 | 192.168.2.228/24 | LAN/management (keep) |
| ha-server-1 | VM 200 | net1 | vmbr1 | 192.168.4.228/29 | Corosync + DRBD (keep) |
| ha-server-1 | VM 200 | **net2 (new)** | **vmbr2** | **192.168.5.1/24** | iSCSI + NFS storage client |
| ha-server-2 | VM 201 | net0 | vmbr0 | 192.168.2.227/24 | LAN/management (keep) |
| ha-server-2 | VM 201 | net1 | vmbr1 | 192.168.4.227/29 | Corosync + DRBD (keep) |
| ha-server-2 | VM 201 | **net2 (new)** | **vmbr2** | **192.168.5.2/24** | iSCSI + NFS storage client |
| docker | CT 105 | net0 | vmbr0 | 192.168.2.225/24 | LAN/management (keep) |
| docker | CT 105 | **net1 (new)** | **vmbr2** | **192.168.5.10/24** | iSCSI + NFS |
**VIP target:** `192.168.5.100/24` on vmbr2. The Pacemaker `vip` resource changes from `ip=192.168.2.229` to `ip=192.168.5.100, nic=<ens20>` (whichever name the new NIC gets inside the HA VMs). The existing `192.168.2.229` LAN VIP can optionally be retained as a separate static alias on ens18 for management-plane access, but should not be the iSCSI portal target.
**iSCSI portal:** Bind to `192.168.5.100:3260` instead of `[::0]:3260`. In targetcli: remove the wildcard portal, add `portals/ create 192.168.5.100`.
**NFS exports:** NFS is a file-level protocol and is fine being accessed over a routed path. After the VIP moves, non-docker LAN clients (proxmox host, pxe-boot, raspi) can reach NFS either via a static route to 192.168.5.0/24 or by keeping a secondary static alias at 192.168.2.229 on ens18 dedicated to NFS. Either approach works — NFS handles reconnect gracefully in ways iSCSI block I/O cannot.
**Corosync second ring (independent, low-disruption improvement):**
Add a second knet link using the LAN addresses as a backup heartbeat path. Edit `corosync.conf` on both nodes:
```
node { ring0_addr: 192.168.4.228; ring1_addr: 192.168.2.228; name: ha-server-1; nodeid: 1; }
node { ring0_addr: 192.168.4.227; ring1_addr: 192.168.2.227; name: ha-server-2; nodeid: 2; }
```
Requires a corosync service restart (brief cluster pause, ~5 seconds), no interface or bridge changes.
**Future physical-host split:**
When ha-server-1 and ha-server-2 move to separate physical machines, vmbr1 and vmbr2 become VLAN-tagged sub-interfaces on a physical trunk (e.g. VLAN 10 → cluster, VLAN 20 → storage). The bridge/subnet/IP layout above is designed so the tag numbers can be layered onto the existing addresses without renumbering.
---
## 4. Gap List
| Gap | Action needed |
|-----|--------------|
| `vmbr2` does not exist on pve1 | Create internal bridge: `/etc/network/interfaces.d/vmbr2.conf` with `bridge-ports none`, `inet manual` |
| VM 200 and VM 201 have no net2 | `qm set 200 --net2 virtio,bridge=vmbr2` / `qm set 201 --net2 virtio,bridge=vmbr2` (hot-plug, no reboot needed) |
| CT 105 has no net1 | `pct set 105 --net1 name=eth1,bridge=vmbr2,ip=192.168.5.10/24` |
| HA VMs have no OS config for the new NIC | NixOS `networking.interfaces.<ens20>` with `ipv4.addresses = [{address="192.168.5.1"; prefixLength=24;}]` per host (name may differ — check `ip link` after hotplug) |
| VIP needs to move to 192.168.5.100 on vmbr2 | `pcs resource update vip ip=192.168.5.100 cidr_netmask=24 nic=<ens20>` |
| iSCSI portal bound to `[::0]` | `targetcli /iscsi/iqn.2026-01.home.sweet:ha-storage/tpg1/portals delete ::0 3260` then `create 192.168.5.100`; save and restart via `pcs resource restart iscsi-target` |
| iSCSI ACLs empty | Get Docker initiator IQN via `iscsiadm -m iface` on CT 105, then add via targetcli `acls/ create <iqn>` |
| Docker host has no iSCSI initiator config | `iscsiadm -m discoverydb -t sendtargets -p 192.168.5.100 -D` then `iscsiadm -m node -l` once ACLs are set |
| Corosync single ring | Add `ring1_addr` entries in `corosync.conf` using LAN IPs; restart corosync cluster-wide (one node at a time) |
| STONITH not configured | Evaluate `fence_pve` (Proxmox API agent); document accepted risk if deferred |
| DRBD still on vmbr1 (optional, separate window) | Stop ms-drbd0 via pcs, edit `/etc/drbd.conf` on both nodes (change `192.168.4.x``192.168.5.x`), restart DRBD, re-enable via pcs |
---
## 5. Migration Notes
### Non-disruptive (no service impact)
- **Create vmbr2 on pve1:** Bridge definition edit only; no effect on existing bridges or guests.
- **Hot-add net2 to VMs 200/201:** Proxmox allows adding a NIC without reboot (`qm set 200 --net2 ...`). The NIC appears inside the VM immediately via QEMU hotplug but will be unconfigured (down) inside NixOS until the NixOS config is deployed — no impact on running services.
- **Add net1 to CT 105:** LXC NIC hotplug works similarly; CT does not need to restart.
- **Add corosync ring1:** Requires `systemctl restart corosync` on both nodes (one at a time). Pacemaker briefly sees corosync go offline and recover; with two nodes and `wait_for_all: 0`, this typically completes in under 5 seconds and resources stay running.
### Disruptive — requires maintenance window
- **Move VIP from 192.168.2.229 to 192.168.5.100:** `pcs resource update vip ip=192.168.5.100` causes Pacemaker to immediately stop the old VIP and start the new one. Any NFS mounts referencing 192.168.2.229 will stall until remounted at the new address (or a static alias is added at 192.168.2.229 on ens18). No iSCSI sessions exist yet, so no iSCSI disruption.
- **Change iSCSI portal from `[::0]` to VIP-specific:** Requires `pcs resource restart iscsi-target` after the targetcli portal change — brief target unavailability. Any initiator sessions (once configured) will need to re-login.
- **Migrate DRBD replication from 192.168.4.x to 192.168.5.x** (optional — only needed to give corosync sole ownership of vmbr1):
1. `pcs resource disable ms-drbd0` — demotes DRBD Primary, stops ha-group (unmounts XFS, stops iSCSI + NFS + VIP)
2. `drbdadm down ha-data` on both nodes
3. Edit `/etc/drbd.conf` on both nodes (change `address` lines)
4. `drbdadm up ha-data` on both nodes
5. `pcs resource enable ms-drbd0` — Pacemaker re-promotes, mounts, starts services
DRBD does **not** require a full resync when only the address changes — the disk data and metadata are unchanged; only the TCP connection endpoint changes. However, the initial sync was in progress at audit time (~20% at ~40 MB/s). Recommend waiting for that sync to complete before scheduling this migration.
+160
View File
@@ -0,0 +1,160 @@
# HA File-Server Cluster
Two `proxmox-ha-server-{1,2}` VMs form an active/passive file-server cluster:
DRBD replicates a block device between nodes; Corosync + Pacemaker manage
failover; XFS, LIO iSCSI, and NFS are brought up as a collocated resource
group on whichever node holds the DRBD Primary role.
NixOS modules: `modules/ha/`. Lifecycle scripts: `scripts/ha/`.
Cluster-wide constants: `variables.nix` (`haServer*` vars).
---
## Network layout
Three subnets — all internal to pve1 (`vmbr0`/`vmbr1`/`vmbr2`):
| Subnet | VLAN | CIDR | Bridge | Purpose |
|---|---|---|---|---|
| LAN | 2 | `192.168.2.0/24` | `vmbr0` | Management, LAN NFS |
| Cluster | 10 | `192.168.10.224/29` | `vmbr1` | Corosync ring0 + DRBD replication |
| Storage-client | 20 | `192.168.20.0/24` | `vmbr2` | NFS + iSCSI for docker/swarm |
Each HA VM has three NICs: `ens18` (LAN/vmbr0), `ens19` (cluster/vmbr1),
`ens20` (storage-client/vmbr2). See `docs/ip-addressing.md` for all IPs.
Corosync ring0 uses the cluster NIC; ring1 (backup heartbeat) uses the LAN
NIC. DRBD replicates over the cluster NIC. No storage traffic crosses the LAN.
---
## Pacemaker resources
All resources run collocated on whichever node is Primary, in this order:
```
ms-drbd0 (promotable DRBD clone)
→ xfs-data (XFS mount on /dev/drbd0 → /srv/ha-data)
→ iscsi-target (targetctl)
→ nfs-server (nfs-server.service)
→ vip-lan (192.168.2.229/24 on vmbr0 — NFS for LAN clients)
→ vip-storage (192.168.20.229/24 on vmbr2 — NFS + iSCSI for VLAN 20)
```
`vip-lan` serves pxe-boot and other LAN-only NFS clients.
`vip-storage` serves docker and any future swarm nodes; iSCSI is available on
VLAN 20 but NFS is preferred for multi-host volume sharing.
---
## DRBD fencing
`fencing resource-only` with `crm-fence-peer.sh`/`crm-unfence-peer.sh`
wrappers (`modules/ha/cluster-config.nix`). The DRBD kernel module invokes
these via the User Mode Helper with a minimal PATH; the wrappers prepend
`/run/current-system/sw/bin` before exec-ing the real handlers so Pacemaker
tools (`cibadmin`, `crm_mon`, etc.) are found.
STONITH is initially disabled (`stonith-enabled: false`,
`no-quorum-policy: ignore`). Enable it once the `fence_pve_ssh` fence agent
(`scripts/ha/fence-pve-ssh.py`) is deployed and authorised:
```bash
scripts/ha/cluster-enable-stonith.sh # run as root on ha-server-1
```
---
## Deploying the cluster from scratch
Use `scripts/ha/deploy.sh` — it orchestrates all phases:
```bash
# Against pve-test (safe — Claude's default target):
scripts/ha/deploy.sh --node "$PVE_TEST_HOST" [--dry-run]
# Against pve1 (production — requires explicit operator go-ahead):
scripts/ha/deploy.sh --node "$PVE1_HOST"
```
Phases (each skippable with `--skip-<phase>`):
1. `ensure-bridge` — creates `vmbr1`/`vmbr2` on the Proxmox node if absent
2. `sync-keys` — generates SSH host keys for both nodes; registers sops recipients
3. `create-vms` — builds disk images, creates VMs via `create-proxmox-resource.sh`
4. `add-hardware` — attaches storage NIC and DRBD data disk to each VM
5. `init-cluster` — runs `scripts/ha/cluster-init.sh` on ha-server-1
`--destroy` runs the teardown sequence.
---
## Day-to-day operations
```bash
# Read-only health check (safe from workstation):
scripts/ha/health.sh
# Graceful failover (prompts for confirmation):
scripts/ha/failover.sh [--to node1|node2]
# Online data-disk growth (no downtime):
scripts/ha/resize-data-disk.sh --size +20G
# Acceptance tests (run after any significant change):
scripts/ha/acceptance-tests.sh
```
---
## Adding FreeIPA host accounts
IPA host registration is automated:
```bash
scripts/ipa/create-nixos-ipa-host-account.sh <hostname>
```
This runs `ipa host-add`, fetches a keytab from the domain controller, and
writes a sops-encrypted `secrets/<hostname>.keytab` in one step. The module
`modules/ipa/client.nix` (imported by every host via
`modules/common/configuration.nix`) consumes the keytab via sops-nix.
---
## Storage layout
```
/srv/ha-data/
docker/
config/ NFS → docker:/mnt/docker/config
databases/ NFS → docker:/mnt/docker/databases
volumes/ NFS → docker:/mnt/docker/volumes
nextcloud-data/ NFS → docker:/mnt/docker/nextcloud-data
proxmox/
iso/ NFS → pve1 ISO storage
lxc/ NFS → pve1 CT template storage
pxe-boot/
images/ NFS → pxe-boot:/srv/pxe/http/images (PXE assets)
raspi/
volumes/ NFS → raspi NFS mounts
iscsi-lun.img iSCSI fileio backstore (VLAN 20 only, not in active use)
```
All shares are defined in `variables.nix` (`vars.nfsShares.*`). The NFS
export list lives in `modules/ha/nfs-exports.nix`.
---
## Key variables
| Variable | Description |
|---|---|
| `vars.haServer1Ip` / `vars.haServer2Ip` | LAN management IPs |
| `vars.haServer1StorageIp` / `vars.haServer2StorageIp` | Cluster NIC IPs (DRBD/Corosync ring0) |
| `vars.haServerLanVip` | Pacemaker `vip-lan` — NFS for LAN (192.168.2.229) |
| `vars.haServerVip` | Pacemaker `vip-storage` — NFS + iSCSI for VLAN 20 (192.168.20.229) |
| `vars.haLanNfsFqdn` | FQDN of `vip-lan`: `ha-vip-lan.sweet.home` |
| `vars.haStorageRoot` | XFS mount point: `/srv/ha-data` |
| `vars.haServerDrbdDisk` | Block device for DRBD backing store |
| `vars.haStorageCidr` | Cluster subnet CIDR (`192.168.10.224/29`) |
| `vars.haClientCidr` | Storage-client subnet CIDR (`192.168.20.0/24`) |
-39
View File
@@ -184,42 +184,3 @@ cannot reach either service on this VIP. The `vip-storage` endpoint is not reach
from the workstation directly (internal bridge only); health checks proxy through the from the workstation directly (internal bridge only); health checks proxy through the
active HA node. active HA node.
---
## Migration reference
Current → target IP for every host being renumbered.
| Host | Current IP | New IP | Config location |
|---|---|---|---|
| router | `192.168.2.254` | `192.168.2.254` | unchanged |
| domain-controller | `192.168.2.138` | `192.168.2.253` | `/etc/sysconfig/network-scripts/ifcfg-eth0` on guest |
| pve1 | `192.168.2.250` | `192.168.2.245` | `/etc/network/interfaces` on Proxmox host |
| pbs | `192.168.2.108` | `192.168.2.244` | static config on PBS host |
| nixos workstation | `192.168.2.119` | `192.168.2.243` | `networking.interfaces` / NetworkManager on guest |
| ha-node1 | — | `192.168.2.228` (LAN), `192.168.10.228` (cluster/VLAN 10), `192.168.20.228` (storage/VLAN 20) | active |
| ha-node2 | — | `192.168.2.227` (LAN), `192.168.10.227` (cluster/VLAN 10), `192.168.20.227` (storage/VLAN 20) | active |
| ha-vip-lan | — | `192.168.2.229` (vmbr0 / Pacemaker `vip-lan`) — NFS endpoint for LAN clients | active |
| ha-vip-storage | — | `192.168.20.229` (vmbr2 / Pacemaker `vip-storage`) — iSCSI endpoint for VLAN 20 clients | active |
| server | `192.168.2.252` | `192.168.2.226` | static config on guest |
| docker | `192.168.2.249` | `192.168.2.225` | static config on guest |
| nix-cache | `192.168.2.120` | `192.168.2.224` | static config on guest |
| pxe-boot | `192.168.2.247` | `192.168.2.223` | static config on guest; update `vars.pxeServerIp` in `variables.nix` ✓ |
| tailscale-router | `192.168.2.121` | `192.168.2.222` | static config on guest |
| tor-relay | `192.168.2.107` | `192.168.2.221` | static config on guest |
| pdm | `192.168.2.248` | `192.168.2.220` | static config on guest |
### Cutover notes
- **Do domain-controller first** — it becomes the DNS server; everything else depends on it
having its new IP and FreeIPA DNS configured before Pi-hole is retired.
- **pve1 last among physical hosts** — changing the Proxmox management IP drops the web UI
briefly; all guests keep running.
- **Update Pi-hole custom.list / FreeIPA DNS A records** to new IPs before flipping any host,
so name resolution stays valid throughout the migration.
- **variables.nix already updated** for `pxeServerIp` (.247→.223), `pbsIp` (.108→.244), and
new `domainControllerIp` (.253). Rebuild affected hosts after renumbering.
- **Router DHCP**: once domain-controller is at .253 and FreeIPA DNS is serving `sweet.home`,
switch router DHCP on with pool .10.59 and DNS option pointing to .253; retire Pi-hole CT.
- **Pi-hole's iPXE dnsmasq config** (`99-ipxe-chainload.conf`) moves to the pxe-boot CT as a
dnsmasq proxy-mode config before Pi-hole is decommissioned.
-448
View File
@@ -1,448 +0,0 @@
# Network Cutover Plan
Moves the LAN from the current flat/Pi-hole-managed state to the new IP scheme
defined in `docs/ip-addressing.md`. Works in five independent stages — each
stage is safe to pause after and resume later. Rollback steps are given at
every point where something can break.
**Before starting anything:** confirm you have
- SSH access to `192.168.2.138` (domain-controller, current IP)
- SSH access to `192.168.2.250` (pve1)
- Browser access to Pi-hole admin at `http://192.168.2.253`
- Browser access to router admin at `http://192.168.2.254`
- The FreeIPA `admin` password to hand
---
## Stage 1 — Prepare FreeIPA DNS (zero downtime)
Everything here is additive. Pi-hole keeps running. Nothing breaks if you stop
mid-stage.
### 1a. Add NextDNS forwarders
```bash
ssh wayne@192.168.2.138
kinit admin # enter FreeIPA admin password when prompted
ipa dnsconfig-mod \
--forwarder=45.90.28.142 \
--forwarder=45.90.30.142 \
--forward-policy=only
```
**Verify external resolution works through FreeIPA before continuing:**
```bash
dig @127.0.0.1 google.com +short # must return an IP, not SERVFAIL
```
### 1b. Add A records for every host at their CURRENT IPs
These represent the live state now. You'll update each record to the new IP
when you renumber that host in Stage 5.
```bash
ipa dnsrecord-add sweet.home pve1 --a-rec 192.168.2.250
ipa dnsrecord-add sweet.home pbs --a-rec 192.168.2.108
ipa dnsrecord-add sweet.home nixos --a-rec 192.168.2.119
ipa dnsrecord-add sweet.home server --a-rec 192.168.2.252
ipa dnsrecord-add sweet.home docker --a-rec 192.168.2.249
ipa dnsrecord-add sweet.home nix-cache --a-rec 192.168.2.120
ipa dnsrecord-add sweet.home pxe-boot --a-rec 192.168.2.247
ipa dnsrecord-add sweet.home tailscale-router --a-rec 192.168.2.121
ipa dnsrecord-add sweet.home tor-relay --a-rec 192.168.2.107
ipa dnsrecord-add sweet.home pdm --a-rec 192.168.2.248
ipa dnsrecord-add sweet.home router --a-rec 192.168.2.254
```
### 1c. Clean up stale reverse-zone PTR records
FreeIPA already has PTR records from an earlier import but some are wrong.
Fix them now so reverse DNS is accurate from day one.
```bash
# Remove stale "win11" entry at .250 (should be pve1)
ipa dnsrecord-del 2.168.192.in-addr.arpa 250 --ptr-rec win11.
ipa dnsrecord-add 2.168.192.in-addr.arpa 250 --ptr-rec pve1.sweet.home.
# Fix unqualified PTR records (missing .sweet.home. suffix)
ipa dnsrecord-mod 2.168.192.in-addr.arpa 108 --ptr-rec pbs.sweet.home.
ipa dnsrecord-mod 2.168.192.in-addr.arpa 248 --ptr-rec pdm.sweet.home.
ipa dnsrecord-mod 2.168.192.in-addr.arpa 249 --ptr-rec docker.sweet.home.
ipa dnsrecord-mod 2.168.192.in-addr.arpa 252 --ptr-rec server.sweet.home.
# Add any missing PTR records
ipa dnsrecord-add 2.168.192.in-addr.arpa 119 --ptr-rec nixos.sweet.home.
ipa dnsrecord-add 2.168.192.in-addr.arpa 120 --ptr-rec nix-cache.sweet.home.
ipa dnsrecord-add 2.168.192.in-addr.arpa 121 --ptr-rec tailscale-router.sweet.home.
ipa dnsrecord-add 2.168.192.in-addr.arpa 247 --ptr-rec pxe-boot.sweet.home.
ipa dnsrecord-add 2.168.192.in-addr.arpa 254 --ptr-rec router.sweet.home.
```
### 1d. Point domain-controller's own DNS at itself
```bash
sudo nmcli connection modify "System eth0" ipv4.dns "127.0.0.1"
sudo nmcli connection up "System eth0"
```
**Verify:**
```bash
dig pve1.sweet.home +short # must return 192.168.2.250
dig google.com +short # must return an IP (NextDNS forwarding)
```
**Rollback 1d:** `sudo nmcli connection modify "System eth0" ipv4.dns "192.168.2.253" && sudo nmcli connection up "System eth0"`
---
## Stage 2 — Move pxe-boot DHCP options off Pi-hole (zero downtime)
Pi-hole's dnsmasq currently serves the iPXE boot options via
`99-ipxe-chainload.conf`. Before Pi-hole is retired, that config must move to
the pxe-boot CT running dnsmasq in proxy mode so PXE boot keeps working.
### 2a. Add dnsmasq proxy config to the pxe-boot NixOS module
In `modules/build-types/pxe-boot.nix`, add:
```nix
services.dnsmasq = {
enable = true;
settings = {
# Proxy mode: respond only to PXE DHCP requests, leave normal leases to router
dhcp-range = [ "192.168.2.0,proxy" ];
# iPXE client detection
dhcp-match = [
"set:ipxe,175"
"set:efi64,option:client-arch,7"
"set:efi64,option:client-arch,9"
];
dhcp-userclass = "set:ipxe,iPXE";
# Boot file selection
dhcp-boot = [
"tag:ipxe,tag:efi64,http://${vars.pxeServerIp}/boot.ipxe"
"tag:ipxe,http://${vars.pxeServerIp}/boot.ipxe"
"tag:efi64,ipxe.efi,,${vars.pxeServerIp}"
"undionly.kpxe,,${vars.pxeServerIp}"
];
};
};
```
### 2b. Rebuild and deploy the pxe-boot CT
```bash
# On pve1 — build the new tarball
nix build .#lxc-pxe-boot.config.system.build.tarball
# Verify dnsmasq starts correctly in the CT after deploy
ssh nixos@192.168.2.247 systemctl status dnsmasq
```
### 2c. Remove the iPXE config from Pi-hole
In the Pi-hole CT, remove `/etc/dnsmasq.d/99-ipxe-chainload.conf` and
restart the FTL service:
```bash
ssh wayne@pve1.sweet.home \
"sudo pct exec 100 -- bash -c 'rm /etc/dnsmasq.d/99-ipxe-chainload.conf && systemctl restart pihole-FTL'"
```
**Verify:** PXE boot a test machine — it should still get an iPXE response and
reach the boot menu.
**Rollback 2c:** restore the file from the Pi-hole config backup at
`/etc/pihole/config_backups/` and restart pihole-FTL.
---
## Stage 3 — DHCP migration: Pi-hole → router (brief maintenance window)
**Do this in the evening.** Existing DHCP leases stay valid during the
switchover so connected devices don't drop — only new lease requests fail
during the gap, which is under 60 seconds if you follow the steps in order.
The key: configure the router's DHCP DNS option to point at `.253` (Pi-hole's
current IP). This way, all new leases issued by the router still get the same
DNS server address — clients never need to change their DNS config. When Pi-hole
is retired and the DC takes `.253` in Stage 4, `.253` just starts answering
differently. No client reconfiguration.
### 3a. Pre-configure router DHCP (do not enable yet)
Log into `http://192.168.2.254`, find the DHCP settings and fill in — but
leave DHCP **disabled** until step 3b:
| Setting | Value |
|---|---|
| Start IP | 192.168.2.10 |
| End IP | 192.168.2.59 |
| Subnet mask | 255.255.255.0 |
| Gateway | 192.168.2.254 |
| Primary DNS | 192.168.2.253 |
| Secondary DNS | *(leave blank)* |
| Lease time | 24h |
Save without enabling.
### 3b. Switchover (do steps in quick succession)
1. **Disable Pi-hole DHCP:** Pi-hole admin UI → Settings → DHCP → uncheck
"DHCP server enabled" → Save
2. **Enable router DHCP** immediately after step 1
### 3c. Verify router DHCP is working
On a phone or laptop, disconnect from WiFi and reconnect (or run
`sudo dhclient -r && sudo dhclient` on a Linux host):
```bash
ip addr show # IP should be in 192.168.2.1059 range
dig google.com # should resolve (Pi-hole DNS still running at .253)
dig pve1.sweet.home # should resolve via FreeIPA at .138 (relayed via Pi-hole)
```
Wait 1015 minutes for the most active devices to renew their leases. There's
no need to wait for all leases to expire before proceeding.
**Rollback 3b:** Re-enable Pi-hole DHCP. Disable router DHCP. Done — existing
leases remain valid so most devices are unaffected.
---
## Stage 4 — Move domain-controller from .138 to .253
Pi-hole lives at `.253`. The DC must take `.253` the moment Pi-hole stops so
clients that still have `.253` as their DNS server don't notice the change.
Script these commands in advance and run them in rapid succession.
**Pre-stage: have this SSH command ready before running step 4a:**
```bash
ssh wayne@192.168.2.138 "
sudo nmcli connection modify 'System eth0' \
ipv4.addresses '192.168.2.253/24' \
ipv4.gateway '192.168.2.254' \
ipv4.dns '127.0.0.1' \
ipv4.method manual && \
sudo nmcli connection up 'System eth0'
"
```
**Also update the Proxmox VM config to match (run from pve1):**
```bash
sudo qm set 108 \
--ipconfig0 ip=192.168.2.253/24,gw=192.168.2.254 \
--nameserver 192.168.2.253
```
### 4a. Stop Pi-hole
```bash
ssh wayne@pve1.sweet.home "sudo pct stop 100"
```
### 4b. Immediately: change DC's IP to .253
Run the pre-staged SSH command from above. You have ~30 seconds before any
client notices Pi-hole is gone. If SSH to `.138` refuses (the IP is already
changing), open a Proxmox console to VM 108 and run the `nmcli` commands
there.
### 4c. Update Proxmox VM config
Run the pre-staged `qm set 108` command from above.
### 4d. Verify
```bash
ssh wayne@192.168.2.253 # must connect (new DC IP)
dig @192.168.2.253 pve1.sweet.home +short # must return 192.168.2.250
dig @192.168.2.253 google.com +short # must return an IP
```
From a client device that renewed its DHCP lease in Stage 3:
```bash
cat /etc/resolv.conf # should show 192.168.2.253
dig pve1.sweet.home # should resolve
```
**Rollback 4:** `ssh wayne@pve1.sweet.home "sudo pct start 100"`. Change DC IP
back to .138 via Proxmox console. This restores full Pi-hole DNS/DHCP service.
Leave Pi-hole CT stopped-but-intact for 48 hours before deleting it.
---
## Stage 5 — Host renumbering (one at a time, any order)
For each host:
1. Update FreeIPA DNS A record and PTR record to the new IP
2. Change the static IP on the host itself
3. Verify SSH to new IP
4. Update `variables.nix` if that host has an IP variable (pxe-boot, pbs — already done in this PR)
**FreeIPA record update template** (run as admin on domain-controller):
```bash
ipa dnsrecord-mod sweet.home <hostname> --a-rec <new-ip>
ipa dnsrecord-del 2.168.192.in-addr.arpa <old-last-octet> --ptr-rec <hostname>.sweet.home.
ipa dnsrecord-add 2.168.192.in-addr.arpa <new-last-octet> --ptr-rec <hostname>.sweet.home.
```
### Renumbering order
| # | Host | Old IP | New IP | How to change IP |
|---|---|---|---|---|
| 1 | nixos workstation | .119 | .243 | NetworkManager on guest; or `nmcli connection modify` |
| 2 | nix-cache | .120 | .224 | `pct set 102 --net0 name=eth0,bridge=vmbr0,ip=192.168.2.224/24,gw=192.168.2.254` then `pct reboot 102` |
| 3 | tailscale-router | .121 | .222 | Static config on guest; check Tailscale ACLs if IP is referenced there |
| 4 | tor-relay | .107 | .221 | `pct set 104 --net0 name=eth0,bridge=vmbr0,ip=192.168.2.221/24,gw=192.168.2.254` then `pct reboot 104` |
| 5 | pdm | .248 | .220 | `pct set 106 --net0 name=eth0,bridge=vmbr0,ip=192.168.2.220/24,gw=192.168.2.254` then `pct reboot 106` |
| 6 | pxe-boot | .247 | .223 | `pct set 103 --net0 name=eth0,bridge=vmbr0,ip=192.168.2.223/24,gw=192.168.2.254` then rebuild NixOS (already updated in variables.nix) |
| 7 | server | .252 | .226 | Static config on guest; NFS clients (docker) lose mounts briefly — they remount automatically |
| 8 | docker | .249 | .225 | Static config on guest; do this after server is at .226 |
| 9 | pbs | .108 | .244 | Static config on PBS host itself; update in `pbsIp` already done in variables.nix |
| 10 | pve1 | .250 | .245 | Edit `/etc/network/interfaces` on the Proxmox host — see below |
### pve1 renumber (step 10 — do last)
All guests keep running; only the Proxmox web UI is briefly unreachable.
```bash
ssh wayne@pve1.sweet.home
# Edit /etc/network/interfaces: change address from .250 to .245
sudo nano /etc/network/interfaces
# Change: address 192.168.2.250/24
# To: address 192.168.2.245/24
sudo systemctl restart networking
# SSH will drop here — reconnect to new IP
```
```bash
ssh wayne@192.168.2.245 # verify
```
Update FreeIPA DNS:
```bash
ipa dnsrecord-mod sweet.home pve1 --a-rec 192.168.2.245
ipa dnsrecord-del 2.168.192.in-addr.arpa 250 --ptr-rec pve1.sweet.home.
ipa dnsrecord-add 2.168.192.in-addr.arpa 245 --ptr-rec pve1.sweet.home.
```
**Rollback any step 5 host:** change the IP back on the guest and update the
FreeIPA record back to the old IP. The old IP is unoccupied so you can
temporarily use either.
---
## Stage 6 — HA storage cutover (docker NFS remount)
> **Prerequisites:**
> - HA cluster fully deployed and `vip-storage` (`nfs.storage.home` → 192.168.20.229) serving NFS ✓
> - DNS configured: `storage.home` zone populated, `nfs.storage.home` resolves to 192.168.20.229 ✓
> - docker CT has eth1 on vmbr2 (`docker.storage.home` → 192.168.20.225) ✓
> - Final rsync from server.sweet.home to `/srv/ha-data` complete before step 6b
docker.sweet.home currently NFS-mounts its persistent volumes from `server.sweet.home`
(`192.168.2.226:/tank/docker/...`). This stage moves those mounts to the HA cluster's
storage VIP so server can be decommissioned.
### 6a. Final rsync from server to HA cluster
Run from server.sweet.home (or over SSH from the workstation) to sync any data written
since the initial rsync:
```bash
# Confirm active HA node and mount point
ssh wayne@192.168.2.228 'sudo findmnt /srv/ha-data' # check which node is active
# rsync each dataset (adjust source paths to match /tank layout on server)
sudo rsync -av --delete /tank/docker/config/ wayne@<active-node-ip>:/srv/ha-data/docker/config/
sudo rsync -av --delete /tank/docker/databases/ wayne@<active-node-ip>:/srv/ha-data/docker/databases/
sudo rsync -av --delete /tank/docker/volumes/ wayne@<active-node-ip>:/srv/ha-data/docker/volumes/
sudo rsync -av --delete /tank/docker/nextcloud-data/ wayne@<active-node-ip>:/srv/ha-data/docker/nextcloud-data/
```
### 6b. Update docker NixOS config to mount from vip-storage
In `hosts/docker/host.nix` (or wherever the NFS mount fileSystems are declared), change
the NFS server from `server.sweet.home` / `192.168.2.226` to `nfs.storage.home`:
```nix
# Before:
fileSystems."/mnt/docker/config" = {
device = "server:/tank/docker/config"; # or 192.168.2.226:...
...
};
# After:
fileSystems."/mnt/docker/config" = {
device = "nfs.storage.home:/srv/ha-data/docker/config";
...
};
```
Using the DNS name (`nfs.storage.home`) rather than the VIP IP means the mount
config survives a future VIP renumber without touching the NixOS config.
Repeat for all four docker shares (`config`, `databases`, `volumes`, `nextcloud-data`).
Then rebuild docker:
```bash
# On the workstation — or via Switch-nix on docker itself
sudo nixos-rebuild switch --no-write-lock-file --refresh \
--flake "git+https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos.git#lxc-docker"
```
### 6c. Verify mounts and container health
```bash
ssh wayne@192.168.2.225 'findmnt | grep 192.168.20' # mounts should show vip-storage
ssh wayne@192.168.2.225 'docker ps' # all containers running
```
Spot-check Nextcloud, Traefik, and any database containers for connectivity.
### 6d. Decommission server.sweet.home
Once docker is confirmed healthy on the HA NFS mounts:
```bash
# Stop server VM on pve1
ssh wayne@192.168.2.245 'sudo qm stop 101'
# (Optional) Archive the ZFS pool snapshot before destroying
# Then after a settling period:
ssh wayne@192.168.2.245 'sudo qm destroy 101 --destroy-unreferenced-disks 1'
```
---
## Stage 7 — Final cleanup
Once all hosts are at their new IPs and verified:
```bash
# Delete the Pi-hole CT (already stopped since Stage 4)
ssh wayne@pve1.sweet.home "sudo pct destroy 100"
# Remove stale FreeIPA records for retired addresses
ipa dnsrecord-del sweet.home pihole --del-all
ipa dnsrecord-del 2.168.192.in-addr.arpa 253 --ptr-rec pihole.sweet.home.
# Rebuild any NixOS hosts that reference pbsIp or pxeServerIp to pick up
# the updated variables.nix values (pxe-boot mandatory; others as convenient)
```
---
## Rollback summary
| What broke | How to roll back |
|---|---|
| FreeIPA DNS not resolving | Check `systemctl status named` on DC; restart if failed |
| FreeIPA DNS unreachable | `pct start 100` on pve1 (restores Pi-hole) |
| Router DHCP not handing out leases | Re-enable Pi-hole DHCP; disable router DHCP |
| DC unreachable after IP change | Proxmox console on VM 108 → `nmcli connection up "System eth0"` with old IP |
| Host unreachable after renumber | Proxmox console → revert IP; or `pct set <id> --net0 ...` old IP and reboot CT |
| pve1 web UI gone after renumber | SSH to .245 and check `/etc/network/interfaces`; if wrong, fix and restart networking |
+22 -20
View File
@@ -6,6 +6,10 @@ let
tftpRoot = "${pxeRoot}/tftp"; tftpRoot = "${pxeRoot}/tftp";
pxeBaseUrl = "http://${vars.pxeServerIp}"; pxeBaseUrl = "http://${vars.pxeServerIp}";
# Base network address extracted from lanCidr (e.g. "192.168.2.0" from
# "192.168.2.0/24") — used by dnsmasq's proxy DHCP range directive.
lanBaseAddr = lib.head (lib.splitString "/" vars.lanCidr);
bootIpxe = pkgs.writeText "boot.ipxe" '' bootIpxe = pkgs.writeText "boot.ipxe" ''
#!ipxe #!ipxe
@@ -73,16 +77,16 @@ let
boot boot
''; '';
# Kickstart file for domain-controller.sweet.home. # Kickstart file for ${vars.ipaServer}.
# Installs Rocky Linux 9, sets a static IP, creates wayne with the # Installs Rocky Linux 9, sets a static IP, creates ${vars.ipaUser} with
# admin SSH key, then on first reboot runs ipa-server-install via a # the admin SSH key, then on first reboot runs ipa-server-install via a
# systemd oneshot service. Passwords are generated at %post time, # systemd oneshot service. Passwords are generated at %post time, written
# written to /root/ipa-credentials.txt (chmod 600), and read back by # to /root/ipa-credentials.txt (chmod 600), and read back by the
# the first-boot script — never hardcoded here or in the repo. # first-boot script — never hardcoded here or in the repo.
rockyFreeIpaKs = pkgs.writeText "rocky-freeipa.ks" '' rockyFreeIpaKs = pkgs.writeText "rocky-freeipa.ks" ''
#version=RHEL9 #version=RHEL9
# Unattended Rocky Linux 9 + FreeIPA install # Unattended Rocky Linux 9 + FreeIPA install
# Target: domain-controller.${vars.homeDomain} ${vars.domainControllerIp} # Target: ${vars.ipaServer} ${vars.domainControllerIp}
url --url=${rockyMirror}/BaseOS/${rockyArch}/os/ url --url=${rockyMirror}/BaseOS/${rockyArch}/os/
repo --name=appstream --baseurl=${rockyMirror}/AppStream/${rockyArch}/os/ repo --name=appstream --baseurl=${rockyMirror}/AppStream/${rockyArch}/os/
@@ -93,14 +97,14 @@ let
# DHCP during install; static IP configured in %post via NM config file # DHCP during install; static IP configured in %post via NM config file
network --bootproto=dhcp --device=link --activate network --bootproto=dhcp --device=link --activate
network --hostname=domain-controller.sweet.home network --hostname=${vars.ipaServer}
selinux --enforcing selinux --enforcing
firewall --enabled --service=ssh firewall --enabled --service=ssh
rootpw --lock rootpw --lock
user --name=wayne --groups=wheel --shell=/bin/bash user --name=${vars.ipaUser} --groups=wheel --shell=/bin/bash
sshkey --username=wayne "${vars.adminSshKey}" sshkey --username=${vars.ipaUser} "${vars.adminSshKey}"
zerombr zerombr
clearpart --all --initlabel --drives=sda clearpart --all --initlabel --drives=sda
@@ -148,7 +152,7 @@ let
# -- /etc/hosts: FQDN must resolve to the real IP (not loopback) for IPA -- # -- /etc/hosts: FQDN must resolve to the real IP (not loopback) for IPA --
sed -i '/domain-controller/d' /etc/hosts sed -i '/domain-controller/d' /etc/hosts
echo '${vars.domainControllerIp} domain-controller.${vars.homeDomain} domain-controller' >> /etc/hosts echo '${vars.domainControllerIp} ${vars.ipaServer} domain-controller' >> /etc/hosts
# -- Generate IPA passwords and store securely -- # -- Generate IPA passwords and store securely --
DM_PASS=$(openssl rand -base64 24 | tr -dc 'A-Za-z0-9' | head -c 24) DM_PASS=$(openssl rand -base64 24 | tr -dc 'A-Za-z0-9' | head -c 24)
@@ -168,13 +172,13 @@ let
ADMIN_PASS=$(grep '^IPA Admin:' /root/ipa-credentials.txt | awk '{print $NF}') ADMIN_PASS=$(grep '^IPA Admin:' /root/ipa-credentials.txt | awk '{print $NF}')
ipa-server-install \ ipa-server-install \
--realm=SWEET.HOME \ --realm=${lib.strings.toUpper vars.homeDomain} \
--domain=sweet.home \ --domain=${vars.homeDomain} \
--hostname=domain-controller.sweet.home \ --hostname=${vars.ipaServer} \
--ds-password="$DM_PASS" \ --ds-password="$DM_PASS" \
--admin-password="$ADMIN_PASS" \ --admin-password="$ADMIN_PASS" \
--setup-dns \ --setup-dns \
--forwarder=192.168.2.253 \ --forwarder=${vars.domainControllerIp} \
--no-dnssec-validation \ --no-dnssec-validation \
--no-ntp \ --no-ntp \
--unattended --unattended
@@ -341,9 +345,7 @@ in
atftpd = { atftpd = {
enable = true; enable = true;
root = tftpRoot; root = tftpRoot;
extraOptions = [ extraOptions = [ "--verbose=5" ];
"--verbose=5"
];
}; };
openssh.settings.PermitRootLogin = "yes"; openssh.settings.PermitRootLogin = "yes";
@@ -428,7 +430,7 @@ in
# Without this dnsmasq tries to bind port 53 which systemd-resolved # Without this dnsmasq tries to bind port 53 which systemd-resolved
# already owns, causing startup failure. # already owns, causing startup failure.
port = 0; port = 0;
dhcp-range = [ "192.168.2.0,proxy" ]; dhcp-range = [ "${lanBaseAddr},proxy" ];
dhcp-match = [ dhcp-match = [
"set:ipxe,175" "set:ipxe,175"
"set:efi64,option:client-arch,7" "set:efi64,option:client-arch,7"
@@ -445,5 +447,5 @@ in
}; };
networking.firewall.allowedTCPPorts = [ vars.ports.pxeBootHttp ]; networking.firewall.allowedTCPPorts = [ vars.ports.pxeBootHttp ];
networking.firewall.allowedUDPPorts = [ vars.ports.pxeBootTftp 67 ]; networking.firewall.allowedUDPPorts = [ vars.ports.pxeBootTftp vars.ports.dhcp ];
} }
+19 -40
View File
@@ -5,13 +5,13 @@ let
sudo nixos-rebuild switch \ sudo nixos-rebuild switch \
--no-write-lock-file \ --no-write-lock-file \
--refresh \ --refresh \
--flake git+https://${vars.lanDomain}/beatzaplenty/nixos.git#$(cat /etc/flake-target) --flake git+https://${vars.giteaDomain}/${vars.giteaRepoPath}.git#$(cat /etc/flake-target)
''; '';
testCmd = '' testCmd = ''
sudo nixos-rebuild test \ sudo nixos-rebuild test \
--no-write-lock-file \ --no-write-lock-file \
--refresh \ --refresh \
--flake git+https://${vars.lanDomain}/beatzaplenty/nixos.git#$(cat /etc/flake-target) --flake git+https://${vars.giteaDomain}/${vars.giteaRepoPath}.git#$(cat /etc/flake-target)
''; '';
buildImageFn = '' buildImageFn = ''
buildImage() { buildImage() {
@@ -39,21 +39,18 @@ in
}; };
interactiveShellInit = buildImageFn; interactiveShellInit = buildImageFn;
}; };
networking.networkmanager.enable = true; networking.networkmanager.enable = true;
# Recommended over the true default (bypasses ZFS's own import safeguards) # Recommended over the true default (bypasses ZFS's own import safeguards)
# per the option's own docs; matches hosts/docker/host.nix and # per the option's own docs; matches hosts/docker/host.nix and
# modules/services/zfs/enable-service.nix, which already set this # modules/services/zfs/enable-service.nix. Harmless no-op on hosts without ZFS.
# explicitly. Harmless no-op on hosts that don't use ZFS at all.
boot.zfs.forceImportRoot = false; boot.zfs.forceImportRoot = false;
# Set your time zone.
time.timeZone = vars.timeZone; time.timeZone = vars.timeZone;
# Enable QEMU agent
services.qemuGuest.enable = true; services.qemuGuest.enable = true;
# Enable docker-compose
environment.systemPackages = with pkgs; [ environment.systemPackages = with pkgs; [
vim vim
btop btop
@@ -63,11 +60,10 @@ in
]; ];
# Secrets shared by every host, decrypted at activation via each host's # Secrets shared by every host, decrypted at activation via each host's
# existing SSH host key (sops-nix derives the age key from # SSH host key (sops-nix derives the age key from
# /etc/ssh/ssh_host_ed25519_key automatically — see modules/common/README # /etc/ssh/ssh_host_ed25519_key automatically). hashedPassword secrets need
# or docs/ for the sops workflow). hashedPassword/hashedPasswordFile need
# neededForUsers so they're available before the normal secret-activation # neededForUsers so they're available before the normal secret-activation
# step, since user creation happens very early in boot. # step user creation happens very early in boot.
sops = { sops = {
defaultSopsFile = ../../secrets/common.yaml; defaultSopsFile = ../../secrets/common.yaml;
@@ -77,9 +73,9 @@ in
"nix-github-token" = { }; "nix-github-token" = { };
}; };
# nix.conf doesn't support a *File-style option for access-tokens, so the # nix.conf has no *File-style option for access-tokens, so the token is
# token is rendered into a runtime-only file (never touches the Nix store) # rendered into a runtime-only file (never touches the Nix store) and
# and pulled in via nix.conf's native !include directive. # pulled in via nix.conf's native !include directive.
templates."nix-github-token.conf".content = '' templates."nix-github-token.conf".content = ''
access-tokens = github.com=${config.sops.placeholder."nix-github-token"} access-tokens = github.com=${config.sops.placeholder."nix-github-token"}
''; '';
@@ -90,12 +86,11 @@ in
''; '';
users = { users = {
# With mutableUsers = false, update-users-groups.pl enforces hashedPasswordFile # mutableUsers = false makes update-users-groups.pl enforce hashedPasswordFile
# on every activation regardless of whether the account already exists in # on every activation, not just on newly-created accounts. Without this, a
# /etc/shadow. The default (true) only applies hashedPasswordFile to newly- # freshly-built proxmox disk image (activation runs without a usable sops key,
# created accounts — which means a freshly-built proxmox disk image (where # so both accounts land in shadow with '!') will never have its passwords fixed
# activation runs without a usable sops key, so both accounts land in shadow # by subsequent boots.
# with !) will never have its passwords fixed by subsequent boots.
mutableUsers = false; mutableUsers = false;
users.root = { users.root = {
@@ -104,39 +99,23 @@ in
users.${vars.primaryUser} = { users.${vars.primaryUser} = {
isNormalUser = true; isNormalUser = true;
extraGroups = [ "wheel" ]; # Enable sudo for the user. extraGroups = [ "wheel" ];
packages = with pkgs; [ packages = with pkgs; [ tree ];
tree
];
hashedPasswordFile = config.sops.secrets."nixos-hashedPassword".path; hashedPasswordFile = config.sops.secrets."nixos-hashedPassword".path;
openssh.authorizedKeys.keys = [ openssh.authorizedKeys.keys = [ vars.adminSshKey ] ++ vars.extraAdminSshKeys;
vars.adminSshKey
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICMJhrfFayLBG+gWtO6oAvgambw5nWWgztiTFEaaaVRH debian@surface"
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIGygkCljN6uKpdJbHTOQtn8ZnH+wKXDLAwrDFbLrE/65 nixos@nixos"
];
}; };
}; };
# Enable the OpenSSH daemon.
services.openssh.enable = true; services.openssh.enable = true;
#Enable flakes
nix.settings = { nix.settings = {
experimental-features = [ "nix-command" "flakes" ]; experimental-features = [ "nix-command" "flakes" ];
auto-optimise-store = true; auto-optimise-store = true;
}; };
programs.git = { programs.git = {
enable = true; enable = true;
package = pkgs.git; package = pkgs.git;
config = { config.credential.helper = "store";
credential.helper = "store";
};
}; };
} }
+2 -6
View File
@@ -42,12 +42,8 @@ let
''; '';
in in
{ {
# Root SSH access — same key set as nixos user so all admin keys can reach root. # Root SSH access — same key set as the nixos user so all admin keys can reach root.
users.users.root.openssh.authorizedKeys.keys = [ users.users.root.openssh.authorizedKeys.keys = [ vars.adminSshKey ] ++ vars.extraAdminSshKeys;
vars.adminSshKey
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICMJhrfFayLBG+gWtO6oAvgambw5nWWgztiTFEaaaVRH debian@surface"
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIGygkCljN6uKpdJbHTOQtn8ZnH+wKXDLAwrDFbLrE/65 nixos@nixos"
];
# Passwordless sudo for wheel — operator SSHes as nixos and uses sudo for # Passwordless sudo for wheel — operator SSHes as nixos and uses sudo for
# cluster management commands (drbdadm, crm*, pcs, etc.) # cluster management commands (drbdadm, crm*, pcs, etc.)
+4 -4
View File
@@ -1,10 +1,10 @@
{ config, lib, vars, ... }: { config, lib, vars, ... }:
let let
# FQDN of the LAN NFS VIP (Pacemaker vip-lan, 192.168.2.229). Using the # FQDN of the LAN NFS VIP (Pacemaker vip-lan, 192.168.2.229). Defined in
# FQDN rather than a raw IP or bare hostname avoids systemd-resolved LLMNR # variables.nix as haLanNfsFqdn; using the FQDN avoids systemd-resolved
# quirks and survives a future VIP renumber via a DNS-only update. # LLMNR quirks and survives a future VIP renumber via a DNS-only update.
nfsServer = "ha-vip-lan.${vars.homeDomain}"; nfsServer = vars.haLanNfsFqdn;
in in
{ {
fileSystems.${vars.nfsShares.pxebootImages.mountpoint} = { fileSystems.${vars.nfsShares.pxebootImages.mountpoint} = {
+18 -25
View File
@@ -3,12 +3,11 @@
{ {
# Run dnsmasq on the LAN interface as a forwarding-only resolver for # Run dnsmasq on the LAN interface as a forwarding-only resolver for
# *.ts.net (Tailscale MagicDNS names). FreeIPA's bind-dyndb-ldap # *.ts.net (Tailscale MagicDNS names). FreeIPA's bind-dyndb-ldap
# cannot reach 100.100.100.100 (Tailscale's internal resolver) directly # cannot reach vars.tailscaleResolverIp directly because the DC is not a
# because the DC is not a Tailscale node. This host IS a Tailscale node # Tailscale node. This host IS a Tailscale node and can reach it via
# and can reach 100.100.100.100 via its tailscale0 interface, so it # tailscale0, so it acts as an intermediary: FreeIPA has a conditional
# acts as an intermediary: FreeIPA has a conditional forward zone for # forward zone for ts.net pointing here (vars.tailscaleRouterIp), and this
# ts.net pointing here (vars.tailscaleRouterIp), and this dnsmasq # dnsmasq instance forwards those queries onward to Tailscale's resolver.
# instance forwards those queries onward to Tailscale's resolver.
# #
# Configure FreeIPA once after deploying this host: # Configure FreeIPA once after deploying this host:
# kinit admin # kinit admin
@@ -29,32 +28,26 @@
resolveLocalQueries = false; resolveLocalQueries = false;
settings = { settings = {
# Listen only on the LAN interface — not tailscale0 or loopback. # Listen only on the LAN interface — not tailscale0 or loopback.
# bind-interfaces prevents dnsmasq from binding to 0.0.0.0 and # bind-interfaces prevents dnsmasq from binding to 0.0.0.0 and then
# then filtering by interface later; combined with `interface` this # filtering by interface later; combined with `interface` this ensures
# ensures it genuinely listens only on eth0. # it genuinely listens only on eth0.
bind-interfaces = true; bind-interfaces = true;
interface = [ vars.lxcLanInterface ]; interface = [ vars.lxcLanInterface ];
# Forward-only: no local /etc/hosts or /etc/resolv.conf reading, # Forward-only: no local /etc/hosts or /etc/resolv.conf reading, no
# no negative caching of NXDOMAIN for names this instance doesn't # negative caching of NXDOMAIN for names this instance doesn't serve.
# serve. All ts.net queries come from FreeIPA's conditional forwarder # All ts.net queries come from FreeIPA's conditional forwarder and must
# and must be answered by Tailscale's resolver. # be answered by Tailscale's resolver.
no-hosts = true; no-hosts = true;
no-resolv = true; no-resolv = true;
# Tailscale's internal "Quad100" resolver — reachable from any # Forward *.tailnetDomain to Tailscale's internal resolver, scoped to
# Tailscale node via the tailscale0 interface. Scoped to the # the tailnet-specific subdomain rather than all of ts.net (FreeIPA
# specific tailnet subdomain (vars.tailnetDomain) rather than # refuses to shadow ts.net, a real public TLD).
# all of ts.net: FreeIPA refuses to shadow ts.net (a real public server = [ "/${vars.tailnetDomain}/${vars.tailscaleResolverIp}" ];
# TLD with DNSimple nameservers) so the conditional forward zone
# in FreeIPA must use the tailnet-specific subdomain instead:
# ipa dnsforwardzone-add ${vars.tailnetDomain} \
# --forwarder=${vars.tailscaleRouterIp} \
# --forward-policy=only
server = [ "/${vars.tailnetDomain}/100.100.100.100" ];
}; };
}; };
networking.firewall.allowedUDPPorts = [ 53 ]; networking.firewall.allowedUDPPorts = [ vars.ports.dns ];
networking.firewall.allowedTCPPorts = [ 53 ]; networking.firewall.allowedTCPPorts = [ vars.ports.dns ];
} }
+173 -147
View File
@@ -1,43 +1,72 @@
{ rec {
# Network / domains # ── Gitea / flake remote ──────────────────────────────────────────────────
lanDomain = "gitea.lan.ddnsgeek.com"; # Gitea/DDNS domain
homeDomain = "sweet.home"; # base LAN domain for service subdomains (pve., docker.) # External Gitea/DDNS domain — used only for the remote flake URL in
tailnetDomain = "tail13f623.ts.net"; # Tailscale MagicDNS suffix # Switch-nix / Test-nix aliases (modules/common/configuration.nix).
lanCidr = "192.168.2.0/24"; # LAN subnet giteaDomain = "gitea.lan.ddnsgeek.com";
lanGateway = "192.168.2.254"; # LAN default gateway (router)
lanPrefixLength = 24; # LAN subnet prefix length (/24 = 255.255.255.0) # Org/repo path within Gitea, combined with giteaDomain to form the
lxcLanInterface = "eth0"; # LAN NIC name in LXC containers (set by Proxmox --net0 name=eth0) # git+https:// URL used by Switch-nix / Test-nix.
lxcStorageInterface = "eth1"; # storage-client NIC name in LXC containers (vmbr2, --net1) giteaRepoPath = "beatzaplenty/nixos";
vmLanInterface = "ens18"; # LAN NIC name in Proxmox VMs (virtio, first NIC)
vmStorageInterface = "ens19"; # cluster-internal NIC in HA VMs (vmbr1 — DRBD + Corosync only) # ── Network ───────────────────────────────────────────────────────────────
# Base LAN domain for service subdomains (pve., docker., nix-cache., …)
homeDomain = "sweet.home";
# Tailscale MagicDNS suffix for this tailnet
tailnetDomain = "tail13f623.ts.net";
lanCidr = "192.168.2.0/24";
lanGateway = "192.168.2.254";
lanPrefixLength = 24;
# NIC names inside guests — determined by the hypervisor/platform, not the OS.
lxcLanInterface = "eth0"; # LAN NIC in LXC containers (Proxmox --net0 name=eth0)
lxcStorageInterface = "eth1"; # storage-client NIC in LXC containers (vmbr2, --net1)
vmLanInterface = "ens18"; # LAN NIC in Proxmox VMs (virtio, first NIC)
vmStorageInterface = "ens19"; # cluster-internal NIC in HA VMs (vmbr1 — DRBD + Corosync)
vmStorageClientInterface = "ens20"; # storage-client NIC in HA VMs (vmbr2 — iSCSI/NFS VIP) vmStorageClientInterface = "ens20"; # storage-client NIC in HA VMs (vmbr2 — iSCSI/NFS VIP)
# ── Host IPs ──────────────────────────────────────────────────────────────
pxeServerIp = "192.168.2.223"; # pxe-boot LXC container LAN IP pxeServerIp = "192.168.2.223"; # pxe-boot LXC container LAN IP
nixCacheIp = "192.168.2.224"; # nix-cache LXC container LAN IP nixCacheIp = "192.168.2.224"; # nix-cache LXC container LAN IP
tailscaleRouterIp = "192.168.2.222"; # tailscale-router LXC container LAN IP tailscaleRouterIp = "192.168.2.222"; # tailscale-router LXC container LAN IP
torRelayIp = "192.168.2.221"; # tor-relay LXC container LAN IP torRelayIp = "192.168.2.221"; # tor-relay LXC container LAN IP
dockerIp = "192.168.2.225"; # docker Proxmox VM LAN IP dockerIp = "192.168.2.225"; # docker Proxmox VM LAN IP
pbsIp = "192.168.2.244"; # Proxmox Backup Server LAN IP (not NixOS-managed) pbsIp = "192.168.2.244"; # Proxmox Backup Server LAN IP (not NixOS-managed)
domainControllerIp = "192.168.2.253"; # FreeIPA domain controller — authoritative DNS for sweet.home (not NixOS-managed) domainControllerIp = "192.168.2.253"; # FreeIPA — authoritative DNS for sweet.home (not NixOS-managed)
ipaServer = "domain-controller.sweet.home"; # FreeIPA server hostname (used by security.ipa and Kerberos; must be a resolvable FQDN, not an IP)
# FreeIPA server FQDN used by security.ipa and Kerberos. Must be a
# resolvable name (not an IP); resolves to domainControllerIp.
ipaServer = "domain-controller.${homeDomain}";
# ── Cross-host references ─────────────────────────────────────────────────
# Cross-host references (LAN hostnames/users other hosts reach over the network)
nixCacheHost = "nix-cache"; # substituter/remote-builder hostname nixCacheHost = "nix-cache"; # substituter/remote-builder hostname
dockerHost = "docker"; # docker-compose stack host dockerHost = "docker"; # docker-compose stack host
# Raspberry Pi's own Tailscale hostname (not fronted by `server` — it # Raspberry Pi's own Tailscale hostname (not fronted by any server — it
# exports its own NFS share directly). Resolved as # exports its own NFS share directly). Resolved as
# "${raspberryPiHost}.${tailnetDomain}" in modules/raspi/mount-data.nix. # "${raspberryPiHost}.${tailnetDomain}" in modules/raspi/mount-data.nix.
raspberryPiHost = "raspberrypi"; raspberryPiHost = "raspberrypi";
remoteBuilderUser = "nixremote"; # remote builder SSH user remoteBuilderUser = "nixremote";
# nix-cache's own SSH host public key (not a secret — the private half # Tailscale's internal "Quad100" DNS resolver, reachable from any Tailscale
# never leaves the host). Wired into every client's # node via tailscale0. Used by modules/tailscale/ts-dns-forwarder.nix to
# programs.ssh.knownHosts by modules/nix-cache/remote-builder-client.nix # forward *.tailnetDomain queries on behalf of FreeIPA's conditional
# so distributed builds don't hit "Host key verification failed" on a # forwarder zone.
# fresh client that has never manually ssh'd to nix-cache before. Update tailscaleResolverIp = "100.100.100.100";
# this if nix-cache's host key is ever rotated or the host is rebuilt
# from scratch. # ── SSH keys ──────────────────────────────────────────────────────────────
# nix-cache's SSH host public key (not a secret — private half never leaves
# the host). Wired into every client's programs.ssh.knownHosts by
# modules/nix-cache/remote-builder-client.nix so distributed builds don't
# hit "Host key verification failed" on a fresh client. Update if nix-cache
# is ever rebuilt with a new host key.
nixCacheHostKey = "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICuHUxGNH6ei3BZD+EfZs3l4X8uJNcjQiOsM/G4yo4O/ lxc-nix-cache"; nixCacheHostKey = "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICuHUxGNH6ei3BZD+EfZs3l4X8uJNcjQiOsM/G4yo4O/ lxc-nix-cache";
# Beszel hub's SSH public key — used by every agent to authenticate the # Beszel hub's SSH public key — used by every agent to authenticate the
@@ -46,8 +75,8 @@
beszelHubKey = "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIFPR9kwtC4TAeTRu46A7+opZsYpxqkRJ+x/ZyB2GWCeG"; beszelHubKey = "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIFPR9kwtC4TAeTRu46A7+opZsYpxqkRJ+x/ZyB2GWCeG";
# Public keys authorized to SSH in as remoteBuilderUser on the nix-cache # Public keys authorized to SSH in as remoteBuilderUser on the nix-cache
# host (modules/nix-cache/server.nix) — one per client host that's allowed # host (modules/nix-cache/server.nix) — one per client host allowed to use
# to use it as a distributed builder. # it as a distributed builder.
remoteBuilderAuthorizedKeys = [ remoteBuilderAuthorizedKeys = [
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIK+ioWPhHixlgCB9KIQ0QTHTz6A+Oo2F3uKiINLip5rO root@docker" "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIK+ioWPhHixlgCB9KIQ0QTHTz6A+Oo2F3uKiINLip5rO root@docker"
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICMJhrfFayLBG+gWtO6oAvgambw5nWWgztiTFEaaaVRH debian@surface" "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICMJhrfFayLBG+gWtO6oAvgambw5nWWgztiTFEaaaVRH debian@surface"
@@ -61,156 +90,162 @@
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIMseQwPpmaa6cgV5U8KhUsiVSYARG85zGa9rho0LJWks wayne@pve1" "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIMseQwPpmaa6cgV5U8KhUsiVSYARG85zGa9rho0LJWks wayne@pve1"
]; ];
# Admin SSH public key, authorized on the primary user of every host and # Primary admin SSH public key, authorized on the primary user of every
# the installer image's nixos/root users. # host and the installer image's nixos/root users.
adminSshKey = "ssh-rsa AAAAB3NzaC1yc2EAAAADAQABAAABgQCq/Q5LvIXlZwO2kdeAN5nLGZ59nZB7JHYMEszHxmNtGMzv1lM31jiPNsr0z2EKVZhE7OOfa2IF9rhWYD7JUA9G0yzdZ4WTXFNGVVOJoOVH6vAF3XCxoVilOEwTc7h2Wiy+rzd0B28/3spffzQQWJhY6GRQVa8j+6xAGF60Fcvl1vLosYT9Bn2ZbK4TCWOwAn2jqXIieGpZdn/UNZbGOeKRiCvhktDfMAzuQzN/9jMu/oF4pkPn2X1UrsQdNlvp0Ci8md612MozIpncQJyAF1ADhunr3sMx0isUXiqD29R5DS4TftpekqLNLak+zcxFa8N7DcRNp3DcKfJvyTkwQrR4r+b7lFLYOLHLagSso9CzeW/paAS2q9I5SBm/2DtE1diLLg2jZikYcstsu/G5RgvbzbKqjiaMwTdXC3AMvDxQrs7U5pDRZFzoofG3cpODbTm+uy3m0kP70z0M1K45UbDG0p+itnTu9x40JbQEgefbx38AItNvAIx1A8HO4I1VX28= wayne@stream"; adminSshKey = "ssh-rsa AAAAB3NzaC1yc2EAAAADAQABAAABgQCq/Q5LvIXlZwO2kdeAN5nLGZ59nZB7JHYMEszHxmNtGMzv1lM31jiPNsr0z2EKVZhE7OOfa2IF9rhWYD7JUA9G0yzdZ4WTXFNGVVOJoOVH6vAF3XCxoVilOEwTc7h2Wiy+rzd0B28/3spffzQQWJhY6GRQVa8j+6xAGF60Fcvl1vLosYT9Bn2ZbK4TCWOwAn2jqXIieGpZdn/UNZbGOeKRiCvhktDfMAzuQzN/9jMu/oF4pkPn2X1UrsQdNlvp0Ci8md612MozIpncQJyAF1ADhunr3sMx0isUXiqD29R5DS4TftpekqLNLak+zcxFa8N7DcRNp3DcKfJvyTkwQrR4r+b7lFLYOLHLagSso9CzeW/paAS2q9I5SBm/2DtE1diLLg2jZikYcstsu/G5RgvbzbKqjiaMwTdXC3AMvDxQrs7U5pDRZFzoofG3cpODbTm+uy3m0kP70z0M1K45UbDG0p+itnTu9x40JbQEgefbx38AItNvAIx1A8HO4I1VX28= wayne@stream";
# Additional SSH keys granted access alongside adminSshKey on every host
# (modules/common/configuration.nix) and on HA cluster root
# (modules/ha/cluster-config.nix). Single definition here prevents the
# two modules from drifting out of sync.
extraAdminSshKeys = [
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAICMJhrfFayLBG+gWtO6oAvgambw5nWWgztiTFEaaaVRH debian@surface"
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIGygkCljN6uKpdJbHTOQtn8ZnH+wKXDLAwrDFbLrE/65 nixos@nixos"
];
# ── Wifi ──────────────────────────────────────────────────────────────────
# Prestaged wifi SSID for the gui host's NetworkManager profile # Prestaged wifi SSID for the gui host's NetworkManager profile
# (modules/networking/wifi.nix). The password is not here -- it's # (modules/networking/wifi.nix). Password is sops-encrypted in
# sops-encrypted in secrets/gui.yaml (wifi-password) instead, since this # secrets/gui.yaml (wifi-password) — not stored here.
# file isn't a secret store.
wifiSsid = "nbn-fttp-net-5G"; wifiSsid = "nbn-fttp-net-5G";
# Bare-metal gui host's two disks for a ZFS RAID0 (striped) root pool # ── Bare-metal GUI host ───────────────────────────────────────────────────
# (modules/disko/baremetal.nix). Only used transiently at disko-format
# time (partitioning); the resulting fileSystems/zpool import reference # Two disks for the ZFS RAID0 (striped) root pool on baremetal-gui
# by-partlabel/by-id paths afterward regardless, same as # (modules/disko/baremetal.nix). Only referenced at disko-format time;
# modules/disko/proxmox.nix's own plain "/dev/sda". # afterward the pool imports by-partlabel/by-id paths regardless.
guiRootDisk1 = "/dev/sda"; guiRootDisk1 = "/dev/sda";
guiRootDisk2 = "/dev/sdb"; guiRootDisk2 = "/dev/sdb";
# System # ── System / users ────────────────────────────────────────────────────────
timeZone = "Australia/Brisbane"; timeZone = "Australia/Brisbane";
# Main interactive user on every host. Every module that grants this user # Main interactive user on every host. Modules that grant this user a
# a group, a home directory, or tmpfiles ownership should reference # group, home directory, or tmpfiles ownership reference this so a rename
# vars.primaryUser rather than the literal "nixos", so renaming it is a # is a one-line change here.
# one-line change.
primaryUser = "nixos"; primaryUser = "nixos";
# Primary IPA/domain user. Home Manager is configured for this user on every # Primary IPA/domain user. Home Manager is configured for this user on
# IPA-enrolled host (see modules/ipa/client.nix) to manage the environment # every IPA-enrolled host (modules/ipa/client.nix).
# that IPA itself doesn't cover: dotfiles, user packages, session variables.
ipaUser = "wayne"; ipaUser = "wayne";
# GID of the IPA "docker-access" group (GID 50010 on the IPA server). # GID of the IPA "docker-access" group (GID 50010 on the IPA server). The
# The local "docker" group is pinned to this GID on every host that runs # local "docker" group is pinned to this GID on every Docker host so IPA
# Docker so that IPA group membership alone grants docker socket access - # group membership alone grants socket access — no per-host
# no per-host users.groups.docker.members entry for the IPA user needed. # users.groups.docker.members entries needed.
dockerAccessGid = 50010; dockerAccessGid = 50010;
# HA file server cluster # ── HA file-server cluster ────────────────────────────────────────────────
# LAN IPs (vmbr0 / ens18) — management only after storage migration. #
# Cluster IPs (vmbr1 / ens19) — VLAN 10 (192.168.10.x), isolated internal bridge, # Three network segments, all internal to pve1:
# DRBD replication and Corosync heartbeat only; never leaves pve1. # LAN VLAN 2 / vmbr0 / 192.168.2.x — management only
# Storage-client IPs (vmbr2 / ens20) — VLAN 20 (192.168.20.x), isolated internal # Cluster VLAN 10 / vmbr1 / 192.168.10.x — DRBD replication + Corosync ring0
# bridge for iSCSI; docker and server VMs connect here instead of crossing vmbr0. # Storage-client VLAN 20 / vmbr2 / 192.168.20.x — iSCSI + NFS client access
# haServerVip: floating virtual IP on vmbr2, managed by Pacemaker IPaddr2; #
# iSCSI clients connect here regardless of which node is Active. # The host octet is consistent across subnets: node1 = .228, node2 = .227,
# Protocol separation: iSCSI on storage-client subnet (VLAN 20) only; # VIP = .229 everywhere.
# NFS on LAN subnet (VLAN 2) only. Enforced by firewall on the HA nodes. #
# Protocol separation (firewall-enforced on HA nodes):
# NFS — both subnets; LAN VIP for pxe-boot/LAN clients, storage VIP for docker
# iSCSI — storage-client subnet only
haServer1Host = "ha-server-1"; haServer1Host = "ha-server-1";
haServer2Host = "ha-server-2"; haServer2Host = "ha-server-2";
haServer1Ip = "192.168.2.228"; # LAN IP, node 1 (vmbr0 / ens18) haServer1Ip = "192.168.2.228"; # LAN IP, node 1 (vmbr0 / ens18)
haServer2Ip = "192.168.2.227"; # LAN IP, node 2 (vmbr0 / ens18) haServer2Ip = "192.168.2.227"; # LAN IP, node 2 (vmbr0 / ens18)
haServer1StorageIp = "192.168.10.228"; # cluster-net IP, node 1 (vmbr1 / ens19, VLAN 10) haServer1StorageIp = "192.168.10.228"; # cluster-net IP, node 1 (vmbr1 / ens19, VLAN 10)
haServer2StorageIp = "192.168.10.227"; # cluster-net IP, node 2 (vmbr1 / ens19, VLAN 10) haServer2StorageIp = "192.168.10.227"; # cluster-net IP, node 2 (vmbr1 / ens19, VLAN 10)
haStorageCidr = "192.168.10.224/29"; # cluster subnet — VLAN 10, internal to pve1 only haStorageCidr = "192.168.10.224/29"; # cluster subnet — VLAN 10, internal to pve1
haStoragePrefixLength = 29; # cluster subnet prefix length (/29) haStoragePrefixLength = 29;
haServer1ClientIp = "192.168.20.228"; # storage-client IP, node 1 (vmbr2 / ens20, VLAN 20) haServer1ClientIp = "192.168.20.228"; # storage-client IP, node 1 (vmbr2 / ens20, VLAN 20)
haServer2ClientIp = "192.168.20.227"; # storage-client IP, node 2 (vmbr2 / ens20, VLAN 20) haServer2ClientIp = "192.168.20.227"; # storage-client IP, node 2 (vmbr2 / ens20, VLAN 20)
haServerVip = "192.168.20.229"; # storage-client floating VIP on vmbr2 (Pacemaker IPaddr2 vip-storage, VLAN 20) haServerVip = "192.168.20.229"; # storage-client floating VIP (Pacemaker vip-storage, VLAN 20)
haServerLanVip = "192.168.2.229"; # LAN floating VIP on vmbr0 (Pacemaker IPaddr2 vip-lan) — NFS access haServerLanVip = "192.168.2.229"; # LAN floating VIP (Pacemaker vip-lan) — NFS for LAN clients
dockerStorageIp = "192.168.20.225"; # docker CT storage-client IP (vmbr2 / eth1, VLAN 20) dockerStorageIp = "192.168.20.225"; # docker CT storage-client IP (vmbr2 / eth1, VLAN 20)
haClientCidr = "192.168.20.0/24"; # storage-client subnet — VLAN 20, internal to pve1 only haClientCidr = "192.168.20.0/24"; # storage-client subnet — VLAN 20, internal to pve1
haClientPrefixLength = 24; # storage-client subnet prefix length (/24) haClientPrefixLength = 24;
haStorageRoot = "/srv/ha-data"; # XFS-over-DRBD mount point on the Active node haStorageRoot = "/srv/ha-data"; # XFS-over-DRBD mount point on the Active node
haStorageNfsFqdn = "nfs.storage.home"; # NFS VIP FQDN (storage.home zone) — resolves to haServerVip; use this in fileSystems device strings
# NFS VIP FQDNs — use these in fileSystems device strings so mounts
# survive a future VIP renumber via a DNS-only update, not a NixOS rebuild.
haStorageNfsFqdn = "nfs.storage.home"; # storage-client VIP (VLAN 20) — docker + future swarm
haLanNfsFqdn = "ha-vip-lan.${homeDomain}"; # LAN VIP (VLAN 2) — pxe-boot + other LAN clients
haIscsiIqn = "iqn.2026-01.home.sweet:ha-storage"; haIscsiIqn = "iqn.2026-01.home.sweet:ha-storage";
# DRBD backing disk — identified by SCSI controller path so it resolves to the
# correct block device regardless of OS-level naming (sda vs sdb can differ # DRBD backing disk — identified by SCSI controller path so it resolves to
# between Proxmox VMs depending on disk-add order). drive-scsi1 is always the # the correct block device regardless of OS-level naming (sda vs sdb can
# dedicated data disk on all HA nodes; drive-scsi0 is the OS disk. # differ between VMs depending on disk-add order). drive-scsi1 is always
# the data disk; drive-scsi0 is the OS disk.
haServerDrbdDisk = "/dev/disk/by-id/scsi-0QEMU_QEMU_HARDDISK_drive-scsi1"; haServerDrbdDisk = "/dev/disk/by-id/scsi-0QEMU_QEMU_HARDDISK_drive-scsi1";
# Storage # ── Storage / NFS ─────────────────────────────────────────────────────────
# NFS share definitions — used by ha-server.nix (exports), docker/mount-data.nix, # NFS share definitions — used by ha-server.nix (exports), docker/mount-data.nix,
# and pxe-boot/mount-pxe-images.nix (mounts). `subpath` is relative to # and pxe-boot/mount-pxe-images.nix (mounts). `subpath` is relative to
# haStorageRoot; `mountpoint` is the absolute local path on each client. # haStorageRoot; `mountpoint` is the absolute local path on each client.
# Renaming a share only needs changing it here — exports and all client # Renaming a share only requires changing it here — exports and all client
# mounts follow automatically. # mounts follow automatically.
nfsShares = { nfsShares = {
options = "(rw,sync,no_subtree_check,no_root_squash)"; options = "(rw,sync,no_subtree_check,no_root_squash)";
dockerConfig = { dockerConfig = { subpath = "docker/config"; mountpoint = "/mnt/docker/config"; };
subpath = "docker/config"; dockerDatabases = { subpath = "docker/databases"; mountpoint = "/mnt/docker/databases"; };
mountpoint = "/mnt/docker/config"; dockerVolumes = { subpath = "docker/volumes"; mountpoint = "/mnt/docker/volumes"; };
}; nextcloudData = { subpath = "docker/nextcloud-data"; mountpoint = "/mnt/nextcloud-data"; };
dockerDatabases = { raspiVolumes = { subpath = "raspi/volumes"; mountpoint = "/mnt/raspi-backup"; };
subpath = "docker/databases"; proxmoxIsos = { subpath = "proxmox/iso"; mountpoint = "/mnt/iso"; };
mountpoint = "/mnt/docker/databases"; proxmoxLxcImages = { subpath = "proxmox/lxc"; mountpoint = "/mnt/lxc"; };
}; pxebootImages = { subpath = "pxe-boot/images"; mountpoint = "/mnt/pxe-images"; };
dockerVolumes = {
subpath = "docker/volumes";
mountpoint = "/mnt/docker/volumes";
};
nextcloudData = {
subpath = "docker/nextcloud-data";
mountpoint = "/mnt/nextcloud-data";
};
raspiVolumes = {
subpath = "raspi/volumes";
mountpoint = "/mnt/raspi-backup";
};
proxmoxIsos = {
subpath = "proxmox/iso";
mountpoint = "/mnt/iso";
};
proxmoxLxcImages = {
subpath = "proxmox/lxc";
mountpoint = "/mnt/lxc";
};
pxebootImages = {
subpath = "pxe-boot/images";
mountpoint = "/mnt/pxe-images";
};
}; };
# The Raspberry Pi's own NFS export — not under storageRoot/nfsServerHost, # The Raspberry Pi's own NFS export — not under haStorageRoot, served
# served directly by the Pi itself over Tailscale (see raspberryPiHost # directly by the Pi over Tailscale (see raspberryPiHost) and mounted by
# above) and mounted at raspiMountpoint by modules/raspi/mount-data.nix. # modules/raspi/mount-data.nix.
raspiNfsPath = "/home/raspi/raspi"; raspiNfsPath = "/home/raspi/raspi";
raspiMountpoint = "/mnt/raspi"; raspiMountpoint = "/mnt/raspi";
# ── Ports ─────────────────────────────────────────────────────────────────
#
# Every literal port referenced from modules/ or hosts/, grouped by the # Every literal port referenced from modules/ or hosts/, grouped by the
# service/host that opens or connects to it — kept as separate entries # service that opens or connects to it. Kept as separate entries even where
# even where two happen to share a number today (e.g. nixCacheHttp and # two share a number today (e.g. nixCacheHttp and pxeBootHttp are both 80)
# pxeBootHttp are both 80) so changing one service's port can never # so changing one service's port never silently changes another.
# silently change an unrelated one.
ports = { ports = {
# nix-cache's nginx reverse proxy in front of nix-serve # nix-cache's nginx reverse proxy in front of nix-serve
# (modules/nix-cache/server.nix). # (modules/nix-cache/server.nix)
nixCacheHttp = 80; nixCacheHttp = 80;
# pxe-boot's nginx asset server, also used to build pxeBaseUrl # pxe-boot's nginx asset server; also used to build pxeBaseUrl
# (modules/build-types/pxe-boot.nix). # (modules/build-types/pxe-boot.nix)
pxeBootHttp = 80; pxeBootHttp = 80;
# pxe-boot's atftpd TFTP server — UDP, not TCP # pxe-boot's atftpd TFTP server — UDP (modules/build-types/pxe-boot.nix)
# (modules/build-types/pxe-boot.nix).
pxeBootTftp = 69; pxeBootTftp = 69;
# `server`'s NFS exports: portmapper (rpcbind), NFS data, and the # DHCP proxy port opened by dnsmasq on the pxe-boot host
# mountd RPC service (used by showmount/NFSv3 mount protocol). # (modules/build-types/pxe-boot.nix)
# Mountd listens on a fixed port so the firewall can whitelist it dhcp = 67;
# explicitly rather than opening all of rpcbind's dynamic range.
# All three need both TCP and UDP (modules/build-types/server.nix and # DNS port opened on tailscale-router for FreeIPA's conditional forwarder
# modules/build-types/ha-server.nix). # (modules/tailscale/ts-dns-forwarder.nix)
dns = 53;
# NFS stack: portmapper (rpcbind), NFS data, and mountd RPC service.
# Mountd is pinned to a fixed port so the firewall can whitelist it
# without opening rpcbind's full dynamic range. All three need TCP + UDP
# (modules/build-types/ha-server.nix).
nfsRpcbind = 111; nfsRpcbind = 111;
nfsd = 2049; nfsd = 2049;
nfsMountd = 20048; nfsMountd = 20048;
# HA cluster ports opened on ha-server-1 and ha-server-2 # HA cluster ports (modules/ha/cluster-config.nix)
# (modules/build-types/ha-server.nix / modules/ha/cluster-config.nix).
haServerDrbd = 7789; # DRBD replication (TCP) haServerDrbd = 7789; # DRBD replication (TCP)
haServerIscsi = 3260; # iSCSI target (TCP) haServerIscsi = 3260; # iSCSI target (TCP)
haServerCorosync1 = 5404; # Corosync totem ring (UDP) haServerCorosync1 = 5404; # Corosync totem ring (UDP)
@@ -219,49 +254,40 @@
haServerPacemakerRemoted = 3121; # pacemaker-remoted (TCP) haServerPacemakerRemoted = 3121; # pacemaker-remoted (TCP)
haServerPcsd = 2224; # pcsd cluster daemon (TCP) haServerPcsd = 2224; # pcsd cluster daemon (TCP)
# Opened on the docker host's firewall for the Traefik-fronted # Docker host — Traefik HTTP/HTTPS listeners plus one additional exposed
# container stack (docker-compose config lives in the separate # service (modules/build-types/docker.nix)
# /home/debian/docker repo, not here): 80/443 are Traefik's own
# HTTP/HTTPS listeners; 8080 is an additional exposed service whose
# exact backend isn't declared in this repo (modules/build-types/docker.nix).
dockerHttp = 80; dockerHttp = 80;
dockerHttps = 443; dockerHttps = 443;
dockerExtra = 8080; dockerExtra = 8080;
# Beszel monitoring hub, reachable at # Beszel monitoring hub on docker.sweet.home, reached by every agent
# http://<dockerHost>.<homeDomain>:<beszelHub> from every agent # (modules/beszel/enable-agent.nix, hosts/nixos/home.nix)
# (modules/beszel/enable-agent.nix, hosts/nixos/home.nix).
beszelHub = 8090; beszelHub = 8090;
# Proxmox VE and Proxmox Backup Server web UIs, opened as desktop # Proxmox VE and PBS web UIs — desktop shortcuts on the gui build type
# shortcuts on the gui build type (hosts/nixos/home.nix). # (hosts/nixos/home.nix, modules/build-types/gui.nix)
pveWeb = 8006; pveWeb = 8006;
pbsWeb = 8007; pbsWeb = 8007;
# Tor relay's ORPort — the port other Tor relays connect to for onion # Tor relay's ORPort (modules/tor/enable-relay.nix). Opened via
# routing traffic (modules/tor/enable-relay.nix). Tor's own conventional # services.tor.openFirewall rather than allowedTCPPorts directly, but
# default; opened via services.tor.openFirewall rather than # kept here so it's not a bare literal if ever referenced elsewhere.
# networking.firewall.allowedTCPPorts directly, but kept here anyway so
# it's not a bare literal duplicated between the relay's settings and
# anything else that ever needs to reference it.
torRelayOrPort = 9001; torRelayOrPort = 9001;
}; };
# ── Build / image settings ────────────────────────────────────────────────
# .raw disk image size for every proxmox-* host's standalone Disko image # .raw disk image size for every proxmox-* host's standalone Disko image
# build (modules/disko/proxmox.nix, config.system.build.diskoImagesScript # build (modules/disko/proxmox.nix — see docs/proxmox-images.md).
# — see docs/proxmox-images.md). Root fills whatever's left after the ESP
# and swap partitions within this total.
proxmoxImageSize = "50G"; proxmoxImageSize = "50G";
# nix-cache's Nix store garbage collection retention # nix-cache Nix store GC retention (modules/nix-cache/server.nix)
# (modules/nix-cache/server.nix).
nixCacheGcMaxAge = "30d"; nixCacheGcMaxAge = "30d";
# Traefik access log rotation, watched on the docker host at # Traefik access log rotation, watched on the docker host at
# nfsShares.dockerVolumes.mountpoint (modules/traefik/rotate-logs.nix). # nfsShares.dockerVolumes.mountpoint (modules/traefik/rotate-logs.nix)
traefikLogRotate = { traefikLogRotate = {
maxSize = "100M"; # rotate once a log file exceeds this size maxSize = "100M"; # rotate once a log file exceeds this size
keep = 20; # number of rotated logs to retain before deleting the oldest keep = 20; # number of rotated logs to retain
}; };
} }