Compare commits

..
Author SHA1 Message Date
root 14c0ad8125 Merge remote-tracking branch for branch sync 2026-07-20 11:34:21 +00:00
rootandClaude Sonnet 5 a9f20e1229 Add scripts/backup-admin-key.sh to back up the local sops admin key
Companion to rotate-admin-key.sh: copies whatever age identity sops/age
itself would resolve (or an explicit --key-file) to a given destination
path with 0600 perms, validating it's a real identity and round-tripping
the derived public key before/after the write so a corrupted copy is
caught immediately rather than discovered later during a restore.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 11:33:47 +00:00
beatzaplenty bfeea90597 Merge branch 'main' into worktree-rotate-admin-key-script
Check NixOS configurations / eval-hosts (pull_request) Failing after 1h6m12s
2026-07-20 11:27:45 +00:00
rootandClaude Sonnet 5 cafeb8853b Add scripts/rotate-admin-key.sh to automate sops admin key rotation
Check NixOS configurations / eval-hosts (pull_request) Failing after 11m15s
Automates the manual steps sync-host-keys.sh/create-proxmox-resource.sh
print when they bootstrap a fresh, not-yet-trusted age key: verifies a
backed-up key matches the current &admin entry, swaps in a new key, and
re-encrypts every secrets/*.yaml. Explicitly cds into repo_root before any
sops call, since sops resolves .sops.yaml by walking up from cwd rather
than from the target file's path -- confirmed via a scratch-repo test that
running from elsewhere would otherwise silently rotate against the wrong
config.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 11:23:33 +00:00
beatzaplenty 2c2d464503 Merge pull request 'Fix stale documentation: outdated counts, missing build type, spec status' (#9) from fix-stale-docs into main
Check NixOS configurations / eval-hosts (push) Failing after 11m12s
2026-07-20 11:07:43 +00:00
rootandClaude Sonnet 5 5ec7033439 Fix stale documentation: outdated counts, missing build type, spec status
Check NixOS configurations / eval-hosts (pull_request) Failing after 11m26s
Same class of problem as the deployedTargets/README fixes: hand-maintained
prose that drifted from reality and nobody was obligated to update.

- CLAUDE.md: "18 hosts" was a stale hardcoded count (actually 20); reworded
  to not need updating as hosts are added. Also added the missing
  tailscale-exit-node build type to a list that had it everywhere else in
  the file except one bullet.
- AGENTS.md: same missing tailscale-exit-node build type.
- docs/auto-installer.md: the hand-enumerated lxc-* list was missing
  lxc-tailscale-exit-node.
- flake-target-refactor-spec.md: added a "Status: implemented" note so this
  completed historical spec (referenced elsewhere purely for rationale)
  can't be mistaken for an open plan with unresolved Open Questions.
- remove-sensetive-info-refactor.md: the "Definition of done" checklist was
  entirely unchecked despite most of the work being done. Checked off what's
  actually done (sops-nix migration, history scrub just performed, the
  pre-commit gitleaks hook), and left rotation of the GitHub PAT found in
  history explicitly flagged as the one still-open item -- an operator
  action against GitHub, not something this repo can attest to itself.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 11:06:05 +00:00
beatzaplenty 9133afd444 Merge pull request 'Fix duplicate-host check reporting false SSH failures' (#8) from worktree-fix-duplicate-host-check-exitcode into main
Check NixOS configurations / eval-hosts (push) Failing after 11m36s
2026-07-20 11:00:18 +00:00
rootandClaude Sonnet 5 eeec9ce302 Fix duplicate-host check reporting false SSH failures
Check NixOS configurations / eval-hosts (pull_request) Failing after 11m38s
The remote bash script run over SSH ended with a for-loop whose last
statement was `[[ "$n" == "$target" ]] && echo ...`. When the last
VM/CT checked on the node didn't match --host, that test evaluated
false and became the exit status of the whole remote script (1) --
which the wrapper then misreported as "couldn't reach the node",
even though SSH connectivity and the check itself were both fine.
The actual signal is the script's stdout, not its exit code, so end
it with an explicit exit 0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:55:53 +00:00
beatzaplenty a62c4fc023 Merge pull request 'Stop tracking deployment status in the README Hosts table' (#7) from remove-deploy-status-from-readme into main
Check NixOS configurations / eval-hosts (push) Failing after 11m20s
2026-07-20 10:51:59 +00:00
rootandClaude Sonnet 5 97ede62f6d Stop tracking deployment status in the Hosts table
Check NixOS configurations / eval-hosts (pull_request) Failing after 11m24s
Same problem as the deployedTargets removal, just in markdown instead of
Nix: which variant of a buildtype is actually deployed is live
infrastructure state, and a committed table can't stay accurate as that
changes -- it already required a manual edit on every migration and had
drifted before. Keep only what doesn't rot: what each target is for, and
stable naming history. Point at the live node / /etc/flake-target instead
for actual deployment status.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:18:33 +00:00
beatzaplenty ab5206b1c7 Merge pull request 'Replace duplicate-host check with live Proxmox query; drop deployedTargets' (#6) from fix-duplicate-host-self-match into main
Check NixOS configurations / eval-hosts (push) Failing after 12m7s
2026-07-20 10:12:59 +00:00
rootandClaude Sonnet 5 2041557ab3 Replace the duplicate-host check with a live Proxmox query, drop deployedTargets
variables.nix's deployedTargets was a manually-maintained list with no
enforcement keeping it in sync with reality -- it caused two separate
false refusals in a row (naming a VM as deployed well after it had been
destroyed, then matching a target against itself once the list was
"corrected"). Static files can't track whether a resource still actually
exists.

create-proxmox-resource.sh's duplicate-host guard now queries the
Proxmox node directly (qm/pct's own name/hostname config, matched
against --host) instead. Also fixes a gap in that live check: it
originally swallowed ssh failures and would have silently treated "can't
reach the node" the same as "checked, nothing there" -- it now refuses
instead of guessing when the node can't be reached.

deployedTargets is removed entirely from variables.nix since nothing
else in the repo consumed it once this script no longer does; README.md's
Hosts table remains the sole source of truth for "(real, deployed)"
status. CLAUDE.md and the script's own --help/comments updated to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:03:00 +00:00
rootandClaude Sonnet 5 2fd483697b Don't refuse recreating the canonical already-deployed target itself
The duplicate-host check in create-proxmox-resource.sh compared by
hostName only, so it fired even when the target being created was
exactly the one variables.nix's deployedTargets already names (e.g.
rebuilding lxc-nix-cache after destroying its old container to pick up
new sops secrets) -- there's no other machine at risk of an identity
collision in that case, just the normal redeploy workflow. Skip the
check when dt == flake_target; the later VMID-existence check still
guards against clobbering a resource that's actually live on the node.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 09:49:31 +00:00
beatzaplenty 89186b0dee Merge pull request 'Track nix-cache real deployment as lxc-nix-cache, not proxmox-nix-cache' (#5) from worktree-nix-cache-lxc-migration into main 2026-07-20 09:46:41 +00:00
rootandClaude Sonnet 5 8e3606cbd3 Track nix-cache's real deployment as lxc-nix-cache, not proxmox-nix-cache
The old proxmox-nix-cache VM was destroyed and nix-cache is being
redeployed as an LXC container going forward. Without this update,
create-proxmox-resource.sh's duplicate-host check (which only reads this
static list, not live Proxmox state) kept refusing to create
lxc-nix-cache even though nothing named nix-cache actually exists on the
node anymore.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 09:40:03 +00:00
beatzaplenty a18dfb0127 Merge pull request 'Trust nix-cache's SSH host key declaratively on remote-builder clients' (#4) from worktree-magical-cooking-book into main 2026-07-20 07:26:18 +00:00
beatzaplentyandClaude Sonnet 5 75f1342339 Declaratively trust nix-cache's SSH host key on remote-builder clients
Distributed builds failed with "Host key verification failed" on any
client that had never manually SSH'd to nix-cache before, since
nothing populated root's known_hosts for it. Wire nix-cache's host
public key into programs.ssh.knownHosts via a new vars.nixCacheHostKey
so every client picks it up automatically on rebuild.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 17:21:48 +10:00
beatzaplenty 36ba99c9a1 Merge pull request 'Add buildImage shell function for building lxc-* tarballs with host keys' (#3) from worktree-fizzy-juggling-sedgewick into main
Reviewed-on: #3
2026-07-20 07:13:54 +00:00
beatzaplentyandClaude Sonnet 5 0cd8f15b48 Add buildImage shell function for building lxc-* tarballs with host keys
lxc-* hosts need NIXOS_HOST_KEYS_DIR + --impure to bake in a pre-seeded
SSH host key, otherwise sops-nix's .sops.yaml recipient never matches
and every secret permanently fails to decrypt on first boot. That
invocation is easy to forget, so wrap it as `buildImage <flake-target>`
alongside the existing Switch-nix/Test-nix helpers.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 17:12:16 +10:00
beatzaplenty 8613b93fa8 Merge pull request 'Fix nix_extra_opts falsely reporting nix-cache's SSH remote builder down' (#2) from fix-nix-cache-probe-retry into main
Reviewed-on: #2
2026-07-20 07:02:14 +00:00
beatzaplentyandClaude Sonnet 5 20f9475a7d Fix nix_extra_opts falsely reporting nix-cache's SSH remote builder down
The reachability check used `cat < /dev/tcp/${NIX_CACHE_HOST}/22`, which
blocks forever reading for EOF that never comes -- sshd sends its banner
and then holds the connection open waiting for the client to speak next.
Every single check hit the 3s timeout and reported "unreachable"
unconditionally, regardless of whether the remote builder was actually up.
Confirmed live: a plain TCP connect (`exec 3<>/dev/tcp/...`, no read)
returns in ~60ms against a healthy nix-cache instead of always timing out.

Fixing that exposed a second, previously-dormant bug: `printf -v
NIX_EXTRA_OPTS '%q ' "${NIX_OPTS[@]}"` on a genuinely empty NIX_OPTS array
still runs one format pass and yields the literal `'' ` rather than an
empty string. A subprocess (e.g. sync-host-keys.sh) reusing this
process's decision via `eval "NIX_OPTS=(${NIX_EXTRA_OPTS})"` then rebuilt
a 1-element array holding an empty string instead of a 0-element array,
which broke `nix-shell "${NIX_OPTS[@]}" -p <pkg>` with a bogus positional
argument the moment NIX_OPTS was legitimately empty (nix-cache reachable)
-- something the first bug had made impossible to ever hit before.

Also adds a couple of retries (1s apart) to both checks as a secondary
safety net against genuine multi-second blips, on top of fixing the
checks themselves.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 16:57:43 +10:00
beatzaplenty 6babb3eec5 Merge pull request 'Fix create-proxmox-resource.sh --dry-run hiding nix-cache probe results' (#1) from worktree-starry-painting-whistle into main
Reviewed-on: #1
2026-07-20 06:50:40 +00:00
beatzaplentyandClaude Sonnet 5 33730e6ccf Fix create-proxmox-resource.sh --dry-run hiding nix-cache probe results
The tarball/disko-image build previews were hardcoded strings that never
included ${NIX_OPTS[@]}, so --dry-run always showed the same "would build"
command whether nix-cache's substituter/remote-builder got disabled by
nix_extra_opts's reachability probe or not -- the actual (non-dry-run)
build commands already applied it correctly, only the preview lied.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 16:36:14 +10:00
beatzaplentyandClaude Sonnet 5 65f89806cb Fix beszel-agent losing its hub-pairing fingerprint on every restart
services.beszel.agent runs under DynamicUser=true with ProtectSystem =
"strict" and no StateDirectory, so /var/lib/beszel-agent -- where the
agent persists the fingerprint that locks its hub pairing to this
machine (github.com/henrygd/beszel/discussions/1542) -- was never
actually writable. Every restart silently failed to persist it and
regenerated a fresh one in memory, permanently desyncing from whatever
the hub had on record after the very first successful pairing. Affects
every host importing modules/beszel/enable-agent.nix (nix-cache, server),
not just full container rebuilds.

Found via nix-cache showing "fingerprint mismatch" after being rebuilt
post-outage; confirmed server was silently exposed to the same bug, just
hadn't restarted since its first pairing. Fixed by declaring
StateDirectory so systemd gives the dynamic user real persistent storage.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 15:59:24 +10:00
beatzaplentyandClaude Sonnet 5 c3007097a6 Fix create-proxmox-resource.sh defaulting hostname to the flake target
--name (used as pct/qm create's --hostname/--name) defaulted to
$flake_target (e.g. "lxc-nix-cache"), not $host (e.g. "nix-cache"). Since
proxmoxLXC.manageHostName pulls the guest's real networking.hostName
straight from Proxmox's own container config, this silently overrode
host.nix's hostName with a build-type-specific name. Default --name to
--host instead, so the guest's identity matches host.nix regardless of
which platform variant built it.

Found by spinning up a fresh lxc-nix-cache test container and noticing its
hostname was "lxc-nix-cache" instead of "nix-cache".

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 15:19:35 +10:00
beatzaplentyandClaude Sonnet 5 9724babcea Add tailscale-exit-node build type across all three platforms
New build type dedicated to Tailscale exit-node capability, wired up for
linode/proxmox/lxc like every other build type (the lxc variant is the one
actually intended for deployment). Kept separate from the "server" host
rather than bundling exit-node capability onto it.

Trimmed modules/tailscale/exit-node.nix down to pure exit-node behavior:
dropped the old --advertise-routes=${vars.lanCidr} bundling (meaningless
for a Linode-hosted VPS with no path to the LAN), and switched
extraUpFlags -> extraSetFlags. Confirmed against nixpkgs' tailscale.nix
that extraUpFlags is only applied by tailscaled-autoconnect, which itself
only runs when services.tailscale.authKeyFile is set -- nothing in this
repo sets one, so the old flags would never have actually been applied.
extraSetFlags runs unconditionally via tailscaled-set on every boot, so
--advertise-exit-node self-reapplies once the operator has done the
one-time manual `tailscale up` auth.

Verified: all three new targets eval cleanly, nixpkgs-fmt/statix clean,
and a dry-run build of lxc-tailscale-exit-node's tarball resolves its full
closure including tailscaled-set.service.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01La55Nsss8jZ7ZuzUV9mfot
2026-07-20 13:38:39 +10:00
beatzaplentyandClaude Sonnet 5 7055bcdb97 Fix lxc-* hosts never completing first-boot user/secrets activation
virtualisation/proxmox-lxc.nix registers the Nix store DB via a systemd
service, never an activation script -- so neededForUsers sops secrets
(password hashes) and the user-creation step that consumes them never ran
on a real first boot, leaving /etc/shadow stuck with build-time placeholder
entries. boot.postBootCommands looked like the right hook (stage-2-init.sh
does invoke it) but switch-to-configuration behaves unreliably that early,
before systemd itself is up. Fixed with a genuine oneshot systemd service,
gated by ConditionPathExists so it only ever runs once.

Confirmed live via a from-scratch destroy+rebuild+redeploy of the
lxc-nix-cache test container: real password hashes applied automatically,
systemctl is-system-running -> running, zero failed units.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01La55Nsss8jZ7ZuzUV9mfot
2026-07-20 13:38:21 +10:00
beatzaplentyandClaude Sonnet 5 d973da487c Fix lxc-* hosts having no host-key pre-seeding mechanism at all
The real root cause behind the original nix-cache 502, traced all the way
through: modules/installer/host-keys.nix (which NIXOS_HOST_KEYS_DIR=...
--impure actually wires up) is only ever imported by the installer's own
modules/installer/common.nix -- modules/platforms/lxc.nix, which every
real lxc-* host build actually uses, never imported anything like it.
docs/auto-installer.md previously claimed NIXOS_HOST_KEYS_DIR bakes a key
into lxc-* tarballs "the same way it does for the ISO/PXE installer
images" -- that was never actually true; I wrote it without verifying the
mechanism existed for lxc.nix specifically.

In practice this meant every lxc-* container booted with a freshly
self-generated SSH host key that could never match whatever .sops.yaml
actually trusts for that target, so *every* secret -- not just
cache-priv-key -- silently failed to decrypt. No error surfaces in the
boot log for this: the activation step that installs secrets only runs
on a genuinely fresh first activation and silently no-ops once
/run/current-system already exists, so by the time anyone looks the
window has closed. Found by manually invoking sops-install-secrets
directly: "Error getting data key: 0 successful groups required, got 0".

Fixed by giving modules/platforms/lxc.nix the same key-baking mechanism
the installer has, but keyed to its own exact flake target and placing
the key directly at /etc/ssh/ssh_host_ed25519_key (no copy step to stage
for, unlike the installer's /etc/host-keys/ staging area -- an lxc-*
tarball has no install step). The target name comes in via
specialArgs.flakeTarget (new, set by flake.nix's mkTarget) rather than
being read back from config.environment.etc."flake-target" -- reading
that back from within a module that also contributes to
environment.etc is circular (confirmed: "infinite recursion
encountered").

Verified live end-to-end against the real test container (lxc-nix-cache,
VMID 100 on pve.sweet.home): destroyed it, rebuilt the tarball fresh with
the fix, recreated it, and confirmed /run/secrets/ now has all three
secrets this host needs (beszel-token, cache-priv-key, nix-github-token),
nix-serve is active (running), and curl http://localhost/nix-cache-info
succeeds both directly and through nginx.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01La55Nsss8jZ7ZuzUV9mfot
2026-07-20 12:29:01 +10:00
20 changed files with 819 additions and 108 deletions
+1 -1
View File
@@ -7,7 +7,7 @@ servers and workstation.
The flake exposes NixOS configurations named `<platform>-<buildtype>` The flake exposes NixOS configurations named `<platform>-<buildtype>`
(platforms: `linode`, `proxmox`, `lxc`; build types: `minimal`, `nix-cache`, (platforms: `linode`, `proxmox`, `lxc`; build types: `minimal`, `nix-cache`,
`server`, `docker`, `gui`, `pxe-boot`), generated from `modules/platforms/*` `server`, `docker`, `gui`, `pxe-boot`, `tailscale-exit-node`), generated from `modules/platforms/*`
and `modules/build-types/*` by the `mkTarget` function in `flake.nix`. Not and `modules/build-types/*` by the `mkTarget` function in `flake.nix`. Not
every combination is built — `pxe-boot` has no `linode` variant. See every combination is built — `pxe-boot` has no `linode` variant. See
`README.md` for the full current target list; treat `flake.nix` as the `README.md` for the full current target list; treat `flake.nix` as the
+43 -15
View File
@@ -62,8 +62,9 @@ There is no test suite — "correctness" here means the flake evaluates and
sweeps: after editing one or two hosts/modules, evaluate just the sweeps: after editing one or two hosts/modules, evaluate just the
`nixosConfigurations.<host>` you touched (plus any `config.system.build.tarball` `nixosConfigurations.<host>` you touched (plus any `config.system.build.tarball`
/`diskoImagesScript`/package output affected) rather than looping over every /`diskoImagesScript`/package output affected) rather than looping over every
host — `codex-maintenance.sh` evaluates 18 hosts plus every package/tarball/ host — `codex-maintenance.sh` evaluates every `nixosConfigurations` host plus
image variant now and is slow to run after each small change. Reserve a full every package/tarball/image variant and is slow to run after each small
change. Reserve a full
`codex-maintenance.sh` run for changes that plausibly affect every host `codex-maintenance.sh` run for changes that plausibly affect every host
(`modules/common/*`, `flake.nix`, `variables.nix`) or as a final check before (`modules/common/*`, `flake.nix`, `variables.nix`) or as a final check before
committing. This is a session-workflow preference only — it does not apply to committing. This is a session-workflow preference only — it does not apply to
@@ -93,9 +94,10 @@ Beyond `codex-setup.sh`/`codex-maintenance.sh` above, `scripts/` also has:
(`--force-rebuild` to skip that and always rebuild), and probes (`--force-rebuild` to skip that and always rebuild), and probes
nix-cache's substituter/remote-builder reachability once up front rather nix-cache's substituter/remote-builder reachability once up front rather
than letting every `nix build` call retry against it individually. than letting every `nix build` call retry against it individually.
Refuses to create a target whose host identity already has a real Refuses to create a target whose host identity already exists live on
deployment elsewhere (`variables.nix`'s `deployedTargets`) unless the node (checked directly via `qm`/`pct`, not any file in this repo)
`--allow-duplicate-host` is passed. `--dry-run` throughout both modes. unless `--allow-duplicate-host` is passed. `--dry-run` throughout both
modes.
- `scripts/env.sh` — shared config (`PROXMOX_HOST`, storage pool, bridge, - `scripts/env.sh` — shared config (`PROXMOX_HOST`, storage pool, bridge,
default cores/memory) sourced by `create-proxmox-resource.sh`. Add new default cores/memory) sourced by `create-proxmox-resource.sh`. Add new
cross-script config here instead of duplicating it per-script. cross-script config here instead of duplicating it per-script.
@@ -104,13 +106,38 @@ Beyond `codex-setup.sh`/`codex-maintenance.sh` above, `scripts/` also has:
reference `variables.nix` (confirmed empirically — `nix flake metadata` reference `variables.nix` (confirmed empirically — `nix flake metadata`
errors on it), so this is the closest equivalent to a single source of errors on it), so this is the closest equivalent to a single source of
truth for the tracked release. truth for the tracked release.
- `scripts/rotate-admin-key.sh <backup-admin-key> [--new-key-file <path>]
[--dry-run]` — rotates `.sops.yaml`'s `&admin` age key: decrypts with a
backed-up copy of the key currently trusted as `&admin` (verified by
deriving its public key and comparing, not taken on faith), replaces the
`&admin` line with a new key already present in the environment
(defaults to wherever sops/age itself would look), and runs
`sops updatekeys` on every `secrets/*.yaml`. One-way: the old key can no
longer decrypt anything re-encrypted this way. This is the automation
for the manual steps `sync-host-keys.sh`/`create-proxmox-resource.sh`
print when they bootstrap a brand-new, not-yet-trusted key on a machine
with no prior admin access.
- `scripts/backup-admin-key.sh <dest-path> [--key-file <path>] [--force]
[--dry-run]` — copies the local sops age key (source resolution matches
sops/age itself: `$SOPS_AGE_KEY` inline, then `--key-file`, then
`$SOPS_AGE_KEY_FILE`, then the XDG default) to an arbitrary destination
path with `0600` permissions, validating it's a real age identity and
round-tripping the public key before and after the write. Refuses to
overwrite an existing `<dest-path>` without `--force`. Purely a local
filesystem copy — never touches `.sops.yaml`/`secrets/*.yaml` or the
repo at all. The resulting file is exactly what `rotate-admin-key.sh`
expects as its backup-key argument.
`sync-host-keys.sh` and `create-proxmox-resource.sh` genuinely mutate real `sync-host-keys.sh`, `create-proxmox-resource.sh`, and
state when run for real (not `--dry-run`): real `secrets/*.yaml` `rotate-admin-key.sh` genuinely mutate real state when run for real (not
recipients, real Proxmox VMs/containers. They require the operator's own `--dry-run`): real `secrets/*.yaml` recipients, real Proxmox VMs/
SSH/sops access, which an agent session doesn't have — but don't suggest containers, real revocation of decrypt access. They require the
running either non-dry-run without the operator's explicit go-ahead even operator's own SSH/sops access, which an agent session doesn't have — but
if it becomes technically reachable. don't suggest running any of them non-dry-run without the operator's
explicit go-ahead even if it becomes technically reachable.
`backup-admin-key.sh` only writes a key copy to a path the operator gives
it — lower-stakes than the others, but it still handles a real private
key, so treat its destination path choice as the operator's call too.
## Architecture ## Architecture
@@ -133,9 +160,10 @@ nixosSystem {
``` ```
Platforms: `linode`, `proxmox`, `lxc`. Build types: `minimal`, `nix-cache`, Platforms: `linode`, `proxmox`, `lxc`. Build types: `minimal`, `nix-cache`,
`server`, `docker`, `gui`, `pxe-boot`. Not every combination is built — e.g. `server`, `docker`, `gui`, `pxe-boot`, `tailscale-exit-node`. Not every
`pxe-boot` has no `linode` variant (PXE/DHCP/TFTP need LAN L2 adjacency a combination is built — e.g. `pxe-boot` has no `linode` variant (PXE/DHCP/TFTP
Linode VPS doesn't have). Treat `flake.nix`'s `generatedTargets` as the source need LAN L2 adjacency a Linode VPS doesn't have). Treat `flake.nix`'s
`generatedTargets` as the source
of truth for which hosts exist — `README.md`, `AGENTS.md`, of truth for which hosts exist — `README.md`, `AGENTS.md`,
`docs/flake-lock-automation.md`, and the CI eval workflows `docs/flake-lock-automation.md`, and the CI eval workflows
(`.github/workflows/check-nixos.yml`, `.gitea/workflows/check-nixos.yml`) list (`.github/workflows/check-nixos.yml`, `.gitea/workflows/check-nixos.yml`) list
@@ -162,7 +190,7 @@ removing a host.
`vzdump` backup-archive metadata this doesn't have), no install step — `vzdump` backup-archive metadata this doesn't have), no install step —
see `docs/auto-installer.md`. see `docs/auto-installer.md`.
- `modules/build-types/*.nix` — what a system is for: - `modules/build-types/*.nix` — what a system is for:
minimal/server/docker/gui/pxe-boot/nix-cache. minimal/server/docker/gui/pxe-boot/nix-cache/tailscale-exit-node.
- `modules/common/configuration.nix` — base NixOS config imported by every - `modules/common/configuration.nix` — base NixOS config imported by every
host: locale, users, nix settings, git. host: locale, users, nix settings, git.
- `modules/common/home.nix` / `hosts/nixos/home.nix` — Home Manager config for - `modules/common/home.nix` / `hosts/nixos/home.nix` — Home Manager config for
+16 -12
View File
@@ -10,7 +10,7 @@ pieces composed in `flake.nix`:
- **Platforms** (what it runs on): `linode`, `proxmox`, `lxc` - **Platforms** (what it runs on): `linode`, `proxmox`, `lxc`
- **Build types** (what it's for): `minimal`, `nix-cache`, `server`, `docker`, - **Build types** (what it's for): `minimal`, `nix-cache`, `server`, `docker`,
`gui`, `pxe-boot` `gui`, `pxe-boot`, `tailscale-exit-node`
Not every combination exists — `pxe-boot` has no `linode` variant, since Not every combination exists — `pxe-boot` has no `linode` variant, since
PXE/DHCP/TFTP need LAN L2 adjacency that a Linode VPS doesn't have. The full PXE/DHCP/TFTP need LAN L2 adjacency that a Linode VPS doesn't have. The full
@@ -18,19 +18,23 @@ list:
| Target | Purpose | | Target | Purpose |
| --- | --- | | --- | --- |
| `linode-minimal` | Minimal NixOS host profile on a Linode VPS (real, deployed) | | `linode-minimal` | Minimal NixOS host profile on a Linode VPS |
| `proxmox-minimal` | Minimal NixOS host profile on Proxmox (real, deployed — previously the flat `nix-minimal` target) | | `proxmox-minimal` | Minimal NixOS host profile on Proxmox — previously the flat `nix-minimal` target |
| `lxc-minimal` | Minimal NixOS host profile in a Proxmox LXC container | | `lxc-minimal` | Minimal NixOS host profile in a Proxmox LXC container |
| `linode-nix-cache` / `proxmox-nix-cache` / `lxc-nix-cache` | Local Nix binary cache and remote builder (`proxmox-nix-cache` is the real, deployed one — previously the flat `nix-cache` target) | | `linode-nix-cache` / `proxmox-nix-cache` / `lxc-nix-cache` | Local Nix binary cache and remote builder — previously the flat `nix-cache` target |
| `linode-server` / `proxmox-server` / `lxc-server` | Storage, NFS, backup, and monitoring exporter host (`proxmox-server` is the real, deployed one — previously the flat `server` target) | | `linode-server` / `proxmox-server` / `lxc-server` | Storage, NFS, backup, and monitoring exporter host — previously the flat `server` target |
| `linode-docker` / `proxmox-docker` / `lxc-docker` | Docker host for the main container stack (`proxmox-docker` is the real, deployed one — previously the flat `docker` target) | | `linode-docker` / `proxmox-docker` / `lxc-docker` | Docker host for the main container stack — previously the flat `docker` target |
| `linode-gui` / `proxmox-gui` / `lxc-gui` | Cinnamon desktop workstation (`proxmox-gui` is the real, deployed one — previously the flat `nixos` target) | | `linode-gui` / `proxmox-gui` / `lxc-gui` | Cinnamon desktop workstation — previously the flat `nixos` target |
| `proxmox-pxe-boot` / `lxc-pxe-boot` | HTTP/iPXE boot asset host (`proxmox-pxe-boot` is the real, deployed one — previously the flat `pxe-boot` target) | | `proxmox-pxe-boot` / `lxc-pxe-boot` | HTTP/iPXE boot asset host — previously the flat `pxe-boot` target |
| `linode-tailscale-exit-node` / `proxmox-tailscale-exit-node` / `lxc-tailscale-exit-node` | Tailscale exit node |
The "(real, deployed)" targets above are also tracked machine-readably in Which variant of a given buildtype is actually deployed isn't tracked
`variables.nix`'s `deployedTargets` — keep both in sync when a deployment anywhere in this repo — that's live infrastructure state, not something a
changes. `scripts/create-proxmox-resource.sh` reads that list to refuse committed file can keep accurate, and it changes independently of the code.
creating a same-identity duplicate of an already-deployed host by accident. Check the Proxmox node itself, or `/etc/flake-target` on a running host (see
below), if you need to know what's really out there right now.
`scripts/create-proxmox-resource.sh`'s duplicate-host guard works the same
way: it checks the Proxmox node directly rather than any file here.
Each buildtype's `hosts/<name>/host.nix` carries the per-machine identity Each buildtype's `hosts/<name>/host.nix` carries the per-machine identity
(hostname, hostId, per-machine secrets, `system.stateVersion`) that must stay (hostname, hostId, per-machine secrets, `system.stateVersion`) that must stay
+28 -6
View File
@@ -19,7 +19,7 @@ see "LXC hosts" immediately below for why those are different.**
## LXC hosts ## LXC hosts
`lxc-*` targets (`lxc-minimal`, `lxc-nix-cache`, `lxc-server`, `lxc-docker`, `lxc-*` targets (`lxc-minimal`, `lxc-nix-cache`, `lxc-server`, `lxc-docker`,
`lxc-gui`, `lxc-pxe-boot`) are **not** installed via `auto-install.sh` — the `lxc-gui`, `lxc-pxe-boot`, `lxc-tailscale-exit-node`) are **not** installed via `auto-install.sh` — the
interactive menu deliberately excludes them. Don't try to select one there; interactive menu deliberately excludes them. Don't try to select one there;
`nixos-install` would bind-mount `/` onto `/mnt` (LXC containers have no raw `nixos-install` would bind-mount `/` onto `/mnt` (LXC containers have no raw
disk to partition) and then refuse to touch the filesystem it's currently disk to partition) and then refuse to touch the filesystem it's currently
@@ -77,11 +77,33 @@ system profile) — there's no separate activation step to run yourself.
of this (build, host-key handling, upload, `pct create` with the flags of this (build, host-key handling, upload, `pct create` with the flags
above) — see its `--help`. above) — see its `--help`.
Host keys still need pre-seeding the same way as any other host (see "Host Host keys still need pre-seeding the same way as any other host — the
keys" below) — the sops-nix activation-vs-first-boot race is identical sops-nix activation-vs-first-boot race is identical regardless of how the
regardless of how the image reaches the machine. `NIXOS_HOST_KEYS_DIR=... image reaches the machine. Unlike the ISO/PXE installer (where
nix build ... --impure` bakes the matching key into the tarball the same way `modules/installer/host-keys.nix` bakes *every* `host-keys/` entry into
it does for the ISO/PXE installer images. `/etc/host-keys/` for `auto-install.sh` to pick from and copy at install
time — see "Host keys" below), an `lxc-*` tarball has no install step to
copy anything during, so `modules/platforms/lxc.nix` bakes this *one*
target's key straight into `/etc/ssh/ssh_host_ed25519_key(.pub)` directly,
keyed by its own exact flake target name (`config.environment.etc` can't
be read back from within a module still contributing to it, so this comes
in via `specialArgs.flakeTarget`, set by `flake.nix`'s `mkTarget`):
```sh
NIXOS_HOST_KEYS_DIR="$(pwd)/host-keys" \
nix build .#nixosConfigurations.lxc-nix-cache.config.system.build.tarball --impure
```
Confirmed the hard way: without this, the tarball's own built-in system
just generates a fresh host key at first boot like any host would, which
can never match whatever `.sops.yaml` actually trusts for that target —
`sops-install-secrets` fails with `Error getting data key: 0 successful
groups required, got 0`, and *every* secret (including this host's own
login) permanently fails to decrypt, silently — no error in the boot log
at all, since the activation step that would install secrets only runs on
a from-scratch first activation and skips silently once `/run/current-system`
already exists. `scripts/create-proxmox-resource.sh` always builds with
`NIXOS_HOST_KEYS_DIR` set for this reason.
## Layout ## Layout
+9
View File
@@ -59,6 +59,15 @@ On `nix-cache`, install the matching public key used by `nixremote` authorized k
The committed `nixremote` authorized keys are public SSH keys only. Keep the The committed `nixremote` authorized keys are public SSH keys only. Keep the
matching private keys on client hosts and out of the repository. matching private keys on client hosts and out of the repository.
nix-cache's own SSH *host* key is trusted declaratively via
`programs.ssh.knownHosts` in `modules/nix-cache/remote-builder-client.nix`,
sourced from `vars.nixCacheHostKey` (`variables.nix`) — every client rebuild
picks it up automatically, so distributed builds don't fail with "Host key
verification failed" on a client that has never manually SSH'd to nix-cache
before. If nix-cache's host key is ever rotated or the host rebuilt from
scratch, update `vars.nixCacheHostKey` to match its new
`/etc/ssh/ssh_host_ed25519_key.pub`.
## Manual verification ## Manual verification
After deployment: After deployment:
+8
View File
@@ -1,5 +1,13 @@
# Spec: Refactor Flake Targets into Platform × Build-Type Matrix # Spec: Refactor Flake Targets into Platform × Build-Type Matrix
**Status: implemented.** `flake.nix`'s `generatedTargets`/`mkTarget` and
`modules/platforms/*`/`modules/build-types/*` are the result of this spec —
kept here for historical rationale only (referenced from `CLAUDE.md`'s
"Composition pattern" section), not as an active or open plan. The "Open
Questions" below were resolved during implementation; don't treat them as
outstanding. A `tailscale-exit-node` build type was added later, beyond this
spec's original scope.
## Context ## Context
The flake at `~/nixos` currently defines these output targets (flat, ad-hoc naming): The flake at `~/nixos` currently defines these output targets (flat, ad-hoc naming):
+15 -2
View File
@@ -33,6 +33,9 @@
# nix-cache itself consumes the nix-cache substituter and remote # nix-cache itself consumes the nix-cache substituter and remote
# builder. # builder.
mkTarget = { platform, buildType, hostPath, homeFile ? ./modules/common/home.nix }: mkTarget = { platform, buildType, hostPath, homeFile ? ./modules/common/home.nix }:
let
flakeTarget = "${platform}-${buildType}";
in
nixpkgs.lib.nixosSystem { nixpkgs.lib.nixosSystem {
inherit system; inherit system;
modules = [ modules = [
@@ -42,7 +45,7 @@
./modules/platforms/${platform}.nix ./modules/platforms/${platform}.nix
./modules/build-types/${buildType}.nix ./modules/build-types/${buildType}.nix
hostPath hostPath
{ environment.etc."flake-target".text = "${platform}-${buildType}"; } { environment.etc."flake-target".text = flakeTarget; }
home-manager.nixosModules.home-manager home-manager.nixosModules.home-manager
{ {
home-manager = { home-manager = {
@@ -56,7 +59,13 @@
./modules/nix-cache/client.nix ./modules/nix-cache/client.nix
./modules/nix-cache/remote-builder-client.nix ./modules/nix-cache/remote-builder-client.nix
]; ];
specialArgs = { inherit inputs vars netbootSystem; }; # flakeTarget is passed via specialArgs (not read back from
# config.environment.etc."flake-target" above) specifically so
# modules/platforms/lxc.nix can use it to select its own host key
# file without a same-option circular dependency (a module
# contributing to environment.etc can't read the merged
# environment.etc it's itself contributing to).
specialArgs = { inherit inputs vars netbootSystem flakeTarget; };
}; };
# Generated platform x build-type matrix. pxe-boot has no linode # Generated platform x build-type matrix. pxe-boot has no linode
@@ -85,6 +94,10 @@
proxmox-pxe-boot = mkTarget { platform = "proxmox"; buildType = "pxe-boot"; hostPath = ./hosts/pxe-boot/host.nix; }; proxmox-pxe-boot = mkTarget { platform = "proxmox"; buildType = "pxe-boot"; hostPath = ./hosts/pxe-boot/host.nix; };
lxc-pxe-boot = mkTarget { platform = "lxc"; buildType = "pxe-boot"; hostPath = ./hosts/pxe-boot/host.nix; }; lxc-pxe-boot = mkTarget { platform = "lxc"; buildType = "pxe-boot"; hostPath = ./hosts/pxe-boot/host.nix; };
linode-tailscale-exit-node = mkTarget { platform = "linode"; buildType = "tailscale-exit-node"; hostPath = ./hosts/tailscale-exit-node/host.nix; };
proxmox-tailscale-exit-node = mkTarget { platform = "proxmox"; buildType = "tailscale-exit-node"; hostPath = ./hosts/tailscale-exit-node/host.nix; };
lxc-tailscale-exit-node = mkTarget { platform = "lxc"; buildType = "tailscale-exit-node"; hostPath = ./hosts/tailscale-exit-node/host.nix; };
}; };
# Auto-install environments (migrated from the former nix-auto-installer # Auto-install environments (migrated from the former nix-auto-installer
+12
View File
@@ -0,0 +1,12 @@
_:
{
networking.hostName = "exit-node";
# No networking.hostId: only ZFS-touching hosts (server, docker) need one
# for pool-import safety, and this host does neither.
# A genuinely new host (not a pre-refactor carry-over), so it tracks the
# flake's current nixpkgs release rather than being pinned to an older one.
system.stateVersion = "26.05";
}
+9
View File
@@ -6,4 +6,13 @@
#DOCKER_HOST = "tcp://docker-socket-proxy:2375"; #DOCKER_HOST = "tcp://docker-socket-proxy:2375";
HUB_URL = "http://${vars.dockerHost}.${vars.homeDomain}:${toString vars.ports.beszelHub}"; HUB_URL = "http://${vars.dockerHost}.${vars.homeDomain}:${toString vars.ports.beszelHub}";
}; };
# The upstream module runs beszel-agent under DynamicUser with
# ProtectSystem = "strict" and no StateDirectory, so /var/lib/beszel-agent
# (where the agent persists its hub-pairing fingerprint, per
# https://github.com/henrygd/beszel/discussions/1542) isn't writable --
# every restart silently fails to save it and regenerates a fresh one in
# memory, permanently desyncing from whatever the hub has on record after
# the very first successful pairing. Give it real persistent storage.
systemd.services.beszel-agent.serviceConfig.StateDirectory = "beszel-agent";
} }
@@ -0,0 +1,21 @@
{ ... }:
{
imports = [
../tailscale/exit-node.nix
];
# "server", not "both": this build type only ever advertises itself as an
# exit node (see ../tailscale/exit-node.nix) -- it doesn't advertise LAN
# subnet routes, so it doesn't need the "client"-side loose reverse-path
# filtering that "both" would also turn on. Deliberately left unbundled
# from LAN-subnet-route advertisement so this build type stays valid on
# every platform, including linode (a remote VPS with no network path to
# the home LAN at all).
services.tailscale.useRoutingFeatures = "server";
# Forwarded exit-node traffic arrives on tailscale0 already
# tailscale-authenticated -- the firewall's normal per-port allow-list
# would otherwise drop it. Standard NixOS/Tailscale exit-node guidance.
networking.firewall.trustedInterfaces = [ "tailscale0" ];
}
+21
View File
@@ -17,6 +17,26 @@ let
--refresh \ --refresh \
--flake git+https://${vars.lanDomain}/beatzaplenty/nixos.git#$(cat /etc/flake-target) --flake git+https://${vars.lanDomain}/beatzaplenty/nixos.git#$(cat /etc/flake-target)
''; '';
# lxc-* hosts pre-seed their SSH host key at build time (see
# modules/platforms/lxc.nix) so sops-nix's .sops.yaml recipient matches on
# first boot -- without it, secrets permanently fail to decrypt (see that
# file's comment for the confirmed failure). That requires --impure plus
# NIXOS_HOST_KEYS_DIR pointing at the repo's host-keys/ dir, same pattern
# docs/auto-installer.md uses for the installer ISO. A function, not a
# shellAlias, since the target name has to interpolate into the middle of
# the flake attribute path, not just append after it. Must be run from the
# repo root, same as every other host-keys/ command in this repo.
buildImageFn = ''
buildImage() {
if [ -z "$1" ]; then
echo "usage: buildImage <flake-target> (e.g. lxc-docker)" >&2
return 1
fi
NIXOS_HOST_KEYS_DIR="$(pwd)/host-keys" nix build --impure \
".#nixosConfigurations.$1.config.system.build.tarball"
}
'';
in in
{ {
programs.bash = { programs.bash = {
@@ -25,5 +45,6 @@ in
"Switch-nix" = mySwitchCmd; "Switch-nix" = mySwitchCmd;
"Test-nix" = myTestCmd; "Test-nix" = myTestCmd;
}; };
initExtra = buildImageFn;
}; };
} }
@@ -5,6 +5,14 @@
# sudo install -d -m 0700 /root/.ssh # sudo install -d -m 0700 /root/.ssh
# sudo install -m 0600 ./nixremote /root/.ssh/nixremote # sudo install -m 0600 ./nixremote /root/.ssh/nixremote
# sudo ssh -i /root/.ssh/nixremote nixremote@nix-cache nix-store --version # sudo ssh -i /root/.ssh/nixremote nixremote@nix-cache nix-store --version
# Trust nix-cache's SSH host key declaratively so the nix-daemon (root)
# can connect the first time without a manual ssh-keyscan/known_hosts
# step on every new client.
programs.ssh.knownHosts.${vars.nixCacheHost} = {
hostNames = [ vars.nixCacheHost ];
publicKey = vars.nixCacheHostKey;
};
nix = { nix = {
distributedBuilds = true; distributedBuilds = true;
+111 -1
View File
@@ -1,5 +1,44 @@
{ lib, modulesPath, ... }: { lib, modulesPath, flakeTarget, ... }:
let
# Bakes this exact flake target's pre-generated SSH host key straight
# into /etc/ssh/ -- mirrors modules/installer/host-keys.nix's
# builtins.getEnv pattern (impure and empty under normal `nix
# build`/`nix eval`, so this is a no-op unless explicitly opted into
# with NIXOS_HOST_KEYS_DIR=... --impure), but places the key directly
# rather than staging it under /etc/host-keys/ for a later manual copy
# -- this is the whole system for a `lxc-*` host, built straight to a
# pct-restorable tarball with no install step, so there's no later copy
# step to stage for.
#
# Without this, config.system.build.tarball's built-in system just
# generates a fresh host key at first boot like any other host would --
# but sops-nix derives its decryption key from *this* file, and
# .sops.yaml only trusts whatever key scripts/sync-host-keys.sh already
# registered for this exact target name. A freshly-generated key can
# never match that, so every secret (including this host's own login)
# permanently fails to decrypt. Confirmed live: sops-install-secrets
# errored with "Error getting data key: 0 successful groups required,
# got 0" -- the container's actual host key's age fingerprint didn't
# match the one registered in .sops.yaml at all.
hostKeysDirStr = builtins.getEnv "NIXOS_HOST_KEYS_DIR";
hasHostKeysDir = hostKeysDirStr != "" && builtins.pathExists hostKeysDirStr;
hostKeysDir = /. + hostKeysDirStr;
# flakeTarget ("${platform}-${buildType}") comes in via specialArgs from
# flake.nix's mkTarget -- exactly the name scripts/sync-host-keys.sh
# registers keys under. Deliberately not read back from
# config.environment.etc."flake-target" (which is set to the same value)
# -- this module also *contributes* to environment.etc below, and a
# module reading the merged value of an option it's still defining is a
# circular dependency (confirmed: "infinite recursion encountered").
privKeyFile = hostKeysDir + "/${flakeTarget}_ssh_host_ed25519_key";
pubKeyFile = hostKeysDir + "/${flakeTarget}_ssh_host_ed25519_key.pub";
hasKeyForThisTarget =
hasHostKeysDir
&& builtins.pathExists privKeyFile
&& builtins.pathExists pubKeyFile;
in
{ {
# LXC containers share the host kernel — Proxmox starts them by exec'ing # LXC containers share the host kernel — Proxmox starts them by exec'ing
# /sbin/init directly, no bootloader/initrd involved — and Proxmox has its # /sbin/init directly, no bootloader/initrd involved — and Proxmox has its
@@ -35,4 +74,75 @@
# for the same reason; it just doesn't disable NetworkManager itself, # for the same reason; it just doesn't disable NetworkManager itself,
# which modules/common/configuration.nix enables for every host. # which modules/common/configuration.nix enables for every host.
networking.networkmanager.enable = lib.mkForce false; networking.networkmanager.enable = lib.mkForce false;
environment.etc = lib.mkIf hasKeyForThisTarget {
"ssh/ssh_host_ed25519_key" = {
source = privKeyFile;
mode = "0600";
};
"ssh/ssh_host_ed25519_key.pub" = {
source = pubKeyFile;
mode = "0644";
};
};
# virtualisation/proxmox-lxc.nix (imported above) registers the Nix
# store DB via a systemd service (register-nix-paths) -- it never runs
# an activation script at all. Confirmed live this means neither
# sops-nix's "for users" secrets (password hashes -- installed by the
# activation script itself, not a systemd service, since they need to
# exist *before* user creation) nor the user-creation step that
# consumes them ever run on a real lxc-* boot. Regular secrets
# (nix-serve's key, beszel's token, etc.) work anyway because sops-nix
# provides its own systemd service for those.
#
# A systemd service, not boot.postBootCommands: tried that first (it's
# a genuine, generally-invoked hook -- nixos/modules/system/boot/stage-2-init.sh,
# which becomes this container's actual /sbin/init, unconditionally
# runs it) but switch-to-configuration behaves differently that early in
# boot (raw stage-2-init.sh, before systemd itself has even started) --
# confirmed live it silently failed to rewrite /etc/shadow from there
# even in "test" mode, despite the exact same command working reliably
# every time when run post-boot (i.e. as a normal systemd service, which
# is what this is). Not fully root-caused why the early context
# specifically breaks it; a real systemd service sidesteps needing to.
#
# /etc/shadow already has PLACEHOLDER entries for every declared user
# baked in at build time (part of constructing the system closure).
# update-users-groups.pl deliberately never overwrites an *existing*
# shadow entry -- a correct safety property in general (don't clobber a
# real user's real password on a config rebuild) -- but on a genuine
# first boot that only means the real hashedPasswordFile-derived hash
# never gets the chance to be applied either, since the placeholder is
# already "seen". Safe to clear here specifically: there is no real
# password yet to protect on a first boot.
#
# "test" mode, not "boot": confirmed live "boot" mode aborts partway
# through (before rewriting /etc/shadow) on a warning that "/boot" is on
# a different filesystem -- a real check for a host with a bootloader to
# update, meaningless for a container that has none
# (boot.loader.{grub,systemd-boot}.enable are both false above), but it
# still aborts the script. "test" runs every activation step without
# touching boot-loader state at all.
#
# ConditionPathExists (systemd-native, not a bash-level check) means
# this only ever runs once, on the genuine first boot -- systemd itself
# skips even starting it on every later boot once the marker exists.
# switch-to-configuration is otherwise the operator's call per this
# repo's own safety rules, not something to run on every boot.
systemd.services.nixos-lxc-first-boot-activate = {
description = "Complete first-boot NixOS activation (users, secrets) for this LXC container";
wantedBy = [ "multi-user.target" ];
unitConfig.ConditionPathExists = "!/var/lib/nixos-lxc-first-boot-activated";
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
};
script = ''
rm -f /etc/shadow
/run/current-system/bin/switch-to-configuration test
mkdir -p /var/lib
touch /var/lib/nixos-lxc-first-boot-activated
'';
};
} }
+11 -12
View File
@@ -1,20 +1,19 @@
{ vars, ... }: _:
{ {
imports = [ ./enable-service.nix ];
services.tailscale = { services.tailscale = {
# Enables the sysctl forwarding settings exit nodes/subnet routers need; enable = true;
# without this, --advertise-exit-node has no effect.
useRoutingFeatures = "server";
# Lets peers reach this node directly over the tailscale UDP port # extraSetFlags (tailscale set, via the always-on tailscaled-set
# instead of relaying through DERP. # service), not extraUpFlags -- extraUpFlags is only ever applied by
openFirewall = true; # tailscaled-autoconnect, which itself only runs when
# services.tailscale.authKeyFile is set (nothing in this repo sets one,
extraUpFlags = [ # so tailscale up is a manual, one-time operator step on every host that
# uses this service). extraSetFlags has no such gate, so
# --advertise-exit-node self-reapplies on every boot once the operator
# has authenticated the node once.
extraSetFlags = [
"--advertise-exit-node" "--advertise-exit-node"
"--advertise-routes=${vars.lanCidr}"
]; ];
}; };
} }
+22 -10
View File
@@ -122,13 +122,25 @@ Add a pre-commit hook (or a `nix flake check` step) running `gitleaks protect --
## Definition of done ## Definition of done
- [ ] Milestone 1 inventory complete and reviewed **Status as of 2026-07-20:** Milestones 13 are done — sops-nix is fully
- [ ] All hosts have per-host age keys; admin key backed up outside the repo wired (`.sops.yaml`, `secrets/*.yaml`, referenced via `hashedPasswordFile`/
- [ ] Every inventoried secret migrated to sops-nix, referenced via `*File`/`sops.secrets.*.path`, nothing plaintext in the working tree `*File`/`sops.secrets.*.path` throughout), and history has been scrubbed
- [ ] `nixos-rebuild dry-build` and at least one real `switch` verified per host with `git-filter-repo` + force-push (this removed a GitHub fine-grained PAT
- [ ] Working-tree scanner sweep clean that had been committed in plaintext in `flake.nix`/`common/home.nix`
- [ ] History rewritten with `git-filter-repo`, force-pushed, full-history scanner sweep clean between 2025-07-16 and 2026-02-09, later migrated to sops but never scrubbed
- [ ] All other clones deleted and re-cloned from the rewritten history from history until now). **Milestone 4 is not confirmed** — whether that PAT
- [ ] Every credential in the original inventory rotated (not just re-encrypted) (or any other historically-plaintext credential) was actually rotated, not
- [ ] Pre-commit secret scanning hook added just re-encrypted, isn't something this repo can attest to; that's an
- [ ] `secrets-inventory.md` deleted from the working directory (never committed) operator action against the issuing service (GitHub, etc.), not a repo
change. Do that before considering this fully closed.
- [x] Milestone 1 inventory complete and reviewed
- [x] All hosts have per-host age keys; admin key backed up outside the repo
- [x] Every inventoried secret migrated to sops-nix, referenced via `*File`/`sops.secrets.*.path`, nothing plaintext in the working tree
- [x] `nixos-rebuild dry-build` and at least one real `switch` verified per host
- [x] Working-tree scanner sweep clean
- [x] History rewritten with `git-filter-repo`, force-pushed, full-history scanner sweep clean
- [ ] All other clones deleted and re-cloned from the rewritten history — every clone that existed before 2026-07-20's rewrite (any other machine, WSL instance, or CI checkout) needs this
- [ ] Every credential in the original inventory rotated (not just re-encrypted) — **the GitHub PAT found in history specifically still needs this**
- [x] Pre-commit secret scanning hook added (`.githooks/pre-commit`, `gitleaks protect --staged`)
- [x] `secrets-inventory.md` deleted from the working directory (never committed)
+148
View File
@@ -0,0 +1,148 @@
#!/usr/bin/env bash
# Backs up the local sops age key (the private key that decrypts
# secrets/*.yaml -- normally the one trusted as &admin) to an arbitrary
# destination path, e.g. a USB drive or other offline storage, so it can
# later be restored and handed to rotate-admin-key.sh if this machine's
# copy is ever lost, or to run either script from a different machine.
#
# Usage:
# scripts/backup-admin-key.sh <dest-path> [--key-file <path>] [--force] [--dry-run]
#
# Source key resolution matches sops/age's own default order:
# $SOPS_AGE_KEY (inline identity text) if set, else
# --key-file if given, else
# $SOPS_AGE_KEY_FILE if set, else
# ${XDG_CONFIG_HOME:-$HOME/.config}/sops/age/keys.txt
set -euo pipefail
repo_root="$(cd "$(dirname "$0")/.." && pwd)"
sops_yaml="${repo_root}/.sops.yaml"
# shellcheck source=env.sh
source "${repo_root}/scripts/env.sh"
# Pin cwd for the same reason rotate-admin-key.sh does: age/sops calls
# below should never depend on wherever the caller's shell happened to be.
cd "$repo_root"
usage() {
cat <<EOF
Usage: $0 <dest-path> [--key-file <path>] [--force] [--dry-run]
<dest-path> Where to write the backup. Parent directories are
created as needed. Written with 0600 permissions.
--key-file <path> Read the key from here instead of the default
sops/age resolution (\$SOPS_AGE_KEY_FILE, then
\${XDG_CONFIG_HOME:-\$HOME/.config}/sops/age/keys.txt).
Ignored if \$SOPS_AGE_KEY is set (that always wins,
same precedence sops/age itself uses).
--force Overwrite <dest-path> if it already exists.
--dry-run Print what would happen; write nothing.
EOF
}
dry_run=0
force=0
key_file="${SOPS_AGE_KEY_FILE:-${XDG_CONFIG_HOME:-$HOME/.config}/sops/age/keys.txt}"
args=()
while [[ $# -gt 0 ]]; do
case "$1" in
--dry-run)
dry_run=1
shift
;;
--force)
force=1
shift
;;
--key-file)
key_file="${2:?--key-file requires a path}"
shift 2
;;
-h | --help)
usage
exit 0
;;
--*)
echo "Unknown option: $1" >&2
usage >&2
exit 1
;;
*)
args+=("$1")
shift
;;
esac
done
if [[ "${#args[@]}" -ne 1 ]]; then
usage >&2
exit 1
fi
dest="${args[0]}"
nix_extra_opts
if [[ -n "${SOPS_AGE_KEY:-}" ]]; then
echo "==> Source: \$SOPS_AGE_KEY (inline identity from the environment)."
src_content="$SOPS_AGE_KEY"
else
[[ -s "$key_file" ]] || {
echo "ERROR: no key found. \$SOPS_AGE_KEY is unset and ${key_file} doesn't exist or is empty." >&2
exit 1
}
echo "==> Source: ${key_file}"
src_content="$(cat "$key_file")"
fi
# Round-trip through a private scratch file (rather than trusting the
# source string as-is) so age-keygen -y validates it's a real identity
# before anything is written to <dest-path>.
scratch="$(mktemp)"
trap 'rm -f "$scratch"' EXIT
( umask 077; printf '%s\n' "$src_content" > "$scratch" )
src_pub="$(nix-shell "${NIX_OPTS[@]}" -p age --run "age-keygen -y '$scratch'")" || {
echo "ERROR: source doesn't look like a valid age identity (age-keygen -y failed)." >&2
exit 1
}
echo " public key: ${src_pub}"
current_admin_pub="$(grep -E '^ - &admin age1' "$sops_yaml" 2>/dev/null | awk '{print $NF}' || true)"
if [[ -n "$current_admin_pub" && "$current_admin_pub" != "$src_pub" ]]; then
echo "NOTE: this key does not match .sops.yaml's current &admin entry (${current_admin_pub})."
echo " Backing it up anyway -- this script doesn't require it to be the admin key."
fi
if [[ -e "$dest" && "$force" -ne 1 ]]; then
echo "ERROR: ${dest} already exists. Pass --force to overwrite." >&2
exit 1
fi
if [[ "$dry_run" -eq 1 ]]; then
echo
echo "[dry-run] would write $(wc -c <"$scratch" | tr -d ' ') bytes to ${dest} (mode 0600)"
[[ -e "$dest" ]] && echo "[dry-run] would overwrite existing file (--force given)"
echo "[dry-run] Nothing was written. Re-run without --dry-run to apply this."
exit 0
fi
mkdir -p "$(dirname "$dest")"
install -m 600 "$scratch" "$dest"
dest_pub="$(nix-shell "${NIX_OPTS[@]}" -p age --run "age-keygen -y '$dest'")"
if [[ "$dest_pub" != "$src_pub" ]]; then
echo "ERROR: ${dest} was written but its public key doesn't match the source -- investigate before relying on this backup." >&2
exit 1
fi
cat <<EOF
Done. Backed up to: ${dest}
public key: ${dest_pub}
This is a private key -- store it somewhere offline/secure, not in this
repo or anywhere it'd get committed. Restore it with:
scripts/rotate-admin-key.sh ${dest}
EOF
+82 -25
View File
@@ -10,7 +10,9 @@
# #
# SAFETY: # SAFETY:
# - The default (create) mode only ever creates a NEW resource -- it # - The default (create) mode only ever creates a NEW resource -- it
# refuses to run if the target VMID already exists on the node. # refuses to run if the target VMID already exists on the node, or if
# a VM/CT identified as --host already exists under any other VMID
# (checked live against the node; --allow-duplicate-host overrides).
# - --modify only ever touches a resource you name explicitly via # - --modify only ever touches a resource you name explicitly via
# --vmid, shows exactly what will change first, and (outside # --vmid, shows exactly what will change first, and (outside
# --dry-run) always requires typing that VMID back to confirm before # --dry-run) always requires typing that VMID back to confirm before
@@ -40,8 +42,12 @@ Create mode (default):
config.networking.hostName (server, docker, config.networking.hostName (server, docker,
nix-cache, nixos, pxe-boot, nix-minimal). Use nix-cache, nixos, pxe-boot, nix-minimal). Use
--list to see what's available for --type. --list to see what's available for --type.
--name <name> Proxmox display name/hostname (default: the flake --name <name> Proxmox display name/hostname (default: --host's
target name, e.g. lxc-server) value, e.g. nix-cache -- for lxc this becomes the
guest's real networking.hostName too, since
proxmoxLXC.manageHostName pulls it from Proxmox's
own container config, so it must match host.nix
regardless of build type)
--vmid <n> Numeric VMID (default: next free, via --vmid <n> Numeric VMID (default: next free, via
\`pvesh get /cluster/nextid\` on the node). \`pvesh get /cluster/nextid\` on the node).
Refuses to run if this ID already exists. Refuses to run if this ID already exists.
@@ -52,10 +58,11 @@ Create mode (default):
--force-rebuild Skip the "does the node already have this --force-rebuild Skip the "does the node already have this
image" check -- always build fresh and image" check -- always build fresh and
overwrite what's there. overwrite what's there.
--allow-duplicate-host Required if --host already has a real --allow-duplicate-host Required if a VM/CT identified as --host
deployment elsewhere (variables.nix's already exists on the node (checked live via
deployedTargets) -- otherwise refused, since qm/pct, not any file in this repo) --
it'd share that host's hostName/hostId. otherwise refused, since it'd share that
host's hostName/hostId.
Modify mode (reconfigure an EXISTING resource -- requires --modify): Modify mode (reconfigure an EXISTING resource -- requires --modify):
--modify Switch to modify mode. --modify Switch to modify mode.
@@ -274,27 +281,66 @@ if [[ -z "$flake_target" ]]; then
exit 1 exit 1
fi fi
[[ -z "$name" ]] && name="$flake_target" # The container/VM's real identity is --host (e.g. "nix-cache"), validated
# above against config.networking.hostName -- not the flake target name
# (e.g. "lxc-nix-cache"), which is build-type-specific and only exists to
# pick which platform variant to build. Defaulting --name to the flake
# target would make lxc's --hostname (which proxmoxLXC.manageHostName
# feeds straight into the guest's real hostname) disagree with host.nix.
[[ -z "$name" ]] && name="$host"
# --- refuse to duplicate a host that's already really deployed ---------- # --- refuse to duplicate a host that's already live on the node ---------
# Checked by hostName, not exact flake target: proxmox-server being # Queries the node itself (qm/pct's own name/hostname config), not any
# deployed also blocks --type lxc --host server, since both would carry # static list in this repo -- a file can't track whether a resource still
# the same hosts/server/host.nix identity (hostName, hostId). # actually exists, and this used to be checked against variables.nix's
if [[ "$allow_duplicate_host" -eq 0 ]]; then # deployedTargets, which drifted stale (it kept naming a VM as "the real
deployed_targets_json="$(nix eval --json --no-use-registries --no-accept-flake-config \ # deployment" well after that VM had been destroyed, blocking its own
--file "${repo_root}/variables.nix" deployedTargets)" # redeploy) until that list was dropped in favour of this live check. This
for dt in $(echo "$deployed_targets_json" | jq -r '.[]'); do # only catches guests identified with the default --name (== --host, what
dt_hostname="$(nix eval --raw --no-use-registries --no-accept-flake-config \ # this script itself always uses unless --name is overridden) -- a guest
"${repo_root}#nixosConfigurations.${dt}.config.networking.hostName" 2>/dev/null || true)" # manually renamed on the node afterwards wouldn't match, but nothing here
if [[ "$dt_hostname" == "$host" ]]; then # creates guests that way.
echo "ERROR: '${host}' already has a real deployment (${dt}, per variables.nix's" >&2 if [[ "$allow_duplicate_host" -eq 1 ]]; then
echo "deployedTargets). Creating ${flake_target} would share its hostName/hostId --" >&2 echo
echo "refusing by default. Pass --allow-duplicate-host if you really mean to spin" >&2 echo "--allow-duplicate-host: skipping the check for an existing '${host}' on ${node}."
echo "up a separate test instance of this host (it'll still get its own distinct" >&2 elif [[ "$dry_run" -eq 1 ]]; then
echo "sops key and VMID, never touching ${dt})." >&2 echo
echo "[dry-run] would check ${node} for an existing VM/CT identified as '${host}'"
else
echo
echo "==> Checking ${node} for an existing VM/CT identified as '${host}'..."
ssh_check_status=0
existing="$(ssh "$ssh_target" bash -s -- "$host" <<'REMOTE_SCRIPT'
target="$1"
for id in $(qm list 2>/dev/null | awk 'NR>1{print $1}'); do
n="$(qm config "$id" 2>/dev/null | grep -oP '^name:\s*\K\S+' || true)"
[[ "$n" == "$target" ]] && echo "vm ${id} ${n}"
done
for id in $(pct list 2>/dev/null | awk 'NR>1{print $1}'); do
n="$(pct config "$id" 2>/dev/null | grep -oP '^hostname:\s*\K\S+' || true)"
[[ "$n" == "$target" ]] && echo "lxc ${id} ${n}"
done
exit 0
REMOTE_SCRIPT
)" || ssh_check_status=$?
if [[ "$ssh_check_status" -ne 0 ]]; then
echo "ERROR: couldn't reach ${node} (ssh exited ${ssh_check_status}) to check for an" >&2
echo "existing '${host}' resource -- refusing to guess. Fix connectivity and retry," >&2
echo "or pass --allow-duplicate-host if you're sure none exists (this skips the" >&2
echo "check entirely)." >&2
exit 1 exit 1
fi fi
if [[ -n "$existing" ]]; then
echo "ERROR: '${host}' already exists on ${node}:" >&2
echo "$existing" | while read -r kind id n; do
echo " - ${kind} VMID ${id} (${n})" >&2
done done
echo "Refusing to create a second resource sharing this identity. Pass" >&2
echo "--allow-duplicate-host to create one anyway (it gets its own distinct" >&2
echo "sops key and VMID -- the existing resource(s) above are left untouched)," >&2
echo "or use --modify to reconfigure the existing one instead." >&2
exit 1
fi
fi fi
echo "Target: ${flake_target} (host=${host}, type=${type}) -> Proxmox resource '${name}'" echo "Target: ${flake_target} (host=${host}, type=${type}) -> Proxmox resource '${name}'"
@@ -383,9 +429,19 @@ else
fi fi
if [[ "$image_already_remote" -eq 0 && -z "$local_image" ]]; then if [[ "$image_already_remote" -eq 0 && -z "$local_image" ]]; then
# Mirrors the real build commands' "${NIX_OPTS[@]}" below -- nix_extra_opts
# (called earlier, once) has already decided whether nix-cache is in play,
# and the dry-run preview needs to reflect that decision instead of always
# printing the same command regardless of outcome.
nix_opts_display=""
if [[ ${#NIX_OPTS[@]} -gt 0 ]]; then
printf -v nix_opts_display '%q ' "${NIX_OPTS[@]}"
nix_opts_display=" ${nix_opts_display% }"
fi
if [[ "$type" == "lxc" ]]; then if [[ "$type" == "lxc" ]]; then
if [[ "$dry_run" -eq 1 ]]; then if [[ "$dry_run" -eq 1 ]]; then
echo "[dry-run] would build: NIXOS_HOST_KEYS_DIR=${repo_root}/host-keys nix build --impure \\" echo "[dry-run] would build: NIXOS_HOST_KEYS_DIR=${repo_root}/host-keys nix build --impure \\"
echo "[dry-run] --no-use-registries --no-accept-flake-config${nix_opts_display} \\"
echo "[dry-run] .#nixosConfigurations.${flake_target}.config.system.build.tarball" echo "[dry-run] .#nixosConfigurations.${flake_target}.config.system.build.tarball"
local_image="<built-tarball>" local_image="<built-tarball>"
else else
@@ -399,7 +455,8 @@ if [[ "$image_already_remote" -eq 0 && -z "$local_image" ]]; then
fi fi
else else
if [[ "$dry_run" -eq 1 ]]; then if [[ "$dry_run" -eq 1 ]]; then
echo "[dry-run] would build: nix build .#nixosConfigurations.${flake_target}.config.system.build.diskoImagesScript" echo "[dry-run] would build: nix build --no-use-registries --no-accept-flake-config${nix_opts_display} \\"
echo "[dry-run] .#nixosConfigurations.${flake_target}.config.system.build.diskoImagesScript"
echo "[dry-run] would run: sudo ./result-${flake_target} \\" echo "[dry-run] would run: sudo ./result-${flake_target} \\"
echo "[dry-run] --pre-format-files host-keys/${flake_target}_ssh_host_ed25519_key /etc/ssh/ssh_host_ed25519_key \\" echo "[dry-run] --pre-format-files host-keys/${flake_target}_ssh_host_ed25519_key /etc/ssh/ssh_host_ed25519_key \\"
echo "[dry-run] --pre-format-files host-keys/${flake_target}_ssh_host_ed25519_key.pub /etc/ssh/ssh_host_ed25519_key.pub \\" echo "[dry-run] --pre-format-files host-keys/${flake_target}_ssh_host_ed25519_key.pub /etc/ssh/ssh_host_ed25519_key.pub \\"
+49 -2
View File
@@ -88,13 +88,60 @@ nix_extra_opts() {
fi fi
export NIX_EXTRA_OPTS_DECIDED=1 export NIX_EXTRA_OPTS_DECIDED=1
NIX_OPTS=() NIX_OPTS=()
if ! curl --silent --fail --max-time 3 "http://${NIX_CACHE_HOST}/nix-cache-info" >/dev/null 2>&1; then
# Retry a couple of times, 1s apart, before believing either check --
# belt-and-suspenders against a genuine multi-second blip (nix-cache
# restarting), on top of the fix below. Worst case (~11s total, host
# genuinely gone) is still nowhere near the 15s+ *per lookup* nix's own
# substituter retries would cost if this check didn't exist at all.
local attempt cache_up=0 builder_up=0
for attempt in 1 2 3; do
if curl --silent --fail --max-time 3 "http://${NIX_CACHE_HOST}/nix-cache-info" >/dev/null 2>&1; then
cache_up=1
break
fi
[[ "$attempt" -lt 3 ]] && sleep 1
done
if [[ "$cache_up" -eq 0 ]]; then
echo "nix-cache (http://${NIX_CACHE_HOST}) is unreachable -- skipping it (substituter + remote builder) for the rest of this run." >&2 echo "nix-cache (http://${NIX_CACHE_HOST}) is unreachable -- skipping it (substituter + remote builder) for the rest of this run." >&2
NIX_OPTS=(--option substituters "https://cache.nixos.org/" --builders "") NIX_OPTS=(--option substituters "https://cache.nixos.org/" --builders "")
elif ! timeout 3 bash -c "cat < /dev/tcp/${NIX_CACHE_HOST}/22" >/dev/null 2>&1; then else
for attempt in 1 2 3; do
# `exec 3<>/dev/tcp/...` just opens the fd and returns -- it does NOT
# read from it. Confirmed live this is load-bearing, not stylistic:
# the previous `cat < /dev/tcp/.../22` blocked forever and always hit
# the timeout even against a perfectly healthy nix-cache, because
# sshd sends its banner and then holds the connection open waiting
# for the client to speak next -- `cat` never sees EOF, so this
# check reported "unreachable" unconditionally, 100% of the time,
# regardless of whether the remote builder was actually up.
if timeout 3 bash -c "exec 3<>/dev/tcp/${NIX_CACHE_HOST}/22" 2>/dev/null; then
builder_up=1
break
fi
[[ "$attempt" -lt 3 ]] && sleep 1
done
if [[ "$builder_up" -eq 0 ]]; then
echo "nix-cache's SSH remote builder (nixremote@${NIX_CACHE_HOST}:22) is unreachable -- disabling remote builds for the rest of this run." >&2 echo "nix-cache's SSH remote builder (nixremote@${NIX_CACHE_HOST}:22) is unreachable -- disabling remote builds for the rest of this run." >&2
NIX_OPTS=(--builders "") NIX_OPTS=(--builders "")
fi fi
fi
# `printf '%q '` with a genuinely empty NIX_OPTS still runs one format
# pass over a missing argument and yields the literal `'' ` rather than
# an empty string (confirmed live) -- a subprocess that later does
# `eval "NIX_OPTS=(${NIX_EXTRA_OPTS})"` (the branch above, for e.g.
# sync-host-keys.sh reusing this process's decision) would then rebuild
# a 1-element array holding an empty string instead of a 0-element
# array, and `nix-shell "${NIX_OPTS[@]}" -p <pkg>` chokes on that stray
# element as a bogus positional argument. Guard the empty case
# explicitly so nix-cache being reachable (NIX_OPTS legitimately empty)
# round-trips as truly empty instead.
if [[ ${#NIX_OPTS[@]} -gt 0 ]]; then
printf -v NIX_EXTRA_OPTS '%q ' "${NIX_OPTS[@]}" printf -v NIX_EXTRA_OPTS '%q ' "${NIX_OPTS[@]}"
else
NIX_EXTRA_OPTS=""
fi
export NIX_EXTRA_OPTS export NIX_EXTRA_OPTS
} }
+190
View File
@@ -0,0 +1,190 @@
#!/usr/bin/env bash
# Rotates the &admin sops age key: decrypts with a backed-up copy of the
# key CURRENTLY trusted as &admin, replaces .sops.yaml's &admin entry with
# a new key already present in this environment, and re-encrypts every
# secrets/*.yaml for the new recipient set. After this runs, the old key
# can no longer decrypt anything -- this is a real, one-way handoff of
# trust, not a preview.
#
# This is the automation for the manual steps create-proxmox-resource.sh /
# sync-host-keys.sh print when they bootstrap a brand-new, not-yet-trusted
# age key on a machine that's never had admin access before:
#
# scripts/rotate-admin-key.sh /path/to/backed-up/admin/keys.txt
#
# The backup key's *public* key must match .sops.yaml's current &admin
# entry -- this script verifies that by deriving it, it doesn't just trust
# the filename or take it on faith. The new key defaults to wherever sops
# itself would already look ($SOPS_AGE_KEY_FILE, then the XDG default), so
# the common case is just pointing this at the restored backup.
set -euo pipefail
repo_root="$(cd "$(dirname "$0")/.." && pwd)"
sops_yaml="${repo_root}/.sops.yaml"
# shellcheck source=env.sh
source "${repo_root}/scripts/env.sh"
# sops resolves .sops.yaml by walking up from the process's cwd, not from
# the target file's own path -- if this script were invoked from somewhere
# other than the repo root (or from inside another checkout/worktree that
# happens to have its own .sops.yaml), `sops updatekeys` would silently
# re-encrypt against the WRONG config's recipient list instead of this
# repo's. Pin cwd here so every sops/age call below is unambiguous
# regardless of where the caller's shell started out.
cd "$repo_root"
usage() {
cat <<EOF
Usage: $0 <path-to-backed-up-admin-key> [--new-key-file <path>] [--dry-run]
<path-to-backed-up-admin-key> age identity file for the key CURRENTLY
trusted as &admin. Only ever read -- never
copied or modified.
--new-key-file <path> age identity file for the key to promote
to &admin. Defaults to \$SOPS_AGE_KEY_FILE,
then
\${XDG_CONFIG_HOME:-\$HOME/.config}/sops/age/keys.txt
(sops/age's own default resolution order).
--dry-run Print what would change; touches nothing
(.sops.yaml untouched, no sops updatekeys
calls).
EOF
}
dry_run=0
new_key_file="${SOPS_AGE_KEY_FILE:-${XDG_CONFIG_HOME:-$HOME/.config}/sops/age/keys.txt}"
args=()
while [[ $# -gt 0 ]]; do
case "$1" in
--dry-run)
dry_run=1
shift
;;
--new-key-file)
new_key_file="${2:?--new-key-file requires a path}"
shift 2
;;
-h | --help)
usage
exit 0
;;
--*)
echo "Unknown option: $1" >&2
usage >&2
exit 1
;;
*)
args+=("$1")
shift
;;
esac
done
if [[ "${#args[@]}" -ne 1 ]]; then
usage >&2
exit 1
fi
backup_key="${args[0]}"
[[ -s "$backup_key" ]] || { echo "ERROR: backup key file not found or empty: ${backup_key}" >&2; exit 1; }
[[ -s "$new_key_file" ]] || { echo "ERROR: new key file not found or empty: ${new_key_file}" >&2; exit 1; }
nix_extra_opts
age_pub() {
nix-shell "${NIX_OPTS[@]}" -p age --run "age-keygen -y '$1'"
}
echo "==> Deriving public keys..."
old_pub="$(age_pub "$backup_key")"
new_pub="$(age_pub "$new_key_file")"
echo " backup (old admin) key: ${old_pub}"
echo " new admin key: ${new_pub}"
if [[ "$old_pub" == "$new_pub" ]]; then
echo "ERROR: backup key and new key are identical -- nothing to rotate." >&2
exit 1
fi
current_admin_line="$(grep -E '^ - &admin age1' "$sops_yaml" || true)"
if [[ -z "$current_admin_line" ]]; then
echo "ERROR: couldn't find a '&admin age1...' line in ${sops_yaml}." >&2
exit 1
fi
current_admin_pub="$(awk '{print $NF}' <<<"$current_admin_line")"
if [[ "$current_admin_pub" != "$old_pub" ]]; then
echo "ERROR: ${backup_key} doesn't match the current &admin key in .sops.yaml." >&2
echo " .sops.yaml &admin: ${current_admin_pub}" >&2
echo " backup key pubkey: ${old_pub}" >&2
echo "Wrong backup file, or .sops.yaml has already moved on -- not touching anything." >&2
exit 1
fi
mapfile -t secrets_files < <(find "${repo_root}/secrets" -maxdepth 1 -name '*.yaml' | sort)
if [[ "${#secrets_files[@]}" -eq 0 ]]; then
echo "ERROR: no secrets/*.yaml files found under ${repo_root}/secrets." >&2
exit 1
fi
echo "==> Confirming the backup key can actually decrypt..."
if ! SOPS_AGE_KEY_FILE="$backup_key" nix-shell "${NIX_OPTS[@]}" -p sops --run \
"sops -d '${secrets_files[0]}'" >/dev/null; then
echo "ERROR: backup key failed to decrypt $(basename "${secrets_files[0]}") -- aborting." >&2
exit 1
fi
echo " OK: decrypted $(basename "${secrets_files[0]}")"
if [[ "$dry_run" -eq 1 ]]; then
echo
echo "[dry-run] would replace .sops.yaml's &admin line:"
echo "[dry-run] - ${current_admin_pub}"
echo "[dry-run] + ${new_pub}"
echo "[dry-run] would then re-encrypt (sops updatekeys --yes) for the new recipient set:"
for f in "${secrets_files[@]}"; do
echo "[dry-run] secrets/$(basename "$f")"
done
echo
echo "[dry-run] Nothing was changed. Re-run without --dry-run to apply this."
exit 0
fi
echo "==> Rotating .sops.yaml's &admin key..."
sed -i "s|^ - &admin age1[a-z0-9]*| - \&admin ${new_pub}|" "$sops_yaml"
grep -qF "$new_pub" "$sops_yaml" || {
echo "ERROR: sed edit didn't take -- .sops.yaml left unchanged, check it by hand." >&2
exit 1
}
echo " Updated."
echo "==> Re-encrypting secrets/*.yaml for the new recipient set..."
for f in "${secrets_files[@]}"; do
echo "==> $(basename "$f")"
SOPS_AGE_KEY_FILE="$backup_key" nix-shell "${NIX_OPTS[@]}" -p sops --run \
"sops updatekeys --yes '${f}'"
done
echo "==> Verifying the new key can decrypt everything..."
for f in "${secrets_files[@]}"; do
if ! SOPS_AGE_KEY_FILE="$new_key_file" nix-shell "${NIX_OPTS[@]}" -p sops --run \
"sops -d '${f}'" >/dev/null; then
echo "ERROR: new key failed to decrypt $(basename "$f") after rotation -- investigate before committing." >&2
exit 1
fi
echo " OK: $(basename "$f")"
done
cat <<EOF
Done. .sops.yaml's &admin key is now:
${new_pub}
The old key (${old_pub}) can no longer decrypt any secrets/*.yaml
re-encrypted above.
Review the diff, then commit:
git add .sops.yaml secrets/*.yaml
git commit -m "Rotate sops admin age key"
EOF
+9 -16
View File
@@ -19,6 +19,15 @@
remoteBuilderUser = "nixremote"; # remote builder SSH user remoteBuilderUser = "nixremote"; # remote builder SSH user
# nix-cache's own SSH host public key (not a secret — the private half
# never leaves the host). Wired into every client's
# programs.ssh.knownHosts by modules/nix-cache/remote-builder-client.nix
# so distributed builds don't hit "Host key verification failed" on a
# fresh client that has never manually ssh'd to nix-cache before. Update
# this if nix-cache's host key is ever rotated or the host is rebuilt
# from scratch.
nixCacheHostKey = "ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIHrMKZlIGUd3pH9G3AqbsruqUGjxIXMAZw52u9MwiBCn lxc-nix-cache";
# Public keys authorized to SSH in as remoteBuilderUser on the nix-cache # Public keys authorized to SSH in as remoteBuilderUser on the nix-cache
# host (modules/nix-cache/server.nix) — one per client host that's allowed # host (modules/nix-cache/server.nix) — one per client host that's allowed
# to use it as a distributed builder. # to use it as a distributed builder.
@@ -147,20 +156,4 @@
keep = 20; # number of rotated logs to retain before deleting the oldest keep = 20; # number of rotated logs to retain before deleting the oldest
}; };
# Flake targets with a real, currently-running deployment somewhere —
# matches README.md's Hosts table "(real, deployed)" annotations; update
# both together. Not consumed by any NixOS module (nothing in the actual
# system config should behave differently because of this) — it's read
# by scripts/create-proxmox-resource.sh to refuse creating a same-identity
# duplicate of an already-deployed host (shared hostName/hostId) unless
# you explicitly pass --allow-duplicate-host.
deployedTargets = [
"linode-minimal"
"proxmox-minimal"
"proxmox-nix-cache"
"proxmox-server"
"proxmox-docker"
"proxmox-gui"
"proxmox-pxe-boot"
];
} }