This repository has been archived on 2026-07-30. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
nixos/CLAUDE.md
T
beatzaplentyandClaude Sonnet 4.6 f22ff7db79
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m59s
refactor(beszel): move KEY and TOKEN fully into enable-agent.nix
KEY (hub's SSH public key) is now stored in variables.nix as beszelHubKey
and set by enable-agent.nix — no host file ever needs to touch it.
TOKEN was already universal; this removes the last per-host beszel config.

Importing enable-agent.nix in a build type is now the only step needed to
add a new host to beszel monitoring.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-29 12:37:54 +10:00

26 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Repo purpose

Flake-based NixOS configuration for Wayne's LAN servers and workstation. There is no application code here — changes are Nix module edits that affect real machines when deployed.

Safety rules (read before touching anything)

  • Never run nixos-rebuild switch|boot|test, nixos-install, parted, mkfs, mkswap, swapon, mount, or any other destructive disk/deploy command from an agent session, even if asked indirectly. Deployment is done manually by the operator on the target host.
  • Validation is limited to evaluation, linting, formatting checks, and nix build --dry-run --no-link.
  • Do not add secrets, tokens, private keys, or new password hashes to the repo.
  • This repo currently contains committed password hashes in modules/installer/common.nix (the auto-installer's own root/nixos login — a deliberate, documented choice, see docs/auto-installer.md, not accidental tech debt) and SSH public keys in variables.nix (vars.adminSshKey, vars.remoteBuilderAuthorizedKeys, vars.beszelHubKey). Don't use the installer's hardcoded hash as a template for a real host — every other host uses sops-nix (hashedPasswordFile, see "Security Notes" in README.md). Flag any new secret-like string you encounter instead of committing it.
  • host-keys/ is gitignored — used only by the auto-installer's own environment for pre-seeding non-LXC host keys before first boot (see docs/auto-installer.md). Never commit its contents; if git status ever shows it as trackable, something is wrong. All deployed hosts use clan vars (vars/per-machine/<target>/openssh/, committed and sops-encrypted) for their SSH host keys — those ARE tracked by git and belong in the repo.

Two Proxmox nodes: pve1.sweet.home (production) and pve-test.sweet.home (sandbox)

There are two SSH-reachable Proxmox nodes on the LAN, both defined in scripts/env.sh (PVE1_HOST / PVE_TEST_HOST), individually targetable via scripts/proxmox/create-proxmox-resource.sh --node <host> or by overriding PROXMOX_HOST. PROXMOX_HOST itself still defaults to PVE1_HOST (production) — that default, and every other script behavior, is unchanged from before pve-test existed; the only thing new is that pve-test can now be reached at all. They are not interchangeable — one is real production infrastructure, the other exists specifically so there's somewhere safe to test. The restriction below is a policy for Claude specifically, not a change to the tooling's own default or anything the operator needs to opt into.

pve1.sweet.home (production — off-limits to Claude)

A real, live Proxmox node hosting production VMs/containers — not a sandbox, and not Claude's to touch by default.

  • Off-limits at all times unless the operator has given explicit, same-session instructions to act on this specific host. That authorization is scoped to the task it was given for — don't carry it forward to unrelated later work in the same conversation, and never assume it from a previous session.
  • Read-only for existing state is always fine, authorization or not. You may SSH in (or use pvesm, qm list, pct list, qm config, pct config, the Proxmox API, etc.) to inspect the node's config, storage, and any existing VM/container — including ones this repo didn't create.
  • Never modify, stop, restart, delete, reconfigure, or create anything on this node (qm set, pct set, qm destroy, pct destroy, qm stop, pct stop, qm create, pct create, snapshot operations, storage changes, etc.) — including scratch/test resources — without that explicit go-ahead. Use pve-test.sweet.home for anything exploratory instead; it exists precisely so pve1 never has to be the answer to "where do I test this."
  • This is a Claude-specific policy, not something the scripts enforce. scripts/env.sh/create-proxmox-resource.sh default to pve1 exactly as they did before pve-test existed, with no extra flag or prompt required — that's deliberate, so the operator's own existing workflows don't change. Claude, however, must never rely on that default: every Proxmox action Claude takes on its own initiative — not explicitly pointed at pve1 by the operator this session — targets pve-test instead (e.g. --node "$PVE_TEST_HOST", or PROXMOX_HOST=$PVE_TEST_HOST). Claude's own default is pve-test, full stop, regardless of what the tooling's own unqualified default happens to be.

pve-test.sweet.home (sandbox — Claude's default target)

A separate Proxmox node set aside for testing. The tooling's default is still production (PROXMOX_HOSTPVE1_HOST, see above) — but Claude's own default is this node: absent an explicit, same-session instruction to use pve1, every Proxmox action Claude initiates targets pve-test. Once targeted, it's safe to create, interrogate, and destroy resources on without asking first.

  • Test VMs/containers are allowed, but must be torn down. Create a scratch VM or container here (e.g. via scripts/proxmox/create-proxmox-resource.sh or raw qm/pct create) to validate something. Anything created this way must be destroyed again in the same session, before ending the task — never leave a test resource running. Use a VMID/name that's obviously scratch (and doesn't collide with a real flake target) so it's unambiguous what's safe to remove.
  • Node-level config is still not yours to change. Creating/destroying your own scratch guests is fine; Proxmox host config, storage pools, and networking on pve-test itself are still the operator's call to make manually, same as on pve1.

Commands

# One-time environment bootstrap (installs Nix if missing, prints hosts)
bash scripts/codex-setup.sh

# Changed-files-only validation: secret grep (whole repo), nixpkgs-fmt --check
# and statix on changed *.nix files, eval of the hosts/packages those changes
# can affect. This is what CI runs on every push/PR.
bash scripts/codex-maintenance.sh

# Full sweep: nixpkgs-fmt --check/statix over the whole tree, eval every host
# and package. Slow (minutes) -- CI never runs this; use it locally before a
# release or after touching modules/common/*, flake.nix, or variables.nix for
# extra confidence beyond the automatic full-fallback those paths already
# trigger in the default mode (see below).
bash scripts/codex-maintenance.sh --full-check

# Either mode, plus a dry-run build (no result symlink) of every host/package
# in whichever scope is active
bash scripts/codex-maintenance.sh --dry-run
bash scripts/codex-maintenance.sh --full-check --dry-run

# List the hosts the flake currently exposes
nix eval --json .#nixosConfigurations --apply builtins.attrNames | jq -r '.[]'

# Evaluate a single host without building (fast sanity check)
nix eval .#nixosConfigurations.<host>.config.system.build.toplevel.drvPath --raw

# Dry-run build a single host
nix build --dry-run --no-link .#nixosConfigurations.<host>.config.system.build.toplevel

Formatting/lint tools (nixpkgs-fmt, statix) are not installed locally; the maintenance script pulls them via nix run github:NixOS/nixpkgs/nixos-25.11#<tool>. There is no test suite — "correctness" here means the flake evaluates and nixpkgs-fmt/statix are clean.

With no flags, codex-maintenance.sh diffs against a base ref (env MAINT_BASE_SHA, else the PR base SHA in CI, else HEAD^ locally) and scopes fmt-check/statix to the changed *.nix files and eval to the hosts/packages those changes can affect — a hosts/<name>/host.nix edit only evals that host's targets, a modules/platforms/<platform>.nix edit only evals that platform's hosts, and so on. A change to flake.nix, flake.lock, variables.nix, modules/common/*, or any other modules/*.nix file outside platforms//build-types/ (whose blast radius isn't safely inferable from the path alone) falls back to evaluating every host and package, same as --full-check would, just without the whole-tree fmt/statix sweep. This exists because the whole-tree sweep is what was timing out CI; CI always runs the plain, no-flag form and never passes --full-check.

The default mode's diff is against the working tree (uncommitted and staged edits included, not just committed ones), so it's already the right tool for an interactive session too: after editing one or two hosts/modules, plain bash scripts/codex-maintenance.sh naturally scopes to just what you touched. Reserve --full-check for changes that plausibly affect every host (modules/common/*, flake.nix, variables.nix — though the default mode already falls back to evaluating everything for those paths, --full-check additionally re-checks fmt/statix over the whole tree) or as a final check before committing.

Scripts

Beyond codex-setup.sh/codex-maintenance.sh above, scripts/ is organized by purpose: scripts/secrets/ (sops/age + SSH host-key management), scripts/proxmox/ (Proxmox deployment), scripts/installer/ (the auto-installer's own shell script, templated into the image — see below), scripts/lib/ (shared helpers, sourced by the scripts below — not run directly), and a handful of repo-wide scripts left at the top level (env.sh, bump-nixpkgs-release.sh, plus codex-setup.sh/ codex-maintenance.sh above). When adding a new script, put it in the matching subfolder rather than the top level, and if it duplicates logic another script already has, lift the shared part into scripts/lib/ instead of copying it.

scripts/installer/

  • scripts/installer/auto-install.sh — the interactive install script baked into the auto-installer image (see docs/auto-installer.md), kept as a real, version-controlled shell file rather than inline in modules/installer/common.nix's Nix. It sources scripts/env.sh itself for LAN_DOMAIN (export LAN_DOMAIN/: "${LAN_DOMAIN:=...}", matching variables.nix's lanDomain — manually kept in sync, same pattern as NIX_CACHE_HOST mirroring nixCacheHost), rather than Nix-level string substitution — that's what makes it work identically whether run straight from a git checkout or from inside the built installer image. common.nix bakes scripts/env.sh in alongside it at a matching relative path (/etc/nixos-installer/env.sh next to /etc/nixos-installer/installer/auto-install.sh) so the script's own source "$(dirname ...)/../env.sh" line resolves the same way in both contexts — this is also why it's invoked from /etc/nixos-installer/installer/auto-install.sh rather than a flat /etc/auto-install.sh. #!/usr/bin/env bash, not #!/run/current-system/sw/bin/bash: the latter only resolves on an already-activated NixOS system, breaking the checked-out-file case entirely (confirmed live: "cannot execute: required file not found" on a non-NixOS box); /usr/bin/env is reliably present on both NixOS (environment.usrbinenv's own default) and any normal Linux distro.

scripts/secrets/

  • scripts/secrets/sync-host-keys.sh — generates/registers SSH host keys and their .sops.yaml/secrets/*.yaml recipients for flake targets, idempotently (--all, <target>, --remove, --regenerate-all-keys, all with --dry-run). Stores keys as clan vars (vars/per-machine/<target>/openssh/, committed and sops-encrypted) for all flake targets. The primary tool for provisioning a new host's secrets access — see "Creating a new machine" in docs/auto-installer.md.
  • scripts/secrets/prepare-host-key.sh — narrower predecessor: generates a key by an arbitrary name without touching .sops.yaml. Still useful to pre-generate a key before its flake target exists yet, since sync-host-keys.sh can only act on targets nixosConfigurations already has.
  • scripts/secrets/rotate-admin-key.sh <backup-admin-key> [--new-key-file <path>] [--dry-run] — rotates .sops.yaml's &admin age key: decrypts with a backed-up copy of the key currently trusted as &admin (verified by deriving its public key and comparing, not taken on faith), replaces the &admin line with a new key already present in the environment (defaults to wherever sops/age itself would look), and runs sops updatekeys on every secrets/*.yaml. One-way: the old key can no longer decrypt anything re-encrypted this way. This is the automation for the manual steps sync-host-keys.sh/create-proxmox-resource.sh print when they bootstrap a brand-new, not-yet-trusted key on a machine with no prior admin access.
  • scripts/secrets/backup-admin-key.sh <dest-path> [--key-file <path>] [--force] [--dry-run] — copies the local sops age key (source resolution matches sops/age itself: $SOPS_AGE_KEY inline, then --key-file, then $SOPS_AGE_KEY_FILE, then the XDG default) to an arbitrary destination path with 0600 permissions, validating it's a real age identity and round-tripping the public key before and after the write. Refuses to overwrite an existing <dest-path> without --force. Purely a local filesystem copy — never touches .sops.yaml/ secrets/*.yaml or the repo at all. The resulting file is exactly what rotate-admin-key.sh expects as its backup-key argument.
  • scripts/secrets/sync-nix-cache-host-key.sh [--check] [--dry-run] [--host <name>] — detects drift between the ed25519 SSH host key nix-cache is actually serving right now (via ssh-keyscan) and vars.nixCacheHostKey (variables.nix), the value modules/nix-cache/remote-builder-client.nix bakes into every real client's declarative programs.ssh.knownHosts and configure-nix-cache-client.sh hardcodes as its own default for non-NixOS clients. That value has no automatic source of truth — it's set once from whatever nix-cache's host key happened to be at the time, and silently goes stale if the host is ever rebuilt/recreated with a new key, breaking every client's distributed-build SSH trust with no error that points back here. --check (used by codex-maintenance.sh, which treats an unreachable nix-cache — e.g. from a non-LAN CI runner — as a silent skip rather than a failure) only reports drift; the no-flags form updates both files in place. Declarative clients still need a rebuild to pick up the fix.

scripts/proxmox/

  • scripts/proxmox/create-proxmox-resource.sh — builds a lxc-*/ proxmox-* target's tarball/disk image and creates it on a real Proxmox node (pct create against the tarball as a CT template / qm create+ importdisk), or reconfigures an existing resource's cores/memory/disk size (--modify, always requires typing the VMID back to confirm). Checks for an already-uploaded image on the node before building (--force-rebuild to skip that and always rebuild), and probes nix-cache's substituter/remote-builder reachability once up front rather than letting every nix build call retry against it individually. Refuses to create a target whose host identity already exists live on the node (checked directly via qm/pct, not any file in this repo) unless --allow-duplicate-host is passed. --dry-run throughout both modes. The first time it has to bootstrap build tooling on a node (i.e. nix wasn't already on its PATH), it also runs scripts/proxmox/configure-nix-cache-client.sh there (non-fatally — a failure just falls back to building from source / cache.nixos.org) so the node substitutes from and can offload builds to nix-cache on every subsequent run, not just this one.
  • scripts/proxmox/configure-nix-cache-client.sh [--dry-run] [--no-remote-builder] [--no-restart] — the non-NixOS equivalent of modules/nix-cache/client.nix/remote-builder-client.nix, for a plain Debian machine with the Nix package manager (not NixOS) already installed: run as root on that machine to add nix-cache as a substituter in /etc/nix/nix.conf (https://cache.nixos.org/ kept as fallback) via extra-substituters/extra-trusted-public-keys so it layers on top of whatever's already there instead of clobbering it, and, if /root/.ssh/nixremote is already present (see docs/nix-cache.md "Remote builder SSH keys"), configures it as a distributed-build machine too and trusts nix-cache's SSH host key in /etc/ssh/ssh_known_hosts. Idempotent (re-running replaces its own marked block rather than duplicating it); restarts nix-daemon by default so the change takes effect immediately.

scripts/lib/

Sourced by the scripts above, never run directly:

  • nix-bootstrap.shNIX_CONFIG/ensure_nix_profile, shared by codex-setup.sh/codex-maintenance.sh and the remote build commands create-proxmox-resource.sh runs over SSH.
  • nix-eval.shNIX_EVAL_FLAGS plus list_flake_targets/ flake_target_hostname flake-introspection helpers.
  • ssh-host-keys.shgenerate_host_ed25519_key/ssh_pubkey_to_age, shared by sync-host-keys.sh and prepare-host-key.sh.
  • sops-age.shage_pubkey_from_identity_file/sops_yaml_admin_pubkey/ sops_updatekeys plus the shared sops/age default key-file resolution, shared by backup-admin-key.sh, rotate-admin-key.sh, and sync-host-keys.sh.
  • confirm.shconfirm_typed, the "type X back to confirm" destructive- action prompt shared by create-proxmox-resource.sh and sync-host-keys.sh.
  • sync-host-keys-edit-sops.py — the .sops.yaml anchor/key_groups editor sync-host-keys.sh shells out to (see that script for why: precise, idempotent YAML edits are impractical in bash).

Top level

  • scripts/env.sh — shared config (PROXMOX_HOST, storage pool, bridge, default cores/memory, NIX_CACHE_HOST, LAN_DOMAIN) sourced by create-proxmox-resource.sh and scripts/installer/auto-install.sh. Add new cross-script config here instead of duplicating it per-script.
  • scripts/bump-nixpkgs-release.sh — bumps flake.nix's nixpkgs.url/ home-manager.url in place. Exists because flake input URLs can't reference variables.nix (confirmed empirically — nix flake metadata errors on it), so this is the closest equivalent to a single source of truth for the tracked release.

sync-host-keys.sh, create-proxmox-resource.sh, and rotate-admin-key.sh genuinely mutate real state when run for real (not --dry-run): real secrets/*.yaml recipients, real Proxmox VMs/ containers, real revocation of decrypt access. They require the operator's own SSH/sops access, which an agent session doesn't have — but don't suggest running any of them non-dry-run without the operator's explicit go-ahead even if it becomes technically reachable. backup-admin-key.sh only writes a key copy to a path the operator gives it — lower-stakes than the others, but it still handles a real private key, so treat its destination path choice as the operator's call too.

Architecture

flake.nix is the single entry point. It generates one nixosConfigurations.<platform>-<buildtype> attribute per target via the mkTarget function, composed from:

nixosSystem {
  modules = [
    disko.nixosModules.disko
    sops-nix.nixosModules.sops
    ./modules/common/configuration.nix
    ./modules/platforms/${platform}.nix        # what it runs on
    ./modules/build-types/${buildType}.nix     # what it's for
    hostPath                                   # hosts/<name>/host.nix — per-machine identity
    home-manager.nixosModules.home-manager { ... }
  ] ++ (client-only modules, for every buildType except "nix-cache" itself)
}

Platforms: linode, proxmox, lxc, baremetal. Build types: minimal, nix-cache, server, docker, gui, pxe-boot, tailscale-exit-node, tor-relay. Not every combination is built — e.g. pxe-boot has no linode variant (PXE/DHCP/TFTP need LAN L2 adjacency a Linode VPS doesn't have), tor-relay currently only exists as lxc-tor-relay, and baremetal currently only exists as baremetal-gui (the real gui-host hardware — see hosts/nixos/host.nix and modules/platforms/baremetal.nix). Treat flake.nix's generatedTargets as the source of truth for which hosts exist — README.md, AGENTS.md, docs/flake-lock-automation.md, and the CI eval workflows (.github/workflows/check-nixos.yml, .gitea/workflows/check-nixos.yml) list hosts by hand (or, for the CI workflows, evaluate the flake dynamically) and can drift from it, so re-check them against flake.nix when adding or removing a host.

Composition pattern

  • hosts/<name>/host.nix — per-machine identity only: hostname, hostId, per-machine secrets, system.stateVersion. These files carry no imports of their own — all shared behavior comes from the platform/build-type modules composed in flake.nix, not from the host file.
  • modules/platforms/{linode,proxmox,lxc,baremetal}.nix — platform-specific config: boot method, guest tooling, and the hardware config, imported directly by the platform module itself — not wired in from flake.nix. VM platforms use ../hardware-configuration/vm/{proxmox,linode}.nix; baremetal.nix uses ../hardware-configuration/baremetal.nix (adapted from a real nixos-generate-config run on the actual hardware, not a vm/ file, since it isn't a VM) plus hardware.enableRedistributableFirmware = true for real wifi/GPU/microcode firmware that VMs never needed. lxc.nix has no hardware-configuration counterpart since containers share the host kernel; instead it imports nixpkgs' own virtualisation/proxmox-lxc.nix, which gives every lxc-* host a config.system.build.tarball output — a plain rootfs tarball, used as a pct create ... vztmpl CT template (not pct restore, which expects vzdump backup-archive metadata this doesn't have), no install step — see docs/auto-installer.md.
  • modules/build-types/*.nix — what a system is for: minimal/server/docker/gui/pxe-boot/nix-cache/tailscale-exit-node/tor-relay.
  • modules/common/configuration.nix — base NixOS config imported by every host: locale, users, nix settings, git.
  • modules/common/home.nix / hosts/nixos/home.nix — Home Manager config for the nixos user; the nixos workstation (gui build type) has its own, other hosts share modules/common/home.nix.
  • modules/disko/proxmox.nix — declarative disk layout (GPT: ESP + swap + ext4 root) via disko, used by all Proxmox-VM hosts (proxmox-*, not lxc-*). Also carries imageSize/imageName, letting every proxmox-* host be built as a standalone, qm importdisk-ready .raw image with no install step — see docs/proxmox-images.md.
  • modules/disko/linode.nixlinode-*'s disko config, deliberately different in kind from the Proxmox one: Linode provisions and sizes /dev/sda//dev/sdb itself as whole, unpartitioned devices before the OS boots, so this declares them with destroy = false (disko never wipes them) and a bare filesystem/swap content type instead of a partition table — idempotent against an already-provisioned disk, never destructive.
  • modules/disko/baremetal.nixbaremetal-gui's disko config: a ZFS RAID0 (striped, no redundancy — disko's zpool mode defaults to "", which is a plain stripe rather than "mirror"/"raidz") root pool across two disks, ESP + systemd-boot on the first. Device paths (vars.guiRootDisk1/guiRootDisk2) are placeholders — fill in stable /dev/disk/by-id/... paths before running disko for real. modules/platforms/baremetal.nix also imports modules/services/zfs/enable-service.nix for this (the zfs_unstable package, autoScrub/autoSnapshot/trim) — the only other importer today is server's NFS data pool, an unrelated non-root ZFS use.
  • modules/boot/efi.nix — systemd-boot + EFI vars, paired with the disko module.
  • modules/installer/ — the auto-installer environment (ISO, also served as PXE netboot): common.nix (shared config + the generated auto-install.sh), iso.nix, host-keys.nix (optionally bakes host-keys/ into the image under --impure). See docs/auto-installer.md.
  • modules/pxe-boot/stage-installer-artifacts.nix — builds the installer's netboot image and stages it on the pxe-boot host so its iPXE menu can chain straight to it. See docs/pxe-boot.md.
  • modules/nix-cache/{client,server,remote-builder-client}.nix — binary cache substituter + SSH remote-builder wiring; see docs/nix-cache.md for the full design (per-host local stores, no shared /nix/store, and how the nixremote signing/SSH keys fit together).
  • modules/beszel/enable-agent.nix — enables beszel-agent, sets HUB_URL, fixes the upstream StateDirectory bug, and wires the universal beszel-token sops secret (from secrets/common.yaml) into the agent's environmentFile; see docs/beszel.md for the full setup guide.
  • modules/tailscale/, modules/docker/, modules/networking/, modules/traefik/, modules/tor/, modules/services/* — single-purpose, single-host feature modules (e.g. docker/enable-service.nix, services/zfs/enable-service.nix). Grep modules/build-types/*.nix for each build type's imports list to see which modules apply where.

New host = new hosts/<name>/host.nix + a matching mkTarget { platform; buildType; hostPath; } entry added to flake.nix's generatedTargets, composed from existing modules/* pieces rather than duplicating config.

Other docs worth reading before touching these areas

  • docs/nix-cache.md — nix-cache binary cache/remote-builder design and key handling.
  • docs/pxe-boot.md — the pxe-boot host's iPXE/TFTP/HTTP boot chain and directory layout under /srv/pxe.
  • docs/auto-installer.md — the installer environment (ISO/netboot/Proxmox LXC), host-keys/ and the sops-nix pre-seeding problem it solves, and why lxc-* hosts are deliberately excluded from its menu.
  • docs/proxmox-images.md — building proxmox-* hosts as standalone .raw disk images (disko's image builder) instead of installing, and deploying the result to Proxmox.
  • docs/flake-lock-automation.md — how flake.lock updates flow through CI (scheduled nix flake update PR + host-eval-on-PR workflow) and why hosts should track the committed lock file rather than nixos-rebuild --upgrade-all.