Commit Graph
100 Commits
Author SHA1 Message Date
beatzaplentyandClaude Sonnet 4.6 c7bbf88dce fix(tailscale-router): stop dnsmasq from intercepting host DNS queries
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m30s
NixOS's dnsmasq module defaults resolveLocalQueries to true, which adds
127.0.0.1 to networking.nameservers and binds dnsmasq to listen-address=127.0.0.1.
This made the host route all its own DNS through dnsmasq, which had
no-resolv=true and no upstream for anything outside the tailnet domain —
so every non-tailscale DNS query from the host itself (including SSSD
resolving the IPA server FQDN after the IPA client module was added) failed.

Setting resolveLocalQueries=false limits dnsmasq to its intended role: a
forwarding proxy reachable on the LAN interface for IPA's conditional
forwarder. The host uses domainControllerIp directly (already set in
networking.nameservers in host.nix).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-28 09:26:50 +10:00
beatzaplenty d57145b31e add tailscale-router to domain
Check NixOS configurations / eval-hosts (push) Failing after 9m44s
2026-07-28 09:17:29 +10:00
beatzaplenty 1cbe80c0ca secrets: add IPA keytab for tailscale-router 2026-07-28 09:15:21 +10:00
beatzaplenty 01ee261cf8 Merge pull request 'fix(ipa): stream keytab via sudo cat instead of scp' (#78) from worktree-ipa-client-module into main
Check NixOS configurations / eval-hosts (push) Successful in 10m22s
Reviewed-on: #78
2026-07-27 23:14:41 +00:00
beatzaplentyandClaude Sonnet 4.6 f6f30c675f fix(ipa): stream keytab via sudo cat instead of scp
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m25s
ipa-getkeytab runs as root via sudo so the temp file is root-owned;
scp as wayne gets Permission denied. Pipe through `sudo cat` over SSH
instead, which reads as root but writes locally as the invoking user.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-28 09:08:19 +10:00
beatzaplenty 60b80cbd96 Merge pull request 'fix(ipa): SSH as wayne with sudo instead of root on domain controller' (#77) from worktree-ipa-client-module into main
Check NixOS configurations / eval-hosts (push) Successful in 10m21s
Reviewed-on: #77
2026-07-27 23:06:49 +00:00
beatzaplentyandClaude Sonnet 4.6 27a8c7fad9 fix(ipa): SSH as wayne with sudo instead of root on domain controller
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m23s
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-28 09:02:32 +10:00
beatzaplenty d9cee0a674 Merge pull request 'feat(ipa): add create-nixos-ipa-host-account script' (#76) from worktree-ipa-client-module into main
Check NixOS configurations / eval-hosts (push) Successful in 10m22s
Reviewed-on: #76
2026-07-27 22:58:46 +00:00
beatzaplentyandClaude Sonnet 4.6 6c1891812e feat(ipa): add create-nixos-ipa-host-account script
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m24s
Single command to enroll a NixOS host in FreeIPA and produce a
sops-encrypted keytab at secrets/<hostname>.keytab:
  - Adds the .sops.yaml creation rule automatically (with all registered
    platform-variant age keys as recipients)
  - SSHes to the domain controller to run ipa host-add + ipa-getkeytab
  - Refreshes the admin Kerberos ticket via `ssh -t ... kinit admin` if
    missing or expired, so no manual kinit step is needed
  - SCPs the keytab and encrypts it in-place with sops (file must be at
    secrets/<hostname>.keytab before encryption so the path-based creation
    rule matches — the common failure point when doing this manually)

Also adds HOME_DOMAIN and IPA_SERVER to scripts/env.sh, matching
variables.nix's homeDomain/ipaServer (same manual-sync pattern as
NIX_CACHE_HOST/LAN_DOMAIN).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-28 08:57:55 +10:00
beatzaplenty 59316c982e Merge branch 'main' of https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos
Check NixOS configurations / eval-hosts (push) Successful in 10m27s
2026-07-28 07:26:40 +10:00
beatzaplenty 9eca3bd719 Merge pull request 'Worktree ipa client module' (#75) from worktree-ipa-client-module into main
Check NixOS configurations / eval-hosts (push) Failing after 9m45s
Reviewed-on: #75
2026-07-27 21:25:57 +00:00
beatzaplenty f2f0fcf756 certs: add FreeIPA CA certificate
Check NixOS configurations / eval-hosts (pull_request) Failing after 9m45s
2026-07-28 07:25:28 +10:00
beatzaplenty b4fb9c25f2 secrets(nix-cache): add sops-encrypted IPA host keytab 2026-07-28 07:25:25 +10:00
beatzaplentyandClaude Sonnet 4.6 8955840f0a fix(ipa): use in-place sops encryption in module docs
sops matches creation rules against the input file path, so encrypting
/tmp/<host>.keytab directly with stdout redirect fails to find the rule.
Copy to secrets/ first, then use -i to encrypt in-place.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-28 07:19:53 +10:00
beatzaplentyandClaude Sonnet 4.6 a4b49c9909 fix(ipa): correct ipaServer to domain-controller.sweet.home
ipa-getkeytab confirmed ipa.sweet.home doesn't respond to LDAP;
domain-controller.sweet.home is the actual IPA server FQDN.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-28 07:16:53 +10:00
beatzaplenty c0936a10e7 Merge pull request 'feat(ipa): add reusable declarative FreeIPA client module' (#74) from worktree-ipa-client-module into main
Check NixOS configurations / eval-hosts (push) Failing after 9m45s
Reviewed-on: #74
2026-07-27 21:06:28 +00:00
beatzaplentyandClaude Sonnet 4.6 f4bbd6331d feat(ipa): add reusable declarative FreeIPA client module
Check NixOS configurations / eval-hosts (pull_request) Failing after 10m0s
Adds modules/ipa/client.nix — a parameterized module that joins a NixOS host
to the sweet.home FreeIPA domain without ipa-client-install. It configures
security.ipa (SSSD, Kerberos, PAM, NSSwitch) and places a pre-provisioned host
keytab via sops-nix binary secret so enrollment is fully reproducible from the
flake.

- variables.nix: adds ipaServer (FQDN of the FreeIPA KDC; security.ipa.server
  requires a hostname, not an IP, for Kerberos/TLS)
- certs/ipa-ca.crt: placeholder for the IPA CA public certificate (operator
  replaces with: curl http://<ipa-server>/ipa/config/ca.crt)
- secrets/nix-cache.keytab: placeholder binary sops file (operator replaces
  with the encrypted keytab after ipa host-add + ipa-getkeytab)
- .sops.yaml: adds creation rule for secrets/nix-cache.keytab (same recipients
  as secrets/nix-cache.yaml)
- hosts/nix-cache/host.nix: imports the IPA client module; adds
  networking.domain so the host's FQDN resolves correctly

Module header documents the three operator steps needed per host before deploy.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 23:31:21 +10:00
beatzaplenty b99ba87cf6 Merge branch 'main' of https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos 2026-07-27 23:05:35 +10:00
beatzaplenty 2128353f9f Merge pull request 'Fix/static ips' (#73) from fix/static-ips into main
Check NixOS configurations / eval-hosts (push) Failing after 42m27s
Reviewed-on: #73
2026-07-27 12:06:33 +00:00
beatzaplentyandClaude Sonnet 4.6 7b4ce0ab3d fix(server): prevent zfs-init-tank from wiping pool on udev race
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m46s
zpool create -f was called if `zpool import -d /dev/disk/by-id` failed,
which could happen due to a race with systemd-udev-settle. The disk
would then be visible by the time zpool create ran, silently destroying
all data on an otherwise-intact pool.

Fix: locate the data disk first, retry the import directly against it
as a fallback, then check zdb -l for existing ZFS label metadata before
concluding the disk is blank. Remove -f so zpool create refuses rather
than overwrites if a pool is present.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULXzafSDwGhmFGnn3LtDSQ
2026-07-27 21:52:39 +10:00
beatzaplentyandClaude Sonnet 4.6 3123565011 fix(tailscale-router): scope ts.net forwarder to tailnet subdomain
IPA refuses to create a forward zone for ts.net because it's a real
public TLD with DNSimple nameservers. The forward zone must use the
tailnet-specific subdomain (vars.tailnetDomain, e.g. tail13f623.ts.net)
instead. Update dnsmasq server selector and comments to match.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULXzafSDwGhmFGnn3LtDSQ
2026-07-27 19:23:46 +10:00
beatzaplentyandClaude Sonnet 4.6 36f5ebdf86 feat(tailscale-router): serve ts.net DNS forward zone for LAN hosts
FreeIPA (the new authoritative DNS) cannot reach 100.100.100.100
(Tailscale's internal MagicDNS resolver) directly because the DC is not
a Tailscale node. The tailscale-router IS a Tailscale node and can
reach 100.100.100.100 via tailscale0, so it now runs a dnsmasq
instance on its LAN interface that forwards all ts.net queries to
Tailscale's resolver.

After deploying this host, configure FreeIPA with:
  kinit admin
  ipa dnsforwardzone-add ts.net \
    --forwarder=192.168.2.222 \
    --forward-policy=only

This replaces Pi-hole's conditional forwarder for ts.net and restores
resolution of Tailscale MagicDNS names (e.g. raspberrypi.tail13f623.ts.net)
for all LAN hosts using FreeIPA as their DNS server.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULXzafSDwGhmFGnn3LtDSQ
2026-07-27 19:20:12 +10:00
beatzaplenty bba054db85 Merge branch 'main' of https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos
# Conflicts:
#	variables.nix
2026-07-27 18:29:07 +10:00
beatzaplenty b63a529a1d Merge branch 'fix/static-ips' 2026-07-27 18:25:37 +10:00
beatzaplentyandClaude Sonnet 4.6 ee96713a50 feat(networking): declare static IPs for all NixOS-managed hosts in flake
Adds static IP configuration to every NixOS host in the flake that
has a fixed LAN address, and centralises all network primitives
(IPs, gateway, prefix length, interface names) in variables.nix so
there is one place to update if any of them change.

variables.nix additions:
- lanGateway / lanPrefixLength — LAN gateway and /24 prefix, replacing
  every hardcoded 192.168.2.254 / 24 across host files
- lxcLanInterface / vmLanInterface / vmStorageInterface — NIC names for
  LXC containers (eth0), Proxmox VMs (ens18), and the HA storage NIC
  (ens19), used as attribute keys so changing the name is a one-line edit
- haStoragePrefixLength — /29 for the storage subnet, mirrors haStorageCidr
- Per-host IP variables: nixCacheIp (.224), tailscaleRouterIp (.222),
  torRelayIp (.221), serverIp (.226), dockerIp (.225)

host.nix changes:
- tailscale-router, tor-relay, nix-cache, pxe-boot: useDHCP = false,
  static address on eth0 (lxcLanInterface), struct-form defaultGateway
  (required when using systemd-networkd which LXC containers use)
- server, docker: useDHCP = false, static address on ens18 (vmLanInterface),
  struct-form defaultGateway (works for both scripted networking and networkd)
- ha-server-1, ha-server-2: replace hardcoded 192.168.2.254 / 24 / 29
  with the new variables; no functional change for these hosts

modules/build-types/pxe-boot.nix:
- Domain-controller kickstart template: replace hardcoded 192.168.2.138
  and 192.168.2.254 with vars.domainControllerIp / vars.lanGateway /
  vars.lanPrefixLength / vars.homeDomain so the template stays correct
  if the DC IP or domain is ever changed again

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ULXzafSDwGhmFGnn3LtDSQ
2026-07-27 18:14:31 +10:00
beatzaplenty 1ece0c75d9 Merge pull request 'chore(pxe-boot): remove log-dhcp debug flag now that PXE boot is confirmed working' (#86) from fix/pxe-dnsmasq-port-conflict into main
Reviewed-on: #86
2026-07-27 05:52:18 +00:00
beatzaplentyandClaude Sonnet 4.6 548c6f5041 chore(pxe-boot): remove log-dhcp debug flag now that PXE boot is confirmed working
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 15:50:19 +10:00
beatzaplenty bae4c8171f Merge pull request 'fix(pxe-boot): use current pxe-boot IP (.247) in dnsmasq until Stage 5 renumber' (#85) from fix/pxe-dnsmasq-port-conflict into main
Reviewed-on: #85
2026-07-27 05:41:38 +00:00
beatzaplentyandClaude Sonnet 4.6 3748c86049 fix(pxe-boot): use current pxe-boot IP (.247) in dnsmasq until Stage 5 renumber
vars.pxeServerIp was already set to the post-renumber target (.223) but the
pxe-boot CT is still at .247, so dnsmasq was advertising .223 as the TFTP server
and the VM couldn't reach it.  Update to .247 so PXE boot works now; the comment
reminds us to flip it back to .223 when Stage 5 step 6 renumbers the CT.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 15:40:54 +10:00
beatzaplenty 9dd969cf4c Merge pull request 'fix(pxe-boot): open UDP 4011 for PXE boot service discovery' (#84) from fix/pxe-dnsmasq-port-conflict into main
Reviewed-on: #84
2026-07-27 05:22:46 +00:00
beatzaplentyandClaude Sonnet 4.6 7e4b2d33fb fix(pxe-boot): open UDP 4011 for PXE boot service discovery
After the proxy DHCP handshake, the UEFI client queries the PXE boot
server on UDP port 4011 to get the actual boot filename. This port was
not open in the NixOS firewall, so all four PXE discovery attempts from
the test VM silently timed out. dnsmasq was bound to 0.0.0.0:4011 but
never received the packets.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 15:12:12 +10:00
beatzaplenty 288d50fd33 Merge pull request 'Fix/pxe dnsmasq port conflict' (#83) from fix/pxe-dnsmasq-port-conflict into main
Reviewed-on: #83
2026-07-27 05:05:42 +00:00
beatzaplentyandClaude Sonnet 4.6 f0e76f8aff fix(pxe-boot): use pxe-service for proxy DHCP initial boot offers
dnsmasq 2.93 does not send proxy DHCP offers from dhcp-boot alone in proxy
mode — pxe-service is required to trigger proxy offers for the initial boot.

Replace the dhcp-boot-based arch detection with pxe-service entries for
arch 0 (BIOS), 7 (EFI BC / OVMF), and 9 (EFI x86-64), then use dhcp-boot
only for the iPXE second-boot HTTP chain (tag:ipxe clients). Add log-dhcp
for post-deploy verification.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 15:03:01 +10:00
beatzaplenty e80d195284 Merge branch 'fix/pxe-dnsmasq-port-conflict' 2026-07-27 14:32:00 +10:00
beatzaplentyandClaude Sonnet 4.6 c770feebc9 fix(pxe-boot): disable dnsmasq DNS listener to avoid port 53 conflict
dnsmasq binds port 53 by default even when only proxy DHCP is needed.
systemd-resolved already owns port 53 on the pxe-boot container, so
dnsmasq crashes at startup. Setting port=0 disables the DNS listener
entirely — pxe-boot only needs the proxy DHCP path for iPXE chainloading.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 14:29:30 +10:00
beatzaplenty 7fd6d558d5 updated dnsmasq service 2026-07-27 14:07:28 +10:00
beatzaplentyandClaude Sonnet 4.6 58c40292e2 fix(pxe-boot): move dnsmasq out of systemd.services into services.dnsmasq
dnsmasq config was placed inside systemd.services{} block, making Nix
interpret it as systemd.services.dnsmasq.settings (which doesn't exist).
Move to top-level services.dnsmasq and add UDP 67 to firewall for DHCP
proxy mode.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 14:02:41 +10:00
beatzaplenty 94842875d0 Merge branch 'main' of https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos into worktree-network-cutover-plan 2026-07-27 14:01:24 +10:00
beatzaplenty cfa36b97fc added dnsmasq to pxe-boot 2026-07-27 13:44:29 +10:00
beatzaplentyandClaude Sonnet 4.6 4444398cac fix(network): wire correct IPs throughout and add cutover plan
variables.nix:
- HA server LAN IPs: .200/.201/.202 → .228/.227/.229 (from ip-addressing.md)
- Add haServer1StorageIp (.228), haServer2StorageIp (.227) for 192.168.4.0/29
- Add haStorageCidr for firewall rules

ha-server host.nix (both nodes):
- Add ens19 interface on storage subnet (/29)
- Fix defaultGateway: 192.168.2.1 → 192.168.2.254
- Fix nameservers: 192.168.2.1/8.8.8.8 → domainControllerIp (.253)

cluster-config.nix:
- DRBD replication addresses: LAN IPs → storage IPs (keep replication off LAN)
- Corosync ring_addrs: LAN IPs → storage IPs
- Firewall: add haStorageCidr to allowed sources

docs/network-cutover.md: step-by-step cutover plan with rollback at every stage

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 12:32:54 +10:00
beatzaplenty f80378f92f Merge pull request 'docs(network): establish IP addressing scheme and update infra IPs' (#82) from worktree-network-ip-scheme into main
Reviewed-on: #82
2026-07-27 02:23:01 +00:00
beatzaplentyandClaude Sonnet 4.6 cd4997f429 docs(network): establish IP addressing scheme and update infra IPs
Defines the new structured 192.168.2.0/24 layout:
- .10–.59   client DHCP (router-assigned, DNS → .253)
- .220–.229 virtual nodes (VMs / LXC containers)
- .230–.239 expansion buffer
- .240–.249 physical nodes (pve1 at .245, PBS at .244)
- .250–.253 network services (router .254, FreeIPA/DC .253)

Storage network 192.168.4.0/29 defined for HA DRBD replication
(internal vmbr1 bridge, no uplink). Host octet matches LAN throughout.

Updates variables.nix: pxeServerIp .247→.223, pbsIp .108→.244,
adds domainControllerIp .253. Updates pxe-boot.md IP references.
Full migration before/after table in docs/ip-addressing.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 12:19:45 +10:00
beatzaplenty 6ce4784376 Merge pull request 'Worktree ha file server test' (#81) from worktree-ha-file-server-test into main
Reviewed-on: #81
2026-07-27 02:18:06 +00:00
beatzaplentyandClaude Sonnet 4.6 5856d45575 feat(ha): wire sops secrets and disable NetworkManager for HA servers
- cluster-config.nix: add corosync_authkey sops binary secret
  (/etc/corosync/authkey, mode 0400) and force-disable NetworkManager
  (common config enables it; HA nodes need stable static IP networking)
- hosts/ha-server-{1,2}/host.nix: add host-token.nix import for
  sops-managed beszel-token; add KEY placeholder for beszel hub pairing
- .sops.yaml: add creation rules for secrets/ha-server-{1,2}.yaml and
  secrets/ha-corosync-authkey (admin-only until sync-host-keys.sh runs)
- secrets/ha-server-{1,2}.yaml, secrets/ha-corosync-authkey: stub files
  so eval passes before real secrets are provisioned

Bootstrap order (post-merge):
  1. bash scripts/secrets/sync-host-keys.sh proxmox-ha-server-1
  2. bash scripts/secrets/sync-host-keys.sh proxmox-ha-server-2
  3. sops updatekeys secrets/common.yaml  (grants HA nodes common secrets)
  4. sops secrets/ha-server-{1,2}.yaml   (set beszel-token values)
  5. On node1: corosync-keygen; sops -e --input-type binary
     /etc/corosync/authkey > secrets/ha-corosync-authkey; git add/commit

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HaH1cSGvhogRP5ExoF6nD8
2026-07-27 11:45:36 +10:00
beatzaplentyandClaude Sonnet 4.6 e3498b1087 feat(ha): promote HA file server to production flake targets
Adds proxmox-ha-server-1 and proxmox-ha-server-2 as real mkTarget entries
alongside the existing proxmox-server, backed by a new ha-server build type.

New modules
  modules/ha/cluster-config.nix — DRBD resource + corosync nodelist sourced
    from vars (haServer1Host/Ip, haServer2Host/Ip); resource-only fencing for
    production STONITH; HA port firewall rules for DRBD, iSCSI, Corosync, pcsd
  modules/build-types/ha-server.nix — imports pacemaker-stack + iscsi-target
    + cluster-config + beszel; NFS exports from vars.haStorageRoot (XFS-over-DRBD
    mount); nfs-server.service.wantedBy force-cleared so Pacemaker controls
    start/stop on the Active node only

New hosts
  hosts/ha-server-{1,2}/host.nix — static IP from vars, unique hostId; sops
    secrets (beszel, corosync authkey) are TODOs pending sync-host-keys.sh

variables.nix
  haServer1/2Host, haServer1/2Ip, haServerVip, haStorageRoot, haIscsiIqn
  ports.haServerDrbd/Iscsi/Corosync{1,2,Crypto}/PacemakerRemoted/Pcsd

scripts/ha/ (migrated + updated from test-lab/ha/)
  cluster-init.sh — generates corosync authkey, initialises DRBD/XFS/iSCSI,
    creates NFS dataset dirs, configures Pacemaker with DRBD + XFS + iSCSI
    + nfs-server + VIP; STONITH disabled initially (enable separately)
  cluster-enable-stonith.sh — enables fence_pve_ssh STONITH after key deploy
  fence-pve-ssh.py — Proxmox SSH fence agent (node names updated to ha-server-1/2)
  acceptance-tests.sh — T1–T7 production acceptance tests

test-lab/ha/ removed — all Nix config moved to modules/ha/ and
  modules/build-types/; scripts moved to scripts/ha/

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HaH1cSGvhogRP5ExoF6nD8
2026-07-27 11:26:37 +10:00
beatzaplentyandClaude Sonnet 4.6 724d9a45af feat(ha): add NixOS modules for DRBD+XFS+LIO+Corosync+Pacemaker HA stack
All 7 acceptance tests pass on live NixOS 25.11 VMs (VMIDs 200/201 on
pve1).  Failover completes in ~5 s with data integrity verified.

modules/ha/pacemaker-stack.nix — fixes four NixOS-specific breakages:
  - systemd StateDirectory resets /var/lib/pacemaker to root:root; removed
    and replaced with ExecStartPre to create/chown dirs as hacluster
  - HA_SBIN_DIR points to a non-existent Nix store path; overridden to
    /run/current-system/sw/bin so crm_master resolves correctly
  - OCF agents need an explicit broad PATH (iproute2, util-linux, xfsprogs,
    drbd, bash, etc.) — NixOS services have no implicit PATH
  - FUSER=true bypasses the psmisc fuser check_binary call in the
    Filesystem OCF agent (psmisc not installed on minimal hosts)

modules/ha/iscsi-target.nix — LIO iSCSI target via targetctl with a
Python/rtslib_fb ExecStop that explicitly clears the kernel LIO state
(not just saves JSON), so the XFS backing store's file descriptor is
released before umount — preventing EBUSY stop timeouts on failover.
Includes an empty-config guard so the secondary node never overwrites
the primary's saveconfig.json with an empty one.

test-lab/ha/common.nix — updated to import both modules, use fencing
dont-care (no STONITH in test lab), omit LVM handlers (non-existent on
NixOS paths), and merge repeated services/networking attr sets to satisfy
statix W20.  test-lab/ha/acceptance-tests.sh — final v4 with crm_standby
fix (pacemaker 3.x API).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HaH1cSGvhogRP5ExoF6nD8
2026-07-27 09:13:55 +10:00
beatzaplenty 109429c7da Merge pull request 'feat(pxe-boot): add FreeIPA Server (Rocky Linux 9) iPXE menu entry' (#80) from worktree-freeipa-pxe into main
Reviewed-on: #80
2026-07-26 21:36:58 +00:00
beatzaplentyandClaude Sonnet 4.6 40856b2e5e feat(pxe-boot): add FreeIPA Server (Rocky Linux 9) iPXE menu entry
Adds an unattended install option to the PXE boot menu that installs
Rocky Linux 9 and configures FreeIPA on the domain-controller.sweet.home
host without any operator interaction after selecting the menu entry.

How it works:
- fetch-rocky-pxeboot.service downloads the Rocky 9 Anaconda pxeboot
  kernel and initrd from the Rocky mirror on first pxe-boot deploy
  (idempotent, same pattern as fetch-debian-netboot)
- rocky-freeipa.ipxe boots Anaconda with inst.ks pointing at the
  hosted Kickstart and net.ifnames=0 biosdevname=0 for stable eth0
- rocky-freeipa.ks (generated, includes vars.adminSshKey) performs:
  - Minimal Rocky 9 install with ipa-server + ipa-server-dns
  - Static IP 192.168.2.138 via NM connection file written in %post
  - /etc/hosts fixed for FreeIPA FQDN requirement
  - Random DM + admin passwords generated and saved to
    /root/ipa-credentials.txt (chmod 600, never hardcoded)
  - freeipa-first-boot.service oneshot enabled to run
    ipa-server-install on the first real boot (~20 min)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Jbvxx4xbHVcx1NkK3vtmK
2026-07-27 07:35:13 +10:00
beatzaplentyandClaude Sonnet 4.6 23634134f0 fix(ha-test): remove Debian LVM handlers and fix corosync authkey length
- Remove before/after-resync-target LVM handlers (snapshot-resync-target-lvm.sh
  doesn't exist in NixOS; exit code 127 caused DRBD to drop connections on sync)
- Remove split-brain handler pointing to /usr/lib/drbd/ (Debian path)
- Fix testAuthKey from 126 to 128 bytes (corosync minimum is 1024 bits)
- Fix cluster-init.sh quorum check: corosync-quorumtool has no -q flag;
  use -s | grep 'Quorate: Yes' instead

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HaH1cSGvhogRP5ExoF6nD8
2026-07-27 06:40:57 +10:00
beatzaplentyandClaude Sonnet 4.6 34c55f27ca test-lab/ha: nixpkgs-fmt formatting
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 06:18:42 +10:00
beatzaplentyandClaude Sonnet 4.6 dc36a47ac9 test-lab/ha: add session key + enable qemu-guest-agent
Add nixos@nixos session key so the Claude Code session can SSH into
test VMs directly.  Also enable services.qemuGuest.enable so
qm guest exec works as a fallback for key injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 06:06:53 +10:00
beatzaplentyandClaude Sonnet 4.6 750121e9dd test-lab: add two-node HA file-server test cluster config
Disposable test VMs (ha-test-node1 / ha-test-node2, VMIDs 200/201 on pve1)
to evaluate whether the DRBD + XFS + LIO + Corosync + Pacemaker stack
runs correctly on NixOS before deciding NixOS vs Debian for production.

Includes:
- test-lab/ha/disko.nix: 20G boot disk layout (smaller than production)
- test-lab/ha/common.nix: shared HA stack (drbd, corosync, pacemaker,
  targetcli-fb, xfsprogs), OCF PATH workaround for nixpkgs#207891
- test-lab/ha/node1.nix / node2.nix: per-node hostname + static IP
- test-lab/ha/fence-pve-ssh.py: Proxmox SSH fence agent for STONITH
- test-lab/ha/cluster-init.sh: one-shot cluster bootstrap script
- test-lab/ha/cluster-enable-stonith.sh: enables STONITH post-key-deploy
- flake.nix: adds ha-test-node1 / ha-test-node2 nixosConfigurations
  (bypasses mkTarget / clan-core / sops-nix — test-only)

These VMs must be destroyed once acceptance testing is complete.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 05:46:04 +10:00
beatzaplenty 479444d26a Merge pull request 'fix(sops): hard-fail on missing admin key and expand literal ~ in key path' (#79) from worktree-structured-nibbling-nova into main
Reviewed-on: #79
2026-07-26 19:43:57 +00:00
beatzaplentyandClaude Sonnet 4.6 8c19ee9d72 fix(sops): hard-fail on missing admin key and expand literal ~ in key path
SOPS_AGE_KEY_FILE was set in hosts/nixos/home.nix sessionVariables with a
literal ~ that Home Manager injects as-is into the environment.  In bash,
tilde expansion does not happen inside double-quoted variable references, so
DEFAULT_SOPS_AGE_KEY_FILE resolved to ~/... literally and the -s file-existence
check in ensure_admin_decrypt_key silently failed.  The script then generated
a brand-new age key (to ~/... relative to the repo root) while the real admin
key at ~/.config/sops/age/keys.txt went untouched -- making it appear the key
was lost when it was actually still intact.

Fix the home.nix root cause by using config.home.homeDirectory so the path
is fully resolved.  Add tilde expansion in ensure_admin_decrypt_key as a
belt-and-suspenders guard for any caller whose environment has the same issue.

Also replace the auto-generate-a-new-key fallback with a hard failure: auto-
generating a new admin key is never useful (it cannot decrypt existing secrets)
and created serious confusion about whether the original key was lost.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 05:42:32 +10:00
beatzaplenty 2514d3bc89 Merge pull request 'fix(lxc-pxe-boot): make privileged so NFS mounts work' (#78) from worktree-debian-pxe into main
Reviewed-on: #78
2026-07-26 19:29:25 +00:00
beatzaplentyandClaude Sonnet 4.6 48d6d6f7a2 fix(lxc): auto-derive privileged from NFS fileSystems, not hostname list
Replace the hardcoded hostname check (docker, pxe-boot) with a check
on config.fileSystems: any lxc-* host whose NixOS config declares an
NFS fileSystem entry is automatically made privileged. The script
already reads proxmoxLXC.privileged dynamically via
flake_target_lxc_privileged, so no logic change is needed there —
only the comment is updated to describe the new derivation.

Result: lxc-docker and lxc-pxe-boot (the two with NFS mounts) evaluate
as privileged=true; lxc-nix-cache, lxc-minimal, lxc-server,
lxc-tailscale-router, lxc-tor-relay evaluate as privileged=false.
Any future lxc-* host that declares an NFS mount gets the correct
privilege level for free without a separate manual edit.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 05:27:56 +10:00
beatzaplentyandClaude Sonnet 4.6 cf0d62696f fix(lxc-pxe-boot): make privileged so NFS mounts work
The kernel's NFS client (FS_USERNS_MOUNT not set) rejects NFS mounts
from inside any unprivileged container's user namespace with EPERM —
AppArmor's mount=nfs feature only whitelists the AppArmor layer; the
VFS-level rejection happens before AppArmor is consulted.

lxc-pxe-boot mounts server.sweet.home:/tank/pxe-boot/images at
/mnt/pxe-images so nginx can serve large ISOs without filling the
container's root disk. Same pattern as lxc-docker (already privileged
for the same reason since it also mounts several NFS shares).

Operator action required: VMID 103 must be recreated as a privileged
container (the UID mapping on disk differs between privileged and
unprivileged; changing it in-place with pct set is unsafe). Rebuild
the tarball and use create-proxmox-resource.sh to replace it.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 05:23:22 +10:00
beatzaplenty b2a3d5fbdd Merge pull request 'docs(pxe-boot): fix NFS path, note nesting=1 requirement for LXC' (#77) from worktree-debian-pxe into main
Reviewed-on: #77
2026-07-26 19:18:48 +00:00
beatzaplentyandClaude Sonnet 4.6 86e55ff954 docs(pxe-boot): fix NFS path, clarify NFSv3/NFSv4 split, document nesting=1 requirement
- Correct the NFS path from /tank/proxmox/pxe-images to /tank/pxe-boot/images
  (matches variables.nix's proxmoxPxeImages.subpath)
- Clarify that LXC uses NFSv3+nolock while VM uses NFSv4.2+automount
- Add explicit note that lxc-pxe-boot needs features: nesting=1,mount=nfs and
  why: nesting=1 is required by systemd 260+ for userns/credential isolation
  (AppArmor denies userns_create without it), mount=nfs for NFSv3 access.
  pct set replaces the whole features string — include both or the container
  will fail to boot.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 05:13:21 +10:00
beatzaplenty d3d6382360 Merge pull request 'fix: wire vars.nfsShares.options and correct stale share key names' (#76) from worktree-tidy-giggling-pizza into main
Reviewed-on: #76
2026-07-26 18:45:53 +00:00
beatzaplentyandClaude Sonnet 4.6 b8d21d78d9 fix: wire vars.nfsShares.options and correct stale share key names
- Filter non-attrset values from nfsShares in server.nix so the
  poolDatasets loop skips the new `options` string entry
- Fix typo proxomoxLxcImages → proxmoxLxcImages in server.nix exports
- Rename proxmoxPxeImages → pxebootImages in mount-pxe-images.nix and
  pxe-boot.nix to match the actual key in variables.nix

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 04:45:16 +10:00
beatzaplenty dcce023b14 Merge branch 'main' of https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos 2026-07-27 04:34:53 +10:00
beatzaplenty 4c5ade5605 updated NFS shares and added vars.nfsShares.options 2026-07-27 04:34:32 +10:00
beatzaplenty b232daf5e1 Merge pull request 'fix(pxe-boot): fix NFS mount in LXC — NFSv3+nolock and skip rpc_pipefs' (#75) from worktree-debian-pxe into main
Reviewed-on: #75
2026-07-26 18:15:58 +00:00
beatzaplentyandClaude Sonnet 4.6 b65736c0dc fix(pxe-boot): fix NFS mount in LXC — NFSv3+nolock and skip rpc_pipefs
Proxmox LXC containers block the sunrpc filesystem (rpc_pipefs) via AppArmor
unless the container has `features: mount=nfs` set. NFSv4 requires rpc_pipefs
for client state management, so the mount fails outright in a default LXC.

Two fixes for the LXC case (config.boot.isContainer):
- Switch from nfsvers=4.2 to nfsvers=3,proto=tcp,nolock,nofail: NFSv3 doesn't
  need rpc_pipefs at the protocol level, and nofail keeps boot clean if the
  NFS server is unreachable.
- Add ConditionVirtualization=!container to var-lib-nfs-rpc_pipefs.mount via
  systemd drop-in: NixOS pulls this unit into nfs-client.target for any NFS
  fileSystems entry. With the condition, systemd skips (not fails) the unit in
  containers, keeping nfs-client.target green and activation reporting clean.

Proxmox VM hosts (not isContainer) continue to use nfsvers=4.2 with
x-systemd.automount unchanged.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 04:14:18 +10:00
beatzaplenty 4f54a1f0cd Merge branch 'main' of https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos 2026-07-27 04:05:15 +10:00
beatzaplenty 2d85ecec8f update pxe-boot images share path 2026-07-27 04:04:59 +10:00
beatzaplenty 34bb14d9f6 Merge pull request 'fix(server): open mountd port 20048 so NFS clients can scan and mount' (#74) from worktree-fix-nfs-mountd-firewall into main
Reviewed-on: #74
2026-07-26 17:57:40 +00:00
beatzaplenty 515da66db9 Merge pull request 'Worktree debian pxe' (#73) from worktree-debian-pxe into main
Reviewed-on: #73
2026-07-26 17:57:31 +00:00
beatzaplentyandClaude Sonnet 4.6 2e9d3da301 fix(server): open mountd port 20048 so NFS clients can scan and mount
showmount and the NFSv3 mount protocol need mountd reachable after querying
portmapper on 111; the server was only opening TCP 111 and 2049, causing
clients (e.g. Proxmox GUI NFS storage scan) to time out connecting to
mountd on 20048. Also adds UDP for all three ports — portmapper, nfsd, and
mountd all use both protocols.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 03:56:39 +10:00
beatzaplentyandClaude Sonnet 4.6 321048626e feat(pxe-boot): mount pxe-images NFS share and symlink to /srv/pxe/http/images
Adds modules/pxe-boot/mount-pxe-images.nix, which mounts
server.sweet.home:/tank/proxmox/pxe-images at /mnt/pxe-images via NFSv4.2
(x-systemd.automount on Proxmox VMs, nofail on LXC containers — same pattern
as docker/mount-data.nix). The pxe-boot build-type now imports this module and
replaces the previous local /srv/pxe/http/images directory rule with an L+
symlink pointing to /mnt/pxe-images, so large images (ISOs, disk images) live
on the NFS share rather than the host's own root disk.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 03:54:37 +10:00
beatzaplenty 16d6baea5f Merge remote-tracking branch 'origin/main' into worktree-debian-pxe 2026-07-27 03:50:56 +10:00
beatzaplenty d082c6a084 updated nfs shares on server build type 2026-07-27 03:40:19 +10:00
beatzaplenty ee93322ae2 add proxmox nfs shares 2026-07-27 03:28:46 +10:00
beatzaplenty d1ce8d3e71 add proxmox iso share 2026-07-27 03:26:31 +10:00
beatzaplentyandClaude Sonnet 4.6 f7c32aff12 feat(pxe-boot): add Debian bookworm minimal netboot menu entry
Adds a `fetch-debian-netboot.service` oneshot that downloads the Debian
bookworm netboot kernel and initrd from deb.debian.org on first boot,
stages them under /srv/pxe/http/debian/, and serves them via a generated
debian.ipxe chain script. The service is idempotent — it skips the
download if both files are already present.

Also merges the previously split systemd.tmpfiles.rules and
systemd.services blocks into a single systemd = { ... } attrset to
satisfy statix's repeated-keys lint.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-27 03:23:14 +10:00
beatzaplenty bf88a6ebb0 added tailscale-router beszel secret
Check NixOS configurations / eval-hosts (push) Successful in 10m22s
2026-07-26 11:14:05 +10:00
beatzaplenty de508141a8 removed tailscale from docker updated secrets
Check NixOS configurations / eval-hosts (push) Successful in 10m33s
2026-07-26 11:12:46 +10:00
beatzaplenty 18cd9c2342 Merge pull request 'feat(beszel): add beszel agent to tailscale-router' (#72) from worktree-beszel-tailscale-router into main
Check NixOS configurations / eval-hosts (push) Failing after 11m27s
Reviewed-on: #72
2026-07-26 01:08:29 +00:00
beatzaplentyandClaude Sonnet 4.6 997918e2f7 feat(beszel): add beszel agent to tailscale-router
Check NixOS configurations / eval-hosts (pull_request) Failing after 11m33s
Wires beszel-agent into all tailscale-router variants (lxc/linode/proxmox)
by importing enable-agent.nix in the build type and host-token.nix in the
host file. Adds the sops creation rule for secrets/tailscale-router.yaml
(all three platform variants as recipients). The secrets file must be
created manually before deploying — see instructions in PR.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-26 11:07:51 +10:00
beatzaplenty 9872b8ff1d Merge pull request 'fix(tailscale-router): add MASQUERADE to POSTROUTING, not nixos-nat-post' (#71) from worktree-lovely-spinning-bubble into main
Check NixOS configurations / eval-hosts (push) Successful in 10m43s
Reviewed-on: #71
2026-07-26 00:01:48 +00:00
beatzaplentyandClaude Sonnet 4.6 e9b225d2d6 fix(tailscale-router): add MASQUERADE to POSTROUTING, not nixos-nat-post
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m31s
extraCommands runs after nixos-nat-post is deleted but before it is
re-created, so -A nixos-nat-post silently fails every time.  POSTROUTING
is a built-in chain that always exists; target it directly instead.

The -C idempotency check prevents duplicate rules on firewall reloads.
Drop networking.nat.enable -- it was only needed for the sub-chain that
turned out to be the wrong target.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-26 10:00:47 +10:00
beatzaplenty 2860f750b4 Merge pull request 'fix(tailscale-router): actually insert MASQUERADE rule via extraCommands' (#70) from worktree-lovely-spinning-bubble into main
Check NixOS configurations / eval-hosts (push) Successful in 10m34s
Reviewed-on: #70
2026-07-25 23:56:11 +00:00
beatzaplentyandClaude Sonnet 4.6 781b1d324e fix(tailscale-router): actually insert MASQUERADE rule via extraCommands
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m34s
networking.nat.externalInterface without internalInterfaces creates the
nixos-nat-post chain but inserts no MASQUERADE rule into it — confirmed
by inspecting the live firewall-start script on the deployed host.

Add the rule explicitly via firewall.extraCommands targeting nixos-nat-post,
scoped to LAN source traffic (vars.lanCidr) going out tailscale0.
extraStopCommands removes it on firewall stop.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-26 09:54:33 +10:00
beatzaplenty 3014a45936 Merge pull request 'fix(tailscale-router): masquerade LAN traffic into Tailscale' (#69) from worktree-lovely-spinning-bubble into main
Check NixOS configurations / eval-hosts (push) Successful in 10m30s
Reviewed-on: #69
2026-07-25 23:48:13 +00:00
beatzaplentyandClaude Sonnet 4.6 cda2132d6a fix(tailscale-router): masquerade LAN traffic into Tailscale
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m31s
Without SNAT on tailscale0, Tailscale drops forwarded packets from LAN
source IPs (192.168.2.x) because they are not recognised Tailscale
addresses.  With networking.nat.externalInterface = "tailscale0", all
traffic leaving through the Tailscale tunnel is masqueraded to the
router's own Tailscale IP (100.x.x.x), making it indistinguishable from
locally-originated traffic.  Conntrack handles the return path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-26 09:45:07 +10:00
beatzaplenty 6c1cc821a0 Merge pull request 'feat(tailscale): enable UDP GRO forwarding on subnet router uplink' (#68) from worktree-polished-wandering-toast into main
Check NixOS configurations / eval-hosts (push) Successful in 10m34s
Reviewed-on: #68
2026-07-25 23:14:33 +00:00
beatzaplentyandClaude Sonnet 4.6 4be064572d feat(tailscale): enable UDP GRO forwarding on subnet router uplink
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m42s
Adds a oneshot systemd service that sets ethtool rx-udp-gro-forwarding on
and rx-gro-list off on the default-route interface at boot, silencing
Tailscale's warning about suboptimal UDP GRO forwarding on subnet routers.
Interface is discovered dynamically via `ip route get` so it works on all
platforms regardless of NIC naming.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-26 09:13:28 +10:00
beatzaplenty 89d506180d Merge pull request 'fix(create-proxmox-resource): fix VM disk never attaching after import' (#67) from worktree-warm-discovering-moon into main
Check NixOS configurations / eval-hosts (push) Successful in 10m33s
Reviewed-on: #67
2026-07-25 22:57:21 +00:00
beatzaplentyandClaude Sonnet 4.6 6e1e992652 fix(proxmox): embed SSH host key via NIXOS_HOST_KEYS_DIR so sops can decrypt on first boot
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m30s
--pre-format-files placed the key on the QEMU builder VM's rootfs, not the
target disk. nixos-install chroots into the target and runs sshd-keygen, which
found no key in the chroot and generated a fresh (unregistered) one. sops then
could not decrypt on first boot because the key didn't match .sops.yaml, leaving
both root and nixos with '!' in /etc/shadow even after mutableUsers = false was
set (hashedPasswordFile pointed to paths sops never wrote).

Fix modules/platforms/proxmox.nix to embed the clan SSH host key in
environment.etc via NIXOS_HOST_KEYS_DIR at eval time -- the same pattern
lxc.nix uses. nixos-install's own activation places the key on the target disk,
sshd-keygen finds it already present and skips generation, and sops decrypts
correctly on first boot. Includes the same preserveSshHostKey/restoreSshHostKey
activation scripts as lxc.nix so subsequent nixos-rebuild switch calls (without
NIXOS_HOST_KEYS_DIR) don't remove the key as "obsolete" from environment.etc.

Update create-proxmox-resource.sh: switch VM builds from
  ./result-<target> --pre-format-files ... --build-memory 2048
to
  NIXOS_HOST_KEYS_DIR=$(pwd)/host-keys nix build --impure ... diskoImagesScript
  ./result-<target> --build-memory 2048
matching the LXC build path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uRcikkTp3D5VbXj2DwNpQ
2026-07-26 08:54:39 +10:00
beatzaplentyandClaude Sonnet 4.6 5467c2e140 fix(create-proxmox-resource): fix VM disk never attaching after import
Three bugs combined to leave every VM build with a shell but no boot disk:

1. The remote build script moved the raw image to /var/lib/vz/import/ before
   qm importdisk could use it. If the mv failed (cross-filesystem copy, sudo
   path, or any other reason) the remote script exited non-zero -- but the
   local script's set -e handling of the SSH heredoc was inconsistent, so
   qm create sometimes ran anyway, leaving a diskless VM shell.

   Fix: skip the mv entirely. The diskoImagesScript writes <hostname>.raw into
   its CWD (the remote repo dir, $out = $PWD at invocation). Import directly
   from that path; clean it up after a successful import.

2. The qm importdisk output regex expected "Successfully imported disk as '...'"
   but current Proxmox emits "unusedN: successfully imported disk '...'"
   (lowercase, no "as"). The grep returned no match and exited 1.

3. The disk_id assignment used $(... | grep ...) without || true inside the
   substitution. With set -euo pipefail, a non-zero grep exit aborts the
   script before the fallback could run -- so the VM was always left with an
   unattached unused0 disk.

   Fix: update the primary regex to match the actual PVE format; add || true
   inside the substitution so set -e never fires on a grep miss; add a qm
   config fallback (scan for unusedN: lines) that works regardless of PVE
   output format changes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uRcikkTp3D5VbXj2DwNpQ
2026-07-26 08:54:39 +10:00
beatzaplenty 123cd2b3d7 Merge branch 'main' of https://gitea.lan.ddnsgeek.com/beatzaplenty/nixos
Check NixOS configurations / eval-hosts (push) Successful in 10m28s
2026-07-26 08:23:44 +10:00
beatzaplenty 006dd8097a updated vm HDD size 2026-07-26 08:23:30 +10:00
beatzaplenty 18cd6e884e Merge pull request 'fix(common): set mutableUsers = false to fix password setup on disk images' (#65) from worktree-warm-discovering-moon into main
Check NixOS configurations / eval-hosts (push) Successful in 10m36s
Reviewed-on: #65
2026-07-25 21:37:34 +00:00
beatzaplenty 3102d66337 Merge pull request 'fix(create-proxmox-resource): case-insensitive importdisk parse + warn on --disk-size for VMs' (#64) from worktree-gentle-cuddling-hippo into main
Check NixOS configurations / eval-hosts (push) Successful in 10m22s
Reviewed-on: #64
2026-07-25 21:37:14 +00:00
beatzaplentyandClaude Sonnet 4.6 dfa5452af5 fix(common): set mutableUsers = false to fix password setup on disk images
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m31s
When a proxmox-* disk image is built, activation runs during the image
build without a valid sops age key (the SSH host key doesn't exist yet),
so root and nixos land in /etc/shadow with locked '!' entries. With the
default mutableUsers = true, update-users-groups.pl preserves existing
shadow entries for accounts that already exist, so hashedPasswordFile is
silently ignored on every subsequent boot — passwords are never fixed.

Setting mutableUsers = false forces update-users-groups.pl to apply
hashedPasswordFile unconditionally on every activation. On first real
boot the sops-decrypted hash is now written regardless of whether the
account already existed in shadow from the image build.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uRcikkTp3D5VbXj2DwNpQ
2026-07-26 07:36:06 +10:00
beatzaplentyandClaude Sonnet 4.6 852ba2240f fix(create-proxmox-resource): case-insensitive importdisk parse + warn on --disk-size for VMs
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m22s
qm importdisk in QEMU 11.x outputs lowercase "successfully imported disk
as '...'" rather than the capitalised form the original grep expected.
The case mismatch made disk_id always empty, which caused the script to
exit 1 after qm create had already run -- leaving the VM with only an
EFI disk, no scsi0, and boot order still set to net0.

Fix by adding -i (case-insensitive) to the grep. Both the old capitalised
format (where the disk id had an "unused0:" prefix inside the quotes) and
the new lowercase format are handled correctly: the sed strip of unused0:
is preserved for backward compatibility, and the regex result is identical
either way.

Also add an early warning when --disk-size is passed for --type vm: the
flag is LXC-only for create mode and was silently ignored, leaving users
expecting a different size than the proxmoxImageSize in variables.nix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-26 06:40:54 +10:00
beatzaplenty 096dff4fa0 Merge pull request 'docs(sync-host-keys): fix stale host-keys/ references in comments and usage' (#63) from fix-stale-wording into main
Check NixOS configurations / eval-hosts (push) Successful in 10m23s
Reviewed-on: #63
2026-07-25 14:36:31 +00:00
beatzaplentyandClaude Sonnet 4.6 b5f749daa9 docs(sync-host-keys): fix stale host-keys/ references in comments and usage
Check NixOS configurations / eval-hosts (pull_request) Successful in 10m24s
After the clan vars migration all keys are in vars/per-machine/, not
host-keys/. Update:
- File header: "existing clan var is never overwritten" (not host-keys/ file)
- Header --remove/--regenerate description: mention clan vars as primary
- usage() --remove, --regenerate-all-keys, --dry-run text
- cmd_remove/cmd_regenerate_all empty-guard messages
- README.md vars/per-machine/ row: "all deployed hosts" (not "LXC hosts")

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B2EJ4qTsM5KUqhS5c3GAwx
2026-07-26 00:11:03 +10:00
beatzaplenty adaf53d647 Merge pull request 'fix(sync-host-keys): extend --remove/--regenerate to cover clan vars' (#62) from fix-sync-host-keys-clan-vars into main
Check NixOS configurations / eval-hosts (push) Successful in 10m22s
Reviewed-on: #62
2026-07-25 14:02:11 +00:00