Merge pull request 'Replace duplicate-host check with live Proxmox query; drop deployedTargets' (#6) from fix-duplicate-host-self-match into main
Check NixOS configurations / eval-hosts (push) Failing after 12m7s

This commit is contained in:
2026-07-20 10:12:59 +00:00
4 changed files with 69 additions and 47 deletions
+4 -3
View File
@@ -93,9 +93,10 @@ Beyond `codex-setup.sh`/`codex-maintenance.sh` above, `scripts/` also has:
(`--force-rebuild` to skip that and always rebuild), and probes (`--force-rebuild` to skip that and always rebuild), and probes
nix-cache's substituter/remote-builder reachability once up front rather nix-cache's substituter/remote-builder reachability once up front rather
than letting every `nix build` call retry against it individually. than letting every `nix build` call retry against it individually.
Refuses to create a target whose host identity already has a real Refuses to create a target whose host identity already exists live on
deployment elsewhere (`variables.nix`'s `deployedTargets`) unless the node (checked directly via `qm`/`pct`, not any file in this repo)
`--allow-duplicate-host` is passed. `--dry-run` throughout both modes. unless `--allow-duplicate-host` is passed. `--dry-run` throughout both
modes.
- `scripts/env.sh` — shared config (`PROXMOX_HOST`, storage pool, bridge, - `scripts/env.sh` — shared config (`PROXMOX_HOST`, storage pool, bridge,
default cores/memory) sourced by `create-proxmox-resource.sh`. Add new default cores/memory) sourced by `create-proxmox-resource.sh`. Add new
cross-script config here instead of duplicating it per-script. cross-script config here instead of duplicating it per-script.
+6 -4
View File
@@ -28,10 +28,12 @@ list:
| `proxmox-pxe-boot` / `lxc-pxe-boot` | HTTP/iPXE boot asset host (`proxmox-pxe-boot` is the real, deployed one — previously the flat `pxe-boot` target) | | `proxmox-pxe-boot` / `lxc-pxe-boot` | HTTP/iPXE boot asset host (`proxmox-pxe-boot` is the real, deployed one — previously the flat `pxe-boot` target) |
| `linode-tailscale-exit-node` / `proxmox-tailscale-exit-node` / `lxc-tailscale-exit-node` | Tailscale exit node (no deployed target yet; `lxc-tailscale-exit-node` is the one planned for actual use) | | `linode-tailscale-exit-node` / `proxmox-tailscale-exit-node` / `lxc-tailscale-exit-node` | Tailscale exit node (no deployed target yet; `lxc-tailscale-exit-node` is the one planned for actual use) |
The "(real, deployed)" targets above are also tracked machine-readably in This table is the only place "(real, deployed)" status is tracked — there's
`variables.nix`'s `deployedTargets` keep both in sync when a deployment no separate machine-readable copy to keep in sync. `scripts/create-proxmox-resource.sh`
changes. `scripts/create-proxmox-resource.sh` reads that list to refuse guards against creating a same-identity duplicate of an already-deployed host
creating a same-identity duplicate of an already-deployed host by accident. by checking the Proxmox node itself (live `qm`/`pct` state) rather than any
file in this repo, since a static list can't track whether a resource still
actually exists.
Each buildtype's `hosts/<name>/host.nix` carries the per-machine identity Each buildtype's `hosts/<name>/host.nix` carries the per-machine identity
(hostname, hostId, per-machine secrets, `system.stateVersion`) that must stay (hostname, hostId, per-machine secrets, `system.stateVersion`) that must stay
+59 -24
View File
@@ -10,7 +10,9 @@
# #
# SAFETY: # SAFETY:
# - The default (create) mode only ever creates a NEW resource -- it # - The default (create) mode only ever creates a NEW resource -- it
# refuses to run if the target VMID already exists on the node. # refuses to run if the target VMID already exists on the node, or if
# a VM/CT identified as --host already exists under any other VMID
# (checked live against the node; --allow-duplicate-host overrides).
# - --modify only ever touches a resource you name explicitly via # - --modify only ever touches a resource you name explicitly via
# --vmid, shows exactly what will change first, and (outside # --vmid, shows exactly what will change first, and (outside
# --dry-run) always requires typing that VMID back to confirm before # --dry-run) always requires typing that VMID back to confirm before
@@ -56,10 +58,11 @@ Create mode (default):
--force-rebuild Skip the "does the node already have this --force-rebuild Skip the "does the node already have this
image" check -- always build fresh and image" check -- always build fresh and
overwrite what's there. overwrite what's there.
--allow-duplicate-host Required if --host already has a real --allow-duplicate-host Required if a VM/CT identified as --host
deployment elsewhere (variables.nix's already exists on the node (checked live via
deployedTargets) -- otherwise refused, since qm/pct, not any file in this repo) --
it'd share that host's hostName/hostId. otherwise refused, since it'd share that
host's hostName/hostId.
Modify mode (reconfigure an EXISTING resource -- requires --modify): Modify mode (reconfigure an EXISTING resource -- requires --modify):
--modify Switch to modify mode. --modify Switch to modify mode.
@@ -286,25 +289,57 @@ fi
# feeds straight into the guest's real hostname) disagree with host.nix. # feeds straight into the guest's real hostname) disagree with host.nix.
[[ -z "$name" ]] && name="$host" [[ -z "$name" ]] && name="$host"
# --- refuse to duplicate a host that's already really deployed ---------- # --- refuse to duplicate a host that's already live on the node ---------
# Checked by hostName, not exact flake target: proxmox-server being # Queries the node itself (qm/pct's own name/hostname config), not any
# deployed also blocks --type lxc --host server, since both would carry # static list in this repo -- a file can't track whether a resource still
# the same hosts/server/host.nix identity (hostName, hostId). # actually exists, and this used to be checked against variables.nix's
if [[ "$allow_duplicate_host" -eq 0 ]]; then # deployedTargets, which drifted stale (it kept naming a VM as "the real
deployed_targets_json="$(nix eval --json --no-use-registries --no-accept-flake-config \ # deployment" well after that VM had been destroyed, blocking its own
--file "${repo_root}/variables.nix" deployedTargets)" # redeploy) until that list was dropped in favour of this live check. This
for dt in $(echo "$deployed_targets_json" | jq -r '.[]'); do # only catches guests identified with the default --name (== --host, what
dt_hostname="$(nix eval --raw --no-use-registries --no-accept-flake-config \ # this script itself always uses unless --name is overridden) -- a guest
"${repo_root}#nixosConfigurations.${dt}.config.networking.hostName" 2>/dev/null || true)" # manually renamed on the node afterwards wouldn't match, but nothing here
if [[ "$dt_hostname" == "$host" ]]; then # creates guests that way.
echo "ERROR: '${host}' already has a real deployment (${dt}, per variables.nix's" >&2 if [[ "$allow_duplicate_host" -eq 1 ]]; then
echo "deployedTargets). Creating ${flake_target} would share its hostName/hostId --" >&2 echo
echo "refusing by default. Pass --allow-duplicate-host if you really mean to spin" >&2 echo "--allow-duplicate-host: skipping the check for an existing '${host}' on ${node}."
echo "up a separate test instance of this host (it'll still get its own distinct" >&2 elif [[ "$dry_run" -eq 1 ]]; then
echo "sops key and VMID, never touching ${dt})." >&2 echo
exit 1 echo "[dry-run] would check ${node} for an existing VM/CT identified as '${host}'"
fi else
done echo
echo "==> Checking ${node} for an existing VM/CT identified as '${host}'..."
ssh_check_status=0
existing="$(ssh "$ssh_target" bash -s -- "$host" <<'REMOTE_SCRIPT'
target="$1"
for id in $(qm list 2>/dev/null | awk 'NR>1{print $1}'); do
n="$(qm config "$id" 2>/dev/null | grep -oP '^name:\s*\K\S+' || true)"
[[ "$n" == "$target" ]] && echo "vm ${id} ${n}"
done
for id in $(pct list 2>/dev/null | awk 'NR>1{print $1}'); do
n="$(pct config "$id" 2>/dev/null | grep -oP '^hostname:\s*\K\S+' || true)"
[[ "$n" == "$target" ]] && echo "lxc ${id} ${n}"
done
REMOTE_SCRIPT
)" || ssh_check_status=$?
if [[ "$ssh_check_status" -ne 0 ]]; then
echo "ERROR: couldn't reach ${node} (ssh exited ${ssh_check_status}) to check for an" >&2
echo "existing '${host}' resource -- refusing to guess. Fix connectivity and retry," >&2
echo "or pass --allow-duplicate-host if you're sure none exists (this skips the" >&2
echo "check entirely)." >&2
exit 1
fi
if [[ -n "$existing" ]]; then
echo "ERROR: '${host}' already exists on ${node}:" >&2
echo "$existing" | while read -r kind id n; do
echo " - ${kind} VMID ${id} (${n})" >&2
done
echo "Refusing to create a second resource sharing this identity. Pass" >&2
echo "--allow-duplicate-host to create one anyway (it gets its own distinct" >&2
echo "sops key and VMID -- the existing resource(s) above are left untouched)," >&2
echo "or use --modify to reconfigure the existing one instead." >&2
exit 1
fi
fi fi
echo "Target: ${flake_target} (host=${host}, type=${type}) -> Proxmox resource '${name}'" echo "Target: ${flake_target} (host=${host}, type=${type}) -> Proxmox resource '${name}'"
-16
View File
@@ -156,20 +156,4 @@
keep = 20; # number of rotated logs to retain before deleting the oldest keep = 20; # number of rotated logs to retain before deleting the oldest
}; };
# Flake targets with a real, currently-running deployment somewhere —
# matches README.md's Hosts table "(real, deployed)" annotations; update
# both together. Not consumed by any NixOS module (nothing in the actual
# system config should behave differently because of this) — it's read
# by scripts/create-proxmox-resource.sh to refuse creating a same-identity
# duplicate of an already-deployed host (shared hostName/hostId) unless
# you explicitly pass --allow-duplicate-host.
deployedTargets = [
"linode-minimal"
"proxmox-minimal"
"lxc-nix-cache"
"proxmox-server"
"proxmox-docker"
"proxmox-gui"
"proxmox-pxe-boot"
];
} }