Archived
Check NixOS configurations / eval-hosts (push) Successful in 10m26s
Root cause of recurring sops failures on new VM boots: disko builds raw disk images, and create-proxmox-resource.sh only syncs the clan-var SSH host key to the Proxmox node (for proxmox.nix to bake into the image) when it actually builds — reusing a cached image skips sync_remote_host_keys, so destroy+recreate reuses a stale image with the wrong or randomly-generated key baked in. On first boot the VM gets a different key than what .sops.yaml was encrypted for, and sops fails permanently. Fix 1 — deploy.sh Phase 3: always pass --force-rebuild so every VM creation rebuilds the disko image fresh with the current clan-var key baked in via NIXOS_HOST_KEYS_DIR (proxmox.nix already reads this under --impure). Fix 2 — deploy.sh Phase 5.5: after VMs boot, scan their actual ed25519 host keys and, if they drift from clan vars, update the clan var pub-key files, rewrite the .sops.yaml age anchors, and re-encrypt all affected sops files. Defence-in-depth: normally a no-op after Fix 1, but catches any residual mismatch (e.g. --skip-create-vms reuse of an older image). Fix 3 — cluster-init.sh: add crm_standby -v on for both nodes before DRBD metadata init. Without this, Pacemaker's OCF DRBD agent races: it sees drbdadm down as a failure and immediately calls drbdadm up again, leaving /dev/sdb busy when create-md / write-dev-uuid runs (drbdmeta exits 20 with "stdin not a TTY, not waiting for confirmation"). Standby suppresses resource scheduling during init; crm_standby -v off restores it after DRBD is up on both nodes. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
500 lines
21 KiB
Bash
Executable File
500 lines
21 KiB
Bash
Executable File
#!/usr/bin/env bash
|
||
# deploy.sh — Full lifecycle management for the HA file-server cluster.
|
||
#
|
||
# Handles everything from zero (no VMs, no secrets) through a running,
|
||
# tested cluster, and optionally tears it back down.
|
||
#
|
||
# Usage:
|
||
# scripts/ha/deploy.sh [options]
|
||
# scripts/ha/deploy.sh --destroy [options]
|
||
#
|
||
# Phases (all run by default; skip any with --skip-*):
|
||
# 1. ensure-bridge Create storage bridge (vmbr1) on the Proxmox node if absent.
|
||
# 2. sync-keys Generate SSH host keys and register age keys for both nodes.
|
||
# 3. create-vms Build disk images and create both VMs via create-proxmox-resource.sh.
|
||
# 4. add-hardware Attach storage NIC (vmbr1) and DRBD data disk to each VM.
|
||
# 5. boot-wait Start VMs, wait for SSH on both nodes.
|
||
# 6. cluster-init Form the cluster: DRBD, corosync, Pacemaker, NFS, VIP.
|
||
# Also encrypts the generated corosync authkey into the repo.
|
||
# 7. run-tests Run acceptance tests (T1–T7).
|
||
#
|
||
# Options:
|
||
# --node <host> Proxmox host to deploy on (default: pve1.sweet.home)
|
||
# --vmid1 <n> VMID for ha-server-1 (default: 200)
|
||
# --vmid2 <n> VMID for ha-server-2 (default: 201)
|
||
# --storage <pool> Proxmox storage pool (default: local-zfs)
|
||
# --storage-bridge <br> Bridge for HA storage network (default: vmbr1)
|
||
# --drbd-disk-gb <n> DRBD data disk size in GB (default: 32)
|
||
# --memory <MB> RAM per node (default: 4096)
|
||
# --cores <n> vCPUs per node (default: 4)
|
||
# --skip-ensure-bridge Skip storage bridge creation/check
|
||
# --skip-sync-keys Skip sync-host-keys.sh (clan vars already exist)
|
||
# --skip-create-vms Skip VM creation (VMs already exist)
|
||
# --skip-add-hardware Skip net1/scsi1 attachment (already attached)
|
||
# --skip-boot-wait Skip boot/SSH wait (VMs already running)
|
||
# --skip-refresh-sops-keys Skip scanning running VMs for fresh SSH host keys
|
||
# --skip-cluster-init Skip cluster formation (cluster already configured)
|
||
# --skip-tests Skip acceptance tests
|
||
# --force-rebuild Pass --force-rebuild to create-proxmox-resource.sh
|
||
# --destroy Stop and delete both VMs (skip all other phases)
|
||
# --dry-run Print what would run without executing
|
||
# -h|--help Show this message
|
||
#
|
||
# Prerequisites:
|
||
# - SSH access to the Proxmox node as $PROXMOX_SSH_USER (wayne).
|
||
# - sops age key in the standard location (used by sync-host-keys.sh).
|
||
# - For --skip-sync-keys: clan vars already in vars/per-machine/proxmox-ha-server-{1,2}/.
|
||
# - For full tests: secrets/common.yaml decryptable on both nodes (run
|
||
# `sops updatekeys secrets/common.yaml` after sync-keys).
|
||
|
||
set -euo pipefail
|
||
|
||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||
REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
|
||
|
||
# shellcheck source=../env.sh
|
||
source "${REPO_ROOT}/scripts/env.sh"
|
||
|
||
# ── Defaults ──────────────────────────────────────────────────────────────────
|
||
|
||
NODE="${PROXMOX_HOST:-$PVE1_HOST}"
|
||
VMID1=200
|
||
VMID2=201
|
||
STORAGE="${PROXMOX_STORAGE:-local-zfs}"
|
||
STORAGE_BRIDGE="vmbr1"
|
||
DRBD_DISK_GB=32
|
||
MEMORY_MB=4096
|
||
CORES=4
|
||
|
||
SKIP_ENSURE_BRIDGE=false
|
||
SKIP_SYNC_KEYS=false
|
||
SKIP_CREATE_VMS=false
|
||
SKIP_ADD_HARDWARE=false
|
||
SKIP_BOOT_WAIT=false
|
||
SKIP_REFRESH_SOPS_KEYS=false
|
||
SKIP_CLUSTER_INIT=false
|
||
SKIP_TESTS=false
|
||
FORCE_REBUILD=false
|
||
DESTROY=false
|
||
DRY_RUN=false
|
||
|
||
# ── Variables from repo ───────────────────────────────────────────────────────
|
||
|
||
NODE1_HOST="ha-server-1"
|
||
NODE2_HOST="ha-server-2"
|
||
NODE1_IP="192.168.2.228"
|
||
NODE2_IP="192.168.2.227"
|
||
STORAGE_IP1="192.168.4.228"
|
||
STORAGE_IP2="192.168.4.227"
|
||
STORAGE_CIDR="192.168.4.0/29"
|
||
SSH_USER="${PROXMOX_SSH_USER:-wayne}"
|
||
|
||
# ── Argument parsing ──────────────────────────────────────────────────────────
|
||
|
||
usage() {
|
||
sed -n '/^# Usage:/,/^[^#]/{ /^#/{ s/^# \?//; p } }' "$0"
|
||
exit "${1:-0}"
|
||
}
|
||
|
||
while [[ $# -gt 0 ]]; do
|
||
case "$1" in
|
||
--node) NODE="$2"; shift 2 ;;
|
||
--vmid1) VMID1="$2"; shift 2 ;;
|
||
--vmid2) VMID2="$2"; shift 2 ;;
|
||
--storage) STORAGE="$2"; shift 2 ;;
|
||
--storage-bridge) STORAGE_BRIDGE="$2"; shift 2 ;;
|
||
--drbd-disk-gb) DRBD_DISK_GB="$2"; shift 2 ;;
|
||
--memory) MEMORY_MB="$2"; shift 2 ;;
|
||
--cores) CORES="$2"; shift 2 ;;
|
||
--skip-ensure-bridge) SKIP_ENSURE_BRIDGE=true; shift ;;
|
||
--skip-sync-keys) SKIP_SYNC_KEYS=true; shift ;;
|
||
--skip-create-vms) SKIP_CREATE_VMS=true; shift ;;
|
||
--skip-add-hardware) SKIP_ADD_HARDWARE=true; shift ;;
|
||
--skip-boot-wait) SKIP_BOOT_WAIT=true; shift ;;
|
||
--skip-refresh-sops-keys) SKIP_REFRESH_SOPS_KEYS=true; shift ;;
|
||
--skip-cluster-init) SKIP_CLUSTER_INIT=true; shift ;;
|
||
--skip-tests) SKIP_TESTS=true; shift ;;
|
||
--force-rebuild) FORCE_REBUILD=true; shift ;;
|
||
--destroy) DESTROY=true; shift ;;
|
||
--dry-run) DRY_RUN=true; shift ;;
|
||
-h|--help) usage 0 ;;
|
||
*) echo "Unknown option: $1" >&2; usage 1 ;;
|
||
esac
|
||
done
|
||
|
||
# ── Helpers ───────────────────────────────────────────────────────────────────
|
||
|
||
log() { echo "==> $*"; }
|
||
logn() { echo " $*"; }
|
||
err() { echo "ERROR: $*" >&2; exit 1; }
|
||
|
||
run() {
|
||
if $DRY_RUN; then
|
||
echo "[dry-run] $*"
|
||
else
|
||
"$@"
|
||
fi
|
||
}
|
||
|
||
pve() {
|
||
# Run a command on the Proxmox node via SSH.
|
||
if $DRY_RUN; then
|
||
echo "[dry-run] ssh ${SSH_USER}@${NODE} sudo $*"
|
||
else
|
||
ssh -i ~/.ssh/id_ed25519 "${SSH_USER}@${NODE}" "sudo $*"
|
||
fi
|
||
}
|
||
|
||
pve_check() {
|
||
# Run a read-only probe on the Proxmox node — always executes even in dry-run.
|
||
ssh -i ~/.ssh/id_ed25519 "${SSH_USER}@${NODE}" "sudo $*"
|
||
}
|
||
|
||
HA_USER="nixos"
|
||
|
||
n1() {
|
||
# Run a command on ha-server-1 via SSH as nixos with sudo.
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no -o ConnectTimeout=5 "${HA_USER}@${NODE1_IP}" sudo "$@" 2>/dev/null
|
||
}
|
||
|
||
n2() {
|
||
# Run a command on ha-server-2 via SSH as nixos with sudo.
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no -o ConnectTimeout=5 "${HA_USER}@${NODE2_IP}" sudo "$@" 2>/dev/null
|
||
}
|
||
|
||
wait_for_ssh() {
|
||
local ip="$1" label="$2"
|
||
if $DRY_RUN; then
|
||
logn "[dry-run] Skipping SSH wait for ${label} (${ip})"
|
||
return 0
|
||
fi
|
||
local deadline=$(( $(date +%s) + 300 ))
|
||
log "Waiting for SSH on ${label} (${ip}) — up to 5 min..."
|
||
while [[ $(date +%s) -lt $deadline ]]; do
|
||
if ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no -o ConnectTimeout=3 \
|
||
-o BatchMode=yes "${HA_USER}@${ip}" true 2>/dev/null; then
|
||
logn "${label} is up."
|
||
return 0
|
||
fi
|
||
sleep 5
|
||
done
|
||
err "Timed out waiting for SSH on ${label} (${ip})"
|
||
}
|
||
|
||
# ── Destroy mode ─────────────────────────────────────────────────────────────
|
||
|
||
if $DESTROY; then
|
||
log "Destroying HA cluster VMs (${VMID1}=${NODE1_HOST}, ${VMID2}=${NODE2_HOST}) on ${NODE}"
|
||
for vmid in "$VMID1" "$VMID2"; do
|
||
STATUS=$(pve "qm status ${vmid} 2>/dev/null" 2>/dev/null || true)
|
||
if echo "$STATUS" | grep -q "running"; then
|
||
log "Stopping VMID ${vmid}..."
|
||
pve "qm stop ${vmid} --skiplock 1"
|
||
sleep 5
|
||
fi
|
||
if $DRY_RUN || pve "qm config ${vmid} >/dev/null 2>&1"; then
|
||
log "Deleting VMID ${vmid}..."
|
||
run pve "qm destroy ${vmid} --purge 1"
|
||
else
|
||
logn "VMID ${vmid} not found — already gone."
|
||
fi
|
||
done
|
||
log "Done — cluster VMs destroyed."
|
||
exit 0
|
||
fi
|
||
|
||
# ── Phase 1: Storage bridge ───────────────────────────────────────────────────
|
||
|
||
if ! $SKIP_ENSURE_BRIDGE; then
|
||
log "Phase 1: Ensuring storage bridge ${STORAGE_BRIDGE} on ${NODE}"
|
||
if pve_check "test -d /sys/class/net/${STORAGE_BRIDGE}" &>/dev/null; then
|
||
logn "${STORAGE_BRIDGE} already exists — skipping."
|
||
else
|
||
logn "Creating isolated internal bridge ${STORAGE_BRIDGE} (no upstream port, ${STORAGE_CIDR})"
|
||
BRIDGE_CONF="auto ${STORAGE_BRIDGE}
|
||
iface ${STORAGE_BRIDGE} inet manual
|
||
bridge-ports none
|
||
bridge-stp off
|
||
bridge-fd 0"
|
||
if $DRY_RUN; then
|
||
echo "[dry-run] Would write /etc/network/interfaces.d/${STORAGE_BRIDGE}.conf and ifup it"
|
||
else
|
||
ssh -i ~/.ssh/id_ed25519 "${SSH_USER}@${NODE}" \
|
||
"echo '${BRIDGE_CONF}' | sudo tee /etc/network/interfaces.d/${STORAGE_BRIDGE}.conf > /dev/null && sudo ifup ${STORAGE_BRIDGE}"
|
||
logn "${STORAGE_BRIDGE} created and brought up."
|
||
fi
|
||
fi
|
||
fi
|
||
|
||
# ── Phase 2: Sync host keys ───────────────────────────────────────────────────
|
||
|
||
if ! $SKIP_SYNC_KEYS; then
|
||
log "Phase 2: Syncing SSH host keys for both HA targets"
|
||
for target in proxmox-ha-server-1 proxmox-ha-server-2; do
|
||
CLAN_DIR="${REPO_ROOT}/vars/per-machine/${target}/openssh"
|
||
if [[ -d "$CLAN_DIR" ]]; then
|
||
logn "Clan vars for ${target} already exist — skipping."
|
||
else
|
||
logn "Generating host keys for ${target}..."
|
||
run bash "${REPO_ROOT}/scripts/secrets/sync-host-keys.sh" "$target"
|
||
fi
|
||
done
|
||
fi
|
||
|
||
# ── Phase 2.5: Prepare Proxmox node for building ─────────────────────────────
|
||
if ! $SKIP_CREATE_VMS && ! $DRY_RUN; then
|
||
CURRENT_BRANCH="$(git -C "$REPO_ROOT" rev-parse --abbrev-ref HEAD)"
|
||
|
||
# Fix /nix ownership if it exists but belongs to a different UID.
|
||
# pve1's IPA-enrolled wayne (UID 50002) can't write to a store created by
|
||
# another UID — passwordless sudo corrects it once.
|
||
# Use direct SSH (no sudo) for the writability check so we test wayne's own
|
||
# access, not root's.
|
||
local_ssh() { ssh -i ~/.ssh/id_ed25519 "${SSH_USER}@${NODE}" "$*"; }
|
||
if local_ssh "test -d /nix" &>/dev/null && ! local_ssh "test -w /nix" &>/dev/null; then
|
||
logn "/nix exists but not writable by ${SSH_USER} — fixing ownership with sudo (one-time)..."
|
||
local_ssh "sudo chown -R ${SSH_USER} /nix"
|
||
logn "Done."
|
||
fi
|
||
unset -f local_ssh
|
||
|
||
# Ensure the remote clone is on the correct branch so create-proxmox-resource.sh
|
||
# builds from the same commits we're deploying.
|
||
REMOTE_REPO="/home/${SSH_USER}/nixos"
|
||
if pve_check "test -d ${REMOTE_REPO}/.git" &>/dev/null; then
|
||
REMOTE_BRANCH=$(ssh -i ~/.ssh/id_ed25519 "${SSH_USER}@${NODE}" \
|
||
"cd ${REMOTE_REPO} && git rev-parse --abbrev-ref HEAD 2>/dev/null")
|
||
if [[ "$REMOTE_BRANCH" != "$CURRENT_BRANCH" ]]; then
|
||
logn "Remote clone is on '${REMOTE_BRANCH}', switching to '${CURRENT_BRANCH}'..."
|
||
ssh -i ~/.ssh/id_ed25519 "${SSH_USER}@${NODE}" \
|
||
"cd ${REMOTE_REPO} && git fetch origin && git checkout '${CURRENT_BRANCH}' && git pull --ff-only"
|
||
logn "Done."
|
||
fi
|
||
fi
|
||
fi
|
||
|
||
# ── Phase 3: Create VMs ───────────────────────────────────────────────────────
|
||
|
||
if ! $SKIP_CREATE_VMS; then
|
||
log "Phase 3: Building and creating VMs on ${NODE}"
|
||
|
||
CREATE="${REPO_ROOT}/scripts/proxmox/create-proxmox-resource.sh"
|
||
|
||
for spec in "${VMID1}:ha-server-1:proxmox-ha-server-1" "${VMID2}:ha-server-2:proxmox-ha-server-2"; do
|
||
IFS=: read -r vmid host_name flake_target <<< "$spec"
|
||
log "Creating ${flake_target} (VMID ${vmid}) on ${NODE}..."
|
||
# Always --force-rebuild: create-proxmox-resource.sh only calls
|
||
# sync_remote_host_keys (which populates host-keys/ for proxmox.nix to
|
||
# bake the clan-var SSH key into the disko image) when it actually builds.
|
||
# Reusing a cached image skips that step, so destroy+recreate would reuse
|
||
# an image with a stale/random key baked in → sops fails on first boot.
|
||
run bash "$CREATE" \
|
||
--type vm \
|
||
--host "$host_name" \
|
||
--vmid "$vmid" \
|
||
--node "$NODE" \
|
||
--storage "$STORAGE" \
|
||
--memory "$MEMORY_MB" \
|
||
--cores "$CORES" \
|
||
--force-rebuild
|
||
done
|
||
fi
|
||
|
||
# ── Phase 4: Add storage NIC and DRBD disk ────────────────────────────────────
|
||
|
||
if ! $SKIP_ADD_HARDWARE; then
|
||
log "Phase 4: Attaching storage NIC (${STORAGE_BRIDGE}) and DRBD disk (${DRBD_DISK_GB}G) to each VM"
|
||
for vmid in "$VMID1" "$VMID2"; do
|
||
log " VMID ${vmid}: stopping to add hardware..."
|
||
pve "qm stop ${vmid} --skiplock 1 2>/dev/null; sleep 3" || true
|
||
|
||
logn "Adding net1 (${STORAGE_BRIDGE})..."
|
||
pve "qm set ${vmid} --net1 virtio,bridge=${STORAGE_BRIDGE},firewall=0"
|
||
|
||
logn "Adding scsi1 (${STORAGE}:${DRBD_DISK_GB}G for DRBD)..."
|
||
pve "qm set ${vmid} --scsi1 ${STORAGE}:${DRBD_DISK_GB},format=raw"
|
||
|
||
logn "Starting VMID ${vmid}..."
|
||
pve "qm start ${vmid}"
|
||
done
|
||
fi
|
||
|
||
# ── Phase 5: Wait for SSH ─────────────────────────────────────────────────────
|
||
|
||
if ! $SKIP_BOOT_WAIT; then
|
||
log "Phase 5: Waiting for both nodes to come up"
|
||
wait_for_ssh "$NODE1_IP" "$NODE1_HOST"
|
||
wait_for_ssh "$NODE2_IP" "$NODE2_HOST"
|
||
logn "Both nodes are SSHable."
|
||
# Give systemd a few seconds to settle after activation
|
||
sleep 10
|
||
fi
|
||
|
||
# ── Phase 5.5: Refresh sops host-key registrations ───────────────────────────
|
||
#
|
||
# Disko builds raw disk images; each new VM boots with a freshly-generated SSH
|
||
# host key rather than the one pre-seeded in clan vars. This phase scans the
|
||
# actual running VMs, and if their ed25519 host keys differ from what clan vars
|
||
# record: updates the clan var pub-key files, rewrites the .sops.yaml age-key
|
||
# anchors, and re-encrypts all affected sops files so the nodes can decrypt
|
||
# secrets on the next nixos-rebuild. Safe no-op when keys haven't changed.
|
||
|
||
if ! $SKIP_REFRESH_SOPS_KEYS; then
|
||
if $DRY_RUN; then
|
||
logn "[dry-run] Would scan VM host keys and refresh .sops.yaml / secrets if needed"
|
||
else
|
||
log "Phase 5.5: Refreshing sops host-key registrations (disko key drift fix)"
|
||
SOPS_UPDATED=false
|
||
|
||
for spec in \
|
||
"${NODE1_IP}:proxmox-ha-server-1:${NODE1_HOST}" \
|
||
"${NODE2_IP}:proxmox-ha-server-2:${NODE2_HOST}"; do
|
||
IFS=: read -r node_ip flake_target host_name <<< "$spec"
|
||
CLAN_PUB="${REPO_ROOT}/vars/per-machine/${flake_target}/openssh/ssh_host_ed25519_key.pub/value"
|
||
|
||
logn "Scanning ed25519 host key from ${host_name} (${node_ip})..."
|
||
RAW=$(ssh-keyscan -t ed25519 "${node_ip}" 2>/dev/null | grep -v "^#") || true
|
||
if [[ -z "$RAW" ]]; then
|
||
logn "WARNING: no ed25519 key returned by ssh-keyscan for ${node_ip} — skipping"
|
||
continue
|
||
fi
|
||
# ssh-keyscan returns: <ip> ssh-ed25519 <b64key>
|
||
SCANNED_TYPE=$(awk '{print $2}' <<< "$RAW")
|
||
SCANNED_KEY=$(awk '{print $3}' <<< "$RAW")
|
||
SCANNED_PUBKEY="${SCANNED_TYPE} ${SCANNED_KEY} ${host_name}"
|
||
|
||
CURRENT=$(tr -d '\n' < "$CLAN_PUB" 2>/dev/null || true)
|
||
if [[ "$SCANNED_PUBKEY" == "$CURRENT" ]]; then
|
||
logn "${host_name}: clan var matches running key — no update needed"
|
||
continue
|
||
fi
|
||
|
||
logn "${host_name}: key drift detected — updating clan var"
|
||
logn " old: ${CURRENT}"
|
||
logn " new: ${SCANNED_PUBKEY}"
|
||
echo "$SCANNED_PUBKEY" > "$CLAN_PUB"
|
||
SOPS_UPDATED=true
|
||
|
||
# Rewrite the .sops.yaml anchor for this host with the new age key.
|
||
ANCHOR="${flake_target}" # e.g. proxmox-ha-server-1
|
||
NEW_AGE=$(echo "$SCANNED_PUBKEY" | \
|
||
nix run --quiet --no-warn-dirty nixpkgs#ssh-to-age 2>/dev/null)
|
||
if [[ -z "$NEW_AGE" ]]; then
|
||
err "ssh-to-age produced no output for ${host_name} — check nixpkgs#ssh-to-age"
|
||
fi
|
||
logn " new age key: ${NEW_AGE}"
|
||
sed -i "/&${ANCHOR} /s| age[a-z0-9]*$| ${NEW_AGE}|" "${REPO_ROOT}/.sops.yaml"
|
||
done
|
||
|
||
if $SOPS_UPDATED; then
|
||
logn "Running sops updatekeys on affected secrets..."
|
||
SOPS="nix run --quiet --no-warn-dirty nixpkgs#sops --"
|
||
(cd "${REPO_ROOT}" && \
|
||
$SOPS updatekeys -y secrets/common.yaml && \
|
||
$SOPS updatekeys -y secrets/ha-server-1.yaml && \
|
||
$SOPS updatekeys -y secrets/ha-server-2.yaml && \
|
||
$SOPS updatekeys -y secrets/ha-server-1.keytab && \
|
||
$SOPS updatekeys -y secrets/ha-server-2.keytab)
|
||
# Note: ha-corosync-authkey is re-generated and re-encrypted by cluster-init below.
|
||
|
||
logn "Committing refreshed host keys and re-encrypted secrets..."
|
||
(cd "${REPO_ROOT}" && \
|
||
git add \
|
||
vars/per-machine/proxmox-ha-server-1/openssh/ssh_host_ed25519_key.pub/value \
|
||
vars/per-machine/proxmox-ha-server-2/openssh/ssh_host_ed25519_key.pub/value \
|
||
.sops.yaml \
|
||
secrets/common.yaml \
|
||
secrets/ha-server-1.yaml \
|
||
secrets/ha-server-2.yaml \
|
||
secrets/ha-server-1.keytab \
|
||
secrets/ha-server-2.keytab && \
|
||
git commit -m "secrets(ha): refresh sops host-key registrations for new VM instances" || true)
|
||
logn "Sops keys refreshed and committed."
|
||
fi
|
||
fi
|
||
fi
|
||
|
||
# ── Phase 6: Cluster init ─────────────────────────────────────────────────────
|
||
|
||
if ! $SKIP_CLUSTER_INIT; then
|
||
log "Phase 6: Initialising HA cluster"
|
||
|
||
CLUSTER_INIT="${REPO_ROOT}/scripts/ha/cluster-init.sh"
|
||
[[ -x "$CLUSTER_INIT" ]] || chmod +x "$CLUSTER_INIT"
|
||
|
||
if $DRY_RUN; then
|
||
logn "[dry-run] Would generate temp key, authorise on ${NODE2_HOST}, scp cluster-init.sh to ${NODE1_HOST}, and run it as root via sudo"
|
||
else
|
||
# Generate a temp keypair so cluster-init.sh can SSH node1→node2 as ${HA_USER}.
|
||
# Root on node1 has no keys; a temp key bridging node1→node2 nixos solves this.
|
||
TEMP_KEY="${REPO_ROOT}/.tmp-cluster-init-key"
|
||
TEMP_KEY_PUB="${TEMP_KEY}.pub"
|
||
rm -f "$TEMP_KEY" "$TEMP_KEY_PUB"
|
||
ssh-keygen -t ed25519 -f "$TEMP_KEY" -N "" -C "cluster-init-temp-$(date +%s)" -q
|
||
TEMP_PUBKEY=$(cat "$TEMP_KEY_PUB")
|
||
|
||
logn "Authorising temp key on ${NODE2_HOST} for ${HA_USER}..."
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no "${HA_USER}@${NODE2_IP}" \
|
||
"mkdir -p ~/.ssh && chmod 700 ~/.ssh && echo '${TEMP_PUBKEY}' >> ~/.ssh/authorized_keys"
|
||
|
||
logn "Placing temp key on ${NODE1_HOST} for root..."
|
||
scp -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no \
|
||
"$TEMP_KEY" "${HA_USER}@${NODE1_IP}:/tmp/cluster-init-key"
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no "${HA_USER}@${NODE1_IP}" \
|
||
"sudo mkdir -p /root/.ssh && sudo cp /tmp/cluster-init-key /root/.ssh/cluster-init-key && \
|
||
sudo chmod 600 /root/.ssh/cluster-init-key && rm -f /tmp/cluster-init-key"
|
||
|
||
logn "Uploading cluster-init.sh to ${NODE1_HOST}..."
|
||
scp -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no \
|
||
"$CLUSTER_INIT" "${HA_USER}@${NODE1_IP}:/tmp/cluster-init.sh"
|
||
|
||
logn "Running cluster-init.sh on ${NODE1_HOST}..."
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no "${HA_USER}@${NODE1_IP}" \
|
||
"sudo env NODE1=${NODE1_HOST} NODE2=${NODE2_HOST} \
|
||
NODE1_IP=${NODE1_IP} NODE2_IP=${NODE2_IP} \
|
||
VIP=192.168.2.229 XFS_MOUNT=/srv/ha-data \
|
||
ISCSI_IQN=iqn.2026-01.home.sweet:ha-storage \
|
||
VMID_NODE1=${VMID1} VMID_NODE2=${VMID2} \
|
||
HA_USER=${HA_USER} HA_KEY=/root/.ssh/cluster-init-key \
|
||
bash /tmp/cluster-init.sh"
|
||
|
||
logn "Cleaning up temp key from both nodes..."
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no "${HA_USER}@${NODE2_IP}" \
|
||
"sed -i '/cluster-init-temp/d' ~/.ssh/authorized_keys" 2>/dev/null || true
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no "${HA_USER}@${NODE1_IP}" \
|
||
"sudo rm -f /root/.ssh/cluster-init-key" 2>/dev/null || true
|
||
rm -f "$TEMP_KEY" "$TEMP_KEY_PUB"
|
||
|
||
# Encrypt the corosync authkey generated by cluster-init and commit it.
|
||
log " Encrypting corosync authkey into secrets/ha-corosync-authkey..."
|
||
AUTHKEY_TMP="${REPO_ROOT}/secrets/ha-corosync-authkey.tmp"
|
||
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=no "${HA_USER}@${NODE1_IP}" \
|
||
"sudo cat /etc/corosync/authkey" > "$AUTHKEY_TMP"
|
||
if [[ ! -s "$AUTHKEY_TMP" ]]; then
|
||
err "corosync authkey on node1 is empty — cluster-init may have failed."
|
||
fi
|
||
mv "$AUTHKEY_TMP" "${REPO_ROOT}/secrets/ha-corosync-authkey"
|
||
(cd "${REPO_ROOT}" && nix run nixpkgs#sops -- -e --input-type binary -i secrets/ha-corosync-authkey)
|
||
logn "Authkey encrypted. Committing..."
|
||
(cd "${REPO_ROOT}" && git add secrets/ha-corosync-authkey && \
|
||
git commit -m "secrets(ha): encrypt corosync authkey generated by cluster-init")
|
||
logn "Committed."
|
||
fi
|
||
fi
|
||
|
||
# ── Phase 7: Acceptance tests ─────────────────────────────────────────────────
|
||
|
||
if ! $SKIP_TESTS; then
|
||
log "Phase 7: Running acceptance tests (T1–T7)"
|
||
if $DRY_RUN; then
|
||
logn "[dry-run] Would run acceptance-tests.sh against ${NODE1_HOST}/${NODE2_HOST}"
|
||
else
|
||
NODE1="$NODE1_HOST" NODE2="$NODE2_HOST" \
|
||
NODE1_IP="$NODE1_IP" NODE2_IP="$NODE2_IP" \
|
||
VIP="192.168.2.229" \
|
||
bash "${REPO_ROOT}/scripts/ha/acceptance-tests.sh"
|
||
fi
|
||
fi
|
||
|
||
log "Deploy complete."
|