pve-test was briefly clustered with pve1 then deliberately de-clustered so it could move to wifi-primary networking (4addr bridge mode, bonded with a wired LAN backup) - a change not achievable while clustered given corosync's latency requirements. Captures that as a reproducible script plus docs: cluster separation procedure, the wifi network design and the live-cutover pitfalls hit along the way, and node-role/history context. Also adds CLAUDE.md guardrails for pve1 (production) vs pve-test (sandbox) - this repo had none before, despite scripts here being able to make real changes to both. Separately: both nodes' mgmt firewalls were dropping ICMP by default (TCP 22/8006 only), which looked like an outage mid-troubleshooting even though SSH/web UI were fine. Added an explicit ping-allow rule to the firewall template, applied it live on both nodes, and added an audit.sh check so it stays enforced.
2.5 KiB
Proxmox Configuration
Base configuration and hardening toolset for Proxmox VE hosts, plus planning
docs for eventually growing this into a 3-node HA/Ceph cluster. See
docs/00-overview.md for the staging: Stage 1 (base config/hardening,
applies to any host — active) vs. Stage 2 (multi-node HA/Ceph — future,
deferred).
Goals
- Stage 1: a reusable, idempotent base-hardening toolset (
scripts/) that can be run against any new Proxmox host — repo/updates, SSH, firewall, PVE user/access hardening — verified withscripts/audit.sh. - Stage 2 (future): 3-node cluster, quorum via corosync, HA-managed VMs
backed by Ceph. Needs dedicated hardware node 1 (
pve1, an ASUS PN53 mini PC) doesn't have — seedocs/01-hardware-node1.md.
See CLAUDE.md for the guardrails Claude Code follows when working
against these hosts (pve1 is production and off-limits by default;
pve-test is the sandbox).
Repo layout
docs/— planning docs: hardware layout, storage migration, networking, security hardening, node roles. Readdocs/00-overview.mdfirst, thendocs/05-node-roles.mdfor whatpve1/pve-testactually are.scripts/— scripts to apply configuration on a node (SSH hardening, repo switch, firewall, updates, wifi/bond networking, etc). Idempotent, safe to re-run.scripts/bootstrap.shruns the full Stage 1 sequence end to end;scripts/audit.shverifies it (read-only).scripts/setup-wifi-bond-network.shreproducespve-test's wifi network (seedocs/06-pve-test-wifi-network.md).scripts/lib/holds shared helpers (common.sh) sourced by the other scripts.config/— reference config files/snippets to drop onto a node (firewall rules, sshd config, etc.).
Quick start (Stage 1, on a fresh node)
MGMT_CIDR=192.168.2.0/24 ./scripts/bootstrap.sh
./scripts/create-admin-user.sh <username>
# then enable 2FA for that user + root@pam via the web UI
./scripts/audit.sh
Status
pve1 built (ASUS PN53 mini PC, ZFS mirror boot+VM storage, single
2.5GbE NIC) and already running production VMs/CTs. Stage 1 base
hardening applied and verified (scripts/audit.sh all green): enterprise
repos removed, SSH key-only + fail2ban, unattended security upgrades,
PVE firewall (mgmt-only), subscription nag disabled, named admin user
(wayne@pve) created. Remaining manual step: enable 2FA/TOTP for
wayne@pve and root@pam via the web UI. Stage 2 (cluster/Ceph) not
started — needs nodes 2/3 on hardware that can actually support it.