pve-test was briefly clustered with pve1 then deliberately de-clustered so it could move to wifi-primary networking (4addr bridge mode, bonded with a wired LAN backup) - a change not achievable while clustered given corosync's latency requirements. Captures that as a reproducible script plus docs: cluster separation procedure, the wifi network design and the live-cutover pitfalls hit along the way, and node-role/history context. Also adds CLAUDE.md guardrails for pve1 (production) vs pve-test (sandbox) - this repo had none before, despite scripts here being able to make real changes to both. Separately: both nodes' mgmt firewalls were dropping ICMP by default (TCP 22/8006 only), which looked like an outage mid-troubleshooting even though SSH/web UI were fine. Added an explicit ping-allow rule to the firewall template, applied it live on both nodes, and added an audit.sh check so it stays enforced.
5.6 KiB
Overview & Roadmap
Staging
This repo now targets two distinct stages, in order:
- Stage 1 (active) — base configuration and hardening for a single
Proxmox host, applicable to any node regardless of eventual cluster
plans: repo/updates, SSH, firewall, PVE user/access hardening. This is
what
scripts/bootstrap.sh,scripts/audit.sh, and04-security-hardening.mdcover, and what's being built out against node 1 (pve1, an ASUS PN53 mini PC) right now. - Stage 2 (future) — the multi-node HA/Ceph cluster described below
and in
01-hardware-node1.md/02-storage-zfs-ceph.md/03-networking.md. Deliberately deferred:pve1's hardware (2 NVMe already merged into one ZFS mirror used for both boot and VM storage, a single 2.5GbE NIC) can't support the separate boot/Ceph disks or bonded/segregated networking those docs assume. Revisit once dedicated cluster hardware (nodes 2/3) is actually being bought and provisioned; until then treat the content below as a target design, not a description ofpve1.
Background
Old hardware required disabling KVM hardware virtualization for VMs to
start at all (falls back to software emulation — slow). This is a
host/BIOS-level issue, not a Proxmox limitation. Confirmed not present on
pve1 (ASUS PN53, Ryzen 7 7735HS): AMD-V and IOMMU both show enabled at
boot (dmesg | grep -i iommu), no workaround needed. If this hardware line
is reused for nodes 2/3, this should hold there too, but re-verify per
node before assuming it.
Stage 1: base config & hardening (active)
Applies to pve1 now, and to every future node regardless of whether it
ever joins the Stage 2 cluster. Covered by 04-security-hardening.md and
scripts/bootstrap.sh / scripts/audit.sh:
- Fresh PVE install; confirm VT-x/AMD-V + IOMMU per the note above.
- Switch off the enterprise repos, onto no-subscription
(
scripts/switch-to-no-subscription-repo.sh). - SSH hardening: key-only root login + fail2ban
(
scripts/harden-ssh.sh). - Unattended security upgrades, no auto-reboot
(
scripts/setup-unattended-upgrades.sh). - PVE datacenter firewall, default-deny, mgmt-only SSH/8006
(
scripts/deploy-firewall.sh). - Named PVE admin user (Administrator role) + 2FA,
root@pamreserved for emergencies (scripts/create-admin-user.sh, then manual TOTP enrollment via the web UI). - Verify with
scripts/audit.sh.
pve1's actual disk/network layout (single ZFS mirror for boot + VMs, one
2.5GbE NIC) is documented as-is in 01-hardware-node1.md — Stage 1 doesn't
require or assume the split-disk/multi-NIC layout Stage 2 wants.
Stage 2: HA cluster + Ceph (future, deferred)
Everything below this point is the target design for when nodes 2 and 3
are actually being provisioned. Not applicable to pve1 as it stands.
pve-test existing does not mean node 2 exists. A second physical
node (pve-test) does run alongside pve1 — see 05-node-roles.md —
but it's a sandbox/test box, not built to this Stage 2 design, and as of
this writing runs on wifi networking
(06-pve-test-wifi-network.md), which is directly incompatible with
corosync's latency/jitter requirements below. pve-test and pve1 were
briefly clustered and then deliberately de-clustered for exactly this
reason. Don't treat pve-test as progress toward Stage 2 without a
deliberate decision to rebuild its networking first.
End goal
3-node Proxmox VE cluster with HA-managed VMs backed by Ceph — true
distributed shared storage, sync replication, near-zero RPO on failover
(see 02-storage-zfs-ceph.md). Ceph needs 3+ nodes and a fast dedicated
network, so it can't exist until nodes 2 and 3 are up.
Deliberately no intermediate "ZFS + storage replication" HA step. Node 1 runs local ZFS for boot + VM storage with no cluster-wide HA until Ceph goes live — as soon as nodes 2/3 join, VMs move onto Ceph rather than adopting ZFS replication as a stopgap. Simpler end state, one storage model to operate instead of two.
Cluster fundamentals (apply from node 1 onward)
- 3 nodes minimum for real quorum. If starting with 2, add a QDevice (small VM or Raspberry Pi) as tie-breaker.
- Dedicated network for corosync (cluster/quorum traffic) — never shared with VM or storage traffic. Needs low, consistent latency (well under 5ms); jitter matters more than bandwidth.
- All nodes on the same PVE version, NTP-synced, SSH reachable between nodes.
- Set VM CPU type to a portable type (e.g.
x86-64-v2-AESorkvm64) rather thanhostif nodes will ever have different CPUs — needed for clean live migration.
Rollout sequence (Stage 2, once dedicated cluster hardware exists)
- Build node 1 per
01-hardware-node1.md's target design — fresh PVE install on a dedicated ZFS boot mirror, Ceph-earmarked disks left idle (or as a temporary local ZFS pool, to be wiped later — see02-storage-zfs-ceph.md). Note: this assumes hardware with enough disks/NICs to separate boot, Ceph, and network roles —pve1does not have this and stays on Stage 1 only unless rebuilt on different hardware. - Migrate VMs onto the new node via
vzdump→ copy backups →qmrestore(converts disks to ZVOLs). No HA yet — single node. - Stage 1 base hardening already applied; layer on Stage 2 networking
(
03-networking.md) before joining a cluster. - Add nodes 2 and 3 identically (same disk/network layout).
- Join cluster, stand up dedicated corosync network.
- Wipe the Ceph-earmarked disks (if used as temporary ZFS) and initialize Ceph across all 3 nodes.
- Migrate VMs from local ZFS onto Ceph-backed storage, then configure HA groups.