Simplifies the storage model to one long-term tier (Ceph) instead of two. Node 1 runs local ZFS only until nodes 2/3 join, at which point Ceph goes live and VMs migrate onto it directly.
2.5 KiB
Overview & Roadmap
Background
Current hardware requires disabling KVM hardware virtualization for VMs to start at all (falls back to software emulation — slow). This is a host/BIOS-level issue, not a Proxmox limitation, and shouldn't be needed on new hardware. Before reusing that workaround on the new box:
- Confirm VT-x (Intel) / AMD-V (AMD) is enabled in BIOS/UEFI.
- Confirm IOMMU is enabled if passthrough is planned.
- Update BIOS/microcode first.
- Rule out running Proxmox nested inside another hypervisor.
End goal
3-node Proxmox VE cluster with HA-managed VMs backed by Ceph — true
distributed shared storage, sync replication, near-zero RPO on failover
(see 02-storage-zfs-ceph.md). Ceph needs 3+ nodes and a fast dedicated
network, so it can't exist until nodes 2 and 3 are up.
Deliberately no intermediate "ZFS + storage replication" HA step. Node 1 runs local ZFS for boot + VM storage with no cluster-wide HA until Ceph goes live — as soon as nodes 2/3 join, VMs move onto Ceph rather than adopting ZFS replication as a stopgap. Simpler end state, one storage model to operate instead of two.
Cluster fundamentals (apply from node 1 onward)
- 3 nodes minimum for real quorum. If starting with 2, add a QDevice (small VM or Raspberry Pi) as tie-breaker.
- Dedicated network for corosync (cluster/quorum traffic) — never shared with VM or storage traffic. Needs low, consistent latency (well under 5ms); jitter matters more than bandwidth.
- All nodes on the same PVE version, NTP-synced, SSH reachable between nodes.
- Set VM CPU type to a portable type (e.g.
x86-64-v2-AESorkvm64) rather thanhostif nodes will ever have different CPUs — needed for clean live migration.
Rollout sequence
- Build node 1 per
01-hardware-node1.md— fresh PVE install on ZFS boot mirror, Ceph-earmarked disks left idle (or as a temporary local ZFS pool, to be wiped later — see02-storage-zfs-ceph.md). - Migrate VMs off old hardware via
vzdump→ copy backups →qmrestoreonto the new ZFS storage (converts disks to ZVOLs). No HA yet — single node. - Apply hardening (
04-security-hardening.md) and networking (03-networking.md) before exposing the node beyond the LAN. - Add nodes 2 and 3 identically (same disk/network layout).
- Join cluster, stand up dedicated corosync network.
- Wipe the Ceph-earmarked disks (if used as temporary ZFS) and initialize Ceph across all 3 nodes.
- Migrate VMs from local ZFS onto Ceph-backed storage, then configure HA groups.