Archived
pve-test was briefly clustered with pve1 then deliberately de-clustered so it could move to wifi-primary networking (4addr bridge mode, bonded with a wired LAN backup) - a change not achievable while clustered given corosync's latency requirements. Captures that as a reproducible script plus docs: cluster separation procedure, the wifi network design and the live-cutover pitfalls hit along the way, and node-role/history context. Also adds CLAUDE.md guardrails for pve1 (production) vs pve-test (sandbox) - this repo had none before, despite scripts here being able to make real changes to both. Separately: both nodes' mgmt firewalls were dropping ICMP by default (TCP 22/8006 only), which looked like an outage mid-troubleshooting even though SSH/web UI were fine. Added an explicit ping-allow rule to the firewall template, applied it live on both nodes, and added an audit.sh check so it stays enforced.
45 lines
1.8 KiB
Plaintext
45 lines
1.8 KiB
Plaintext
# Example cluster-wide firewall rules for /etc/pve/firewall/cluster.fw
|
|
#
|
|
# Stage 1 (single host, current): only the mgmt IPSET applies. Applied
|
|
# automatically by scripts/deploy-firewall.sh, which fills in <MGMT_CIDR>.
|
|
#
|
|
# Stage 2 (future cluster/Ceph): the corosync and Ceph rules below are
|
|
# commented out placeholders. Uncomment and fill in <COROSYNC_CIDR> /
|
|
# <CEPH_CIDR> when nodes 2/3 join and those networks actually exist -
|
|
# leaving them active on a single node with no corosync/Ceph traffic is
|
|
# just dead config, and a literal `<COROSYNC_CIDR>` is invalid syntax if
|
|
# left uncommented and unfilled.
|
|
#
|
|
# Copy to /etc/pve/firewall/cluster.fw and edit before enabling (or use
|
|
# scripts/deploy-firewall.sh).
|
|
|
|
[OPTIONS]
|
|
enable: 1
|
|
policy_in: DROP
|
|
policy_out: ACCEPT
|
|
|
|
[IPSET mgmt]
|
|
<MGMT_CIDR>
|
|
|
|
[RULES]
|
|
# Web UI + SSH only from the management network
|
|
IN ACCEPT -source +mgmt -p tcp -dport 8006 -log nolog
|
|
IN ACCEPT -source +mgmt -p tcp -dport 22 -log nolog
|
|
|
|
# ICMP echo (ping) from the management network - diagnostic convenience
|
|
# only, nothing else depends on it. Without this, policy_in DROP silently
|
|
# eats ping while SSH/web UI keep working - looks like an outage during
|
|
# troubleshooting when the host is actually fine. See
|
|
# docs/06-pve-test-wifi-network.md for a case this caused real confusion.
|
|
IN ACCEPT -source +mgmt -p icmp -icmp-type echo-request -log nolog
|
|
|
|
# Stage 2: Corosync (cluster quorum) - uncomment once node 2/3 join and
|
|
# the corosync network/VLAN exists.
|
|
# IN ACCEPT -source <COROSYNC_CIDR> -p udp -dport 5404:5405 -log nolog
|
|
|
|
# Stage 2: Ceph (uncomment once Ceph is live; ports: mon 3300,6789,
|
|
# osd/mgr/mds 6800-7300)
|
|
# IN ACCEPT -source <CEPH_CIDR> -p tcp -dport 3300 -log nolog
|
|
# IN ACCEPT -source <CEPH_CIDR> -p tcp -dport 6789 -log nolog
|
|
# IN ACCEPT -source <CEPH_CIDR> -p tcp -dport 6800:7300 -log nolog
|