Archived
pve-test was briefly clustered with pve1 then deliberately de-clustered so it could move to wifi-primary networking (4addr bridge mode, bonded with a wired LAN backup) - a change not achievable while clustered given corosync's latency requirements. Captures that as a reproducible script plus docs: cluster separation procedure, the wifi network design and the live-cutover pitfalls hit along the way, and node-role/history context. Also adds CLAUDE.md guardrails for pve1 (production) vs pve-test (sandbox) - this repo had none before, despite scripts here being able to make real changes to both. Separately: both nodes' mgmt firewalls were dropping ICMP by default (TCP 22/8006 only), which looked like an outage mid-troubleshooting even though SSH/web UI were fine. Added an explicit ping-allow rule to the firewall template, applied it live on both nodes, and added an audit.sh check so it stays enforced.
84 lines
4.1 KiB
Markdown
84 lines
4.1 KiB
Markdown
# Node Roles: pve1 and pve-test
|
|
|
|
Two Proxmox VE hosts exist on the LAN. They are **not interchangeable** —
|
|
different purpose, different risk tolerance, different guardrails. See
|
|
`CLAUDE.md` for the actual rules Claude follows; this doc is the
|
|
background/context those rules assume.
|
|
|
|
## pve1 (production)
|
|
|
|
ASUS PN53 mini PC (Ryzen 7 7735HS, 32GB RAM) — see `01-hardware-node1.md`
|
|
for the full hardware layout. Stage 1 base hardening applied and verified
|
|
(`04-security-hardening.md`).
|
|
|
|
Runs real production workloads:
|
|
|
|
- `server-nixos` (VM) and LXC containers `pihole`, `claude`, `nix-cache`,
|
|
`docker`, `pdm`, `pxe-boot` — DNS, this Claude Code environment itself,
|
|
the LAN's Nix binary cache, etc.
|
|
|
|
Off-limits to Claude by default — see `CLAUDE.md`.
|
|
|
|
## pve-test (sandbox)
|
|
|
|
A separate physical node kept for testing. Has a capable GPU (earmarked
|
|
for future GPU workloads — transcoding, local inference, etc. — nothing
|
|
built on it yet as of this writing). No production guests run here.
|
|
|
|
Safe for Claude to create/destroy scratch VMs and containers on without
|
|
asking — see `CLAUDE.md` for the exact boundary (node-level config is
|
|
still not free-for-all even here).
|
|
|
|
### History: briefly clustered with pve1, then deliberately separated
|
|
|
|
On 2026-07-21/22, pve-test was joined to pve1 in a 2-node Proxmox cluster
|
|
(corosync + shared `/etc/pve`) for a period, then the cluster was
|
|
**intentionally destroyed** at the operator's explicit request — not a
|
|
failure or accident. Rationale: the operator wanted pve-test moved onto
|
|
wifi networking (see `06-pve-test-wifi-network.md`), and wifi is
|
|
fundamentally unsuitable for corosync's latency/jitter requirements (token
|
|
timeouts, flapping membership). Rather than fight that, the cluster was
|
|
torn down first, then wifi was configured on the now-standalone node.
|
|
|
|
Teardown procedure used (safe to repeat if pve-test or pve1 is ever
|
|
reclustered and needs separating again):
|
|
|
|
1. Un-manage any HA-assigned resources first if HA is in use
|
|
(`ha-manager remove <sid>` per resource) — **critical**: if a node has
|
|
HA-managed guests and its watchdog is still armed when it loses
|
|
quorum, HA's self-fence behavior can force-reboot it. Removing the
|
|
resource assignments first (not stopping the guests — just
|
|
un-managing them) prevents this.
|
|
2. Stop `pve-ha-lrm`/`pve-ha-crm` on all nodes once no resources remain
|
|
assigned, confirm via `ha-manager status` that fencing is no longer
|
|
armed.
|
|
3. From a node that will remain in the cluster: `pvecm delnode <name>`
|
|
for the node being removed. In a 2-node cluster with no QDevice, this
|
|
can itself cause the *remaining* node to transiently lose quorum
|
|
(killing a node's live corosync membership drops total votes below
|
|
expected votes before the nodelist file can be rewritten to match) —
|
|
this is expected, not a failure; proceed to step 4 regardless.
|
|
4. On **each** node being separated (including one you just ran `delnode`
|
|
from, if it's also being decoupled): `systemctl stop pve-cluster
|
|
corosync`, `pmxcfs -l` (local mode, bypasses the quorum requirement
|
|
for local writes), remove `/etc/pve/corosync.conf` and
|
|
`/etc/corosync/*`, `killall pmxcfs`, `systemctl start pve-cluster`.
|
|
This is the official "separate a node without reinstalling" procedure
|
|
and doesn't touch any running VM/CT — they're independent OS
|
|
processes, unaffected by corosync/pmxcfs state either way.
|
|
5. Optional cosmetic cleanup: the node(s) staying up may retain a stale
|
|
`/etc/pve/nodes/<removed-node>/` directory (cached VM/CT configs from
|
|
when they were clustered) — safe to `rm -rf` once separation is
|
|
confirmed; it's dead data, not a live reference to anything.
|
|
|
|
Guests were never at risk during this procedure (confirmed: all 7 of
|
|
pve1's guests ran continuously throughout) — the only real hazard is the
|
|
HA/watchdog self-fence path in step 1, and losing corosync-provided quorum
|
|
temporarily blocking `/etc/pve` writes (not guest execution) in step 3.
|
|
|
|
### Current state: standalone, wifi-primary networking
|
|
|
|
pve-test now runs entirely independently of pve1, on the network design
|
|
documented in `06-pve-test-wifi-network.md`. **Do not re-cluster
|
|
pve-test while it's on wifi** — see that doc for why.
|