Archived
Two new NixOS Proxmox VMs (VMIDs 202/203) forming a dual-manager Docker
Swarm on dedicated vmbr3 (192.168.30.0/24, VLAN 30) for gossip and VXLAN,
with NFS via the storage-client network (vmbr2) from the existing HA cluster.
- nixos/variables.nix: add ha-docker IP/interface/port vars and swarm CIDR
- nixos/modules/build-types/ha-docker.nix: new build type — Docker 29,
NFS mounts, beszel-agent, health monitoring, swarm firewall rules with
checkReversePath = "loose" for VXLAN routing mesh
- nixos/hosts/ha-docker-{1,2}/host.nix: per-host identity — three NICs
(LAN, storage, swarm), IPA dyndns pinned to LAN interface
- nixos/flake.nix: add proxmox-ha-docker-{1,2} targets; build-validated
with nix build --dry-run (169 derivations, no errors)
- nixos/docs/ip-addressing.md: document VLAN 30 / swarm.home zone,
ha-docker IP allocations across all three subnets
- nixos/scripts/docker-swarm/deploy.sh: 10-phase lifecycle script
(bridge, keys, IPA, VMs, swarm init, DNS, verify); modelled on
scripts/ha/deploy.sh with --destroy mode
- nixos/docs/internal/docker-swarm-cutover.md: service-by-service
migration guide covering Traefik log rotation, Nextcloud cron sidecar,
docker-health-to-gotify swarm awareness updates, Passbolt/Gitea steps,
DNS cutover, and CT 105 decommission checklist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DASH15okNvWeY1rVJmyJoJ
229 lines
11 KiB
Markdown
229 lines
11 KiB
Markdown
# IP Addressing Scheme
|
||
|
||
## Subnets
|
||
|
||
| Subnet | VLAN | CIDR | Purpose | Routed? |
|
||
|---|---|---|---|---|
|
||
| LAN | 2 (native/untagged) | `192.168.2.0/24` | General LAN — clients and infrastructure | Yes (gateway .254) |
|
||
| Cluster | 10 | `192.168.10.224/29` | HA file server DRBD replication + Corosync heartbeat | No — internal `vmbr1` only, no uplink |
|
||
| Storage client | 20 | `192.168.20.0/24` | HA file server NFS (and iSCSI if needed) — docker and swarm nodes mount from VIP here | No — internal `vmbr2` only, no uplink |
|
||
| Swarm cluster | 30 | `192.168.30.0/24` | Docker Swarm gossip (TCP/UDP 7946) + VXLAN overlay (UDP 4789) | No — internal `vmbr3` only, no uplink |
|
||
|
||
When expanded to a second Proxmox node, VLAN 10 (cluster), VLAN 20 (storage-client), and VLAN 30 (swarm) all share
|
||
the same inter-node trunk NIC via 802.1q VLAN tagging — different VLAN IDs, same physical cable.
|
||
|
||
The cluster and storage-client subnets never leave pve1. `vmbr1` and `vmbr2` are Proxmox Linux
|
||
bridges with no physical port attached; traffic between guests on each bridge stays in-kernel.
|
||
|
||
VLAN IDs match the third octet of each subnet (VLAN 2 → 192.168.**2**.x, VLAN 10 → 192.168.**10**.x,
|
||
VLAN 20 → 192.168.**20**.x). The host octet is consistent across all subnets — e.g. ha-node1
|
||
is always `.228`: `192.168.2.228` (LAN), `192.168.10.228` (cluster), `192.168.20.228` (storage client).
|
||
|
||
**Protocol separation** (enforced by firewall on HA nodes):
|
||
- NFS (ports 111, 2049, 20048): both subnets, each restricted to its own CIDR
|
||
- VLAN 2 only → `vip-lan` (192.168.2.229) — pxe-boot and other LAN clients
|
||
- VLAN 20 only → `vip-storage` (192.168.20.229) — docker, future swarm nodes
|
||
- iSCSI (port 3260): VLAN 20 only — available but not in active use; NFS is preferred
|
||
for multi-host access (shared volumes across a Docker Swarm require a shared filesystem,
|
||
not per-host block devices)
|
||
|
||
---
|
||
|
||
## DNS Zones
|
||
|
||
FreeIPA (domain-controller.sweet.home) is authoritative for all zones.
|
||
|
||
Four zones correspond to the four subnets. All zones are internal only; no external delegation.
|
||
|
||
### sweet.home — VLAN 2 (192.168.2.x)
|
||
|
||
General LAN zone. All infrastructure hostnames live here.
|
||
|
||
| Hostname | A record | Notes |
|
||
|---|---|---|
|
||
| `domain-controller.sweet.home` | `192.168.2.253` | FreeIPA / KDC / DNS |
|
||
| `ha-vip-lan.sweet.home` | `192.168.2.229` | Pacemaker `vip-lan` — NFS for LAN clients |
|
||
| `ha-server-1.sweet.home` | `192.168.2.228` | HA node 1 management NIC |
|
||
| `ha-server-2.sweet.home` | `192.168.2.227` | HA node 2 management NIC |
|
||
| `server.sweet.home` | `192.168.2.226` | Current ZFS/NFS server (retiring) |
|
||
| `docker.sweet.home` | `192.168.2.225` | Docker/Traefik host |
|
||
| `nix-cache.sweet.home` | `192.168.2.224` | Nix binary cache + remote builder |
|
||
| `pxe-boot.sweet.home` | `192.168.2.223` | PXE / TFTP / HTTP netboot |
|
||
| `tailscale-router.sweet.home` | `192.168.2.222` | Tailscale exit node |
|
||
| `tor-relay.sweet.home` | `192.168.2.221` | Tor relay |
|
||
| `pdm.sweet.home` | `192.168.2.220` | Proxmox Deploy Manager |
|
||
| `nixos.sweet.home` | `192.168.2.39` | Bare-metal workstation (DHCP) |
|
||
| `pve1.sweet.home` | `192.168.2.245` | Proxmox VE hypervisor |
|
||
| `pbs.sweet.home` | `192.168.2.244` | Proxmox Backup Server |
|
||
|
||
PTR records exist for all static hosts. The workstation (`nixos.sweet.home`) is
|
||
DHCP-assigned; its PTR is omitted.
|
||
|
||
### cluster.home — VLAN 10 (192.168.10.x)
|
||
|
||
Internal only — Corosync ring0 heartbeat and DRBD replication between HA nodes.
|
||
No VIP exists on this subnet (DRBD/Corosync endpoints are static per-node IPs).
|
||
|
||
| Hostname | A record | Notes |
|
||
|---|---|---|
|
||
| `ha-server-1.cluster.home` | `192.168.10.228` | HA node 1 cluster NIC (ens19 / vmbr1) |
|
||
| `ha-server-2.cluster.home` | `192.168.10.227` | HA node 2 cluster NIC (ens19 / vmbr1) |
|
||
|
||
PTR records exist for both. DNS here is for debugging convenience — DRBD and
|
||
Corosync use the IPs from the NixOS config directly, not DNS.
|
||
|
||
### storage.home — VLAN 20 (192.168.20.x)
|
||
|
||
Internal only — NFS (and iSCSI) client access to the HA storage VIP. NFS clients
|
||
mount from **`nfs.storage.home`** (the Pacemaker floating VIP) so mounts survive
|
||
failover transparently without reconfiguration.
|
||
|
||
| Hostname | A record | Notes |
|
||
|---|---|---|
|
||
| `nfs.storage.home` | `192.168.20.229` | Pacemaker `vip-storage` — NFS + iSCSI VIP |
|
||
| `ha-server-1.storage.home` | `192.168.20.228` | HA node 1 storage-client NIC (ens20 / vmbr2) |
|
||
| `ha-server-2.storage.home` | `192.168.20.227` | HA node 2 storage-client NIC (ens20 / vmbr2) |
|
||
| `docker.storage.home` | `192.168.20.225` | Docker host storage-client NIC (eth1 / vmbr2) |
|
||
| `server.storage.home` | `192.168.20.226` | server VM storage-client NIC (decommissioned — remove DNS record after VM is destroyed) |
|
||
|
||
PTR records exist for all five. Remove `server.storage.home`, `server.sweet.home`,
|
||
and their PTRs from FreeIPA DNS once the server VM is destroyed.
|
||
|
||
---
|
||
|
||
## LAN — 192.168.2.0/24
|
||
|
||
### Address map
|
||
|
||
| Range | Purpose |
|
||
|---|---|
|
||
| .1–.9 | Reserved, never assign |
|
||
| .10–.59 | Client DHCP pool (router-assigned) |
|
||
| .60–.219 | Unallocated buffer |
|
||
| .220–.229 | Virtual nodes (VMs / LXC containers) |
|
||
| .230–.239 | Expansion buffer (reserved, unallocated) |
|
||
| .240–.249 | Physical nodes (bare-metal hosts) |
|
||
| .250–.253 | Network services |
|
||
| .254 | Router / gateway |
|
||
|
||
### Network services (.250–.253)
|
||
|
||
| IP | Hostname | Role |
|
||
|---|---|---|
|
||
| `192.168.2.254` | router | Gateway (TP-Link) |
|
||
| `192.168.2.253` | domain-controller | FreeIPA — authoritative DNS for `sweet.home`, Kerberos, LDAP |
|
||
| `192.168.2.250`–`.252` | — | Reserved for future network services |
|
||
|
||
### Physical nodes (.240–.249)
|
||
|
||
| IP | Hostname | Role |
|
||
|---|---|---|
|
||
| `192.168.2.245` | pve1 | Proxmox VE hypervisor |
|
||
| `192.168.2.244` | pbs | Proxmox Backup Server |
|
||
| `192.168.2.243` | nixos | Bare-metal workstation (`baremetal-gui`) |
|
||
| `192.168.2.246`–`.249` | — | Reserved — second Proxmox node and associated services |
|
||
| `192.168.2.240`–`.242` | — | Reserved |
|
||
|
||
pve1 sits mid-range deliberately so a second Proxmox node can slot in on either side.
|
||
|
||
### Virtual nodes (.220–.229)
|
||
|
||
All VMs and LXC containers run on pve1.
|
||
|
||
| IP | Hostname | Role | Status |
|
||
|---|---|---|---|
|
||
| `192.168.2.229` | ha-vip-lan | HA file server LAN floating VIP (Pacemaker `vip-lan`) — LAN iSCSI + NFS | Active |
|
||
| `192.168.2.228` | ha-node1 | HA file server node 1 — management NIC | Active |
|
||
| `192.168.2.227` | ha-node2 | HA file server node 2 — management NIC | Active |
|
||
| `192.168.2.226` | server | Former NFS/ZFS file server — decommissioned | Removed from flake |
|
||
| `192.168.2.225` | docker | Docker / Traefik stack (CT 105 — existing single-host) | Active |
|
||
| `192.168.2.224` | nix-cache | Nix binary cache + remote builder | Active |
|
||
| `192.168.2.223` | pxe-boot | PXE / TFTP / HTTP netboot server | Active |
|
||
| `192.168.2.222` | tailscale-router | Tailscale exit node / router | Active |
|
||
| `192.168.2.221` | tor-relay | Tor relay | Active |
|
||
| `192.168.2.220` | pdm | Proxmox Deploy Manager | Active |
|
||
| `192.168.2.231` | ha-docker-2 | Docker Swarm node 2 — management NIC | Active |
|
||
| `192.168.2.230` | ha-docker-1 | Docker Swarm node 1 — management NIC | Active |
|
||
|
||
### Client DHCP pool (.10–.59)
|
||
|
||
Assigned by the router. DNS option points to `192.168.2.253` (domain-controller).
|
||
|
||
Devices in this range: phones, laptops, IoT, Canon printer, any non-infrastructure host.
|
||
No static reservations for infrastructure hosts — all infra uses static IP configuration
|
||
on the guest itself (not DHCP reservations), so IPs survive VM recreation regardless of
|
||
MAC address churn.
|
||
|
||
---
|
||
|
||
## Cluster network — VLAN 10 — 192.168.10.224/29
|
||
|
||
Internal to pve1 only. Proxmox bridge `vmbr1`, no physical NIC attached.
|
||
|
||
| IP | Hostname | Interface role |
|
||
|---|---|---|
|
||
| `192.168.10.228` | ha-node1 | DRBD replication + Corosync ring0 (primary heartbeat) |
|
||
| `192.168.10.227` | ha-node2 | DRBD replication + Corosync ring0 (primary heartbeat) |
|
||
| — | no gateway | Isolated — not routed to LAN or internet |
|
||
|
||
Corosync ring1 (backup heartbeat only) uses the LAN IPs (`192.168.2.228` / `192.168.2.227`)
|
||
over `vmbr0` — no additional bridge needed, and DRBD traffic never crosses ring1.
|
||
|
||
---
|
||
|
||
## Storage-client network — VLAN 20 — 192.168.20.0/24
|
||
|
||
Internal to pve1 only. Proxmox bridge `vmbr2`, no physical NIC attached.
|
||
|
||
| IP | Hostname | Interface / role |
|
||
|---|---|---|
|
||
| `192.168.20.229` | ha-vip-storage | Pacemaker floating VIP — NFS + iSCSI endpoint |
|
||
| `192.168.20.228` | ha-node1 | Storage-client NIC (ens20 / vmbr2) |
|
||
| `192.168.20.227` | ha-node2 | Storage-client NIC (ens20 / vmbr2) |
|
||
| `192.168.20.226` | server | Storage-client NIC (ens19 / vmbr2) — decommissioned |
|
||
| `192.168.20.225` | docker | Storage-client NIC (eth1 / vmbr2) — NFS client (CT 105) |
|
||
| `192.168.20.231` | ha-docker-2 | Storage-client NIC (ens19 / vmbr2) — NFS client |
|
||
| `192.168.20.230` | ha-docker-1 | Storage-client NIC (ens19 / vmbr2) — NFS client |
|
||
| — | no gateway | Isolated — not routed to LAN or internet |
|
||
|
||
NFS clients mount from `192.168.20.229` (surviving failover transparently via the VIP).
|
||
Firewall on each HA node restricts NFS and iSCSI ports to `192.168.20.0/24` — LAN hosts
|
||
cannot reach either service on this VIP. The `vip-storage` endpoint is not reachable
|
||
from the workstation directly (internal bridge only); health checks proxy through the
|
||
active HA node.
|
||
|
||
---
|
||
|
||
## Swarm cluster network — VLAN 30 — 192.168.30.0/24
|
||
|
||
Internal to pve1 only. Proxmox bridge `vmbr3`, no physical NIC attached.
|
||
Carries Docker Swarm inter-node traffic only: Raft consensus (TCP 2377),
|
||
Serf gossip (TCP/UDP 7946), and VXLAN overlay data path (UDP 4789).
|
||
Docker Swarm is initialised with `--advertise-addr` and `--data-path-addr`
|
||
both pointing to this subnet so all cluster traffic stays on `vmbr3` and
|
||
never crosses the LAN.
|
||
|
||
| IP | Hostname | Interface / role |
|
||
|---|---|---|
|
||
| `192.168.30.231` | ha-docker-2 | Swarm cluster NIC (ens20 / vmbr3) |
|
||
| `192.168.30.230` | ha-docker-1 | Swarm cluster NIC (ens20 / vmbr3) |
|
||
| — | no gateway | Isolated — not routed to LAN or internet |
|
||
|
||
### DNS zone: `swarm.home` — VLAN 30 (192.168.30.x)
|
||
|
||
| Hostname | A record | Notes |
|
||
|---|---|---|
|
||
| `ha-docker-1.swarm.home` | `192.168.30.230` | Swarm NIC — debugging only |
|
||
| `ha-docker-2.swarm.home` | `192.168.30.231` | Swarm NIC — debugging only |
|
||
|
||
Operators reach the Docker API on the LAN IPs (`192.168.2.230`/`.231`), not these addresses.
|
||
The `swarm.home` records exist for diagnostic convenience (e.g. confirming `vmbr3` routing).
|
||
|
||
### Multi-node Proxmox expansion
|
||
|
||
When a second Proxmox node (pve2) is added, VLAN 10 (cluster), VLAN 20 (storage-client),
|
||
and VLAN 30 (swarm) all extend to pve2 via 802.1q VLAN tagging on the inter-node trunk
|
||
link. All three internal networks share the same physical NIC between hypervisors —
|
||
VLAN tags provide the logical separation.
|
||
|