Archived
fix(ha/cluster-init): keep Pacemaker in standby until DRBD sync completes
Check NixOS configurations / eval-hosts (push) Successful in 10m22s
Check NixOS configurations / eval-hosts (push) Successful in 10m22s
Clearing crm_standby before the initial sync finished caused Pacemaker's OCF DRBD agent to race with the manual drbdadm up/primary calls. The agent saw DRBD in WFConnection or SyncSource and tore it down, driving the resource back to StandAlone and killing the sync in seconds. Move the crm_standby -v off calls to immediately after the sync-complete break, so Pacemaker only resumes once DRBD is UpToDate/UpToDate and safe to hand back. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -232,12 +232,13 @@ log "Bringing up DRBD on both nodes..."
|
|||||||
drbdadm up ha-data 2>/dev/null || true
|
drbdadm up ha-data 2>/dev/null || true
|
||||||
n2_ssh "drbdadm up ha-data" 2>/dev/null || true
|
n2_ssh "drbdadm up ha-data" 2>/dev/null || true
|
||||||
|
|
||||||
log "Clearing Pacemaker standby — DRBD is up, letting Pacemaker resume..."
|
|
||||||
crm_standby -N "$NODE1" -v off 2>/dev/null || true
|
|
||||||
crm_standby -N "$NODE2" -v off 2>/dev/null || true
|
|
||||||
|
|
||||||
log "Forcing $NODE1 to DRBD Primary for initial sync..."
|
log "Forcing $NODE1 to DRBD Primary for initial sync..."
|
||||||
drbdadm primary ha-data --force
|
drbdadm primary ha-data --force
|
||||||
|
# NOTE: Pacemaker standby is intentionally kept ON until after the sync
|
||||||
|
# completes. Clearing it here races with the OCF DRBD agent: Pacemaker
|
||||||
|
# sees DRBD in WFConnection/SyncSource and may call drbdadm-down thinking
|
||||||
|
# something went wrong, killing the sync. Standby is cleared below, after
|
||||||
|
# UpToDate/UpToDate is confirmed.
|
||||||
|
|
||||||
log "Waiting for DRBD initial sync to complete (32 GB may take 10–20 min)..."
|
log "Waiting for DRBD initial sync to complete (32 GB may take 10–20 min)..."
|
||||||
log " (monitor with: watch -n3 cat /proc/drbd)"
|
log " (monitor with: watch -n3 cat /proc/drbd)"
|
||||||
@@ -271,6 +272,10 @@ while true; do
|
|||||||
sleep 3
|
sleep 3
|
||||||
done
|
done
|
||||||
|
|
||||||
|
log "Clearing Pacemaker standby — sync complete, handing DRBD back to Pacemaker..."
|
||||||
|
crm_standby -N "$NODE1" -v off 2>/dev/null || true
|
||||||
|
crm_standby -N "$NODE2" -v off 2>/dev/null || true
|
||||||
|
|
||||||
# ── 3. XFS filesystem ─────────────────────────────────────────────────────
|
# ── 3. XFS filesystem ─────────────────────────────────────────────────────
|
||||||
log "Creating XFS on ${DRBD_DEVICE}..."
|
log "Creating XFS on ${DRBD_DEVICE}..."
|
||||||
if ! xfs_info "${DRBD_DEVICE}" &>/dev/null; then
|
if ! xfs_info "${DRBD_DEVICE}" &>/dev/null; then
|
||||||
|
|||||||
Reference in New Issue
Block a user