9. Legacy appendix — Heartbeat / openais / crm shell (historical reference only)¶
This section covers the older Ubuntu 10.04–12.04 stack (crm shell, openais
compatibility mode, DRBD for shared storage) from the OSCAR-cluster notes. It's
obsolete for new builds — use §2–§6 instead — but preserved here because the DRBD
split-brain recovery steps in particular are still useful if you ever touch one of
these old boxes again.
Old-style corosync.conf (openais compatibility, multicast)¶
totem {
version: 2
secauth: off
interface {
ringnumber: 0
bindnetaddr: 10.0.1.0
mcastaddr: 226.94.1.1
mcastport: 5405
}
}
service {
ver: 1
name: pacemaker
}
Generate the shared authkey once, on the primary node only, then copy it to the peer(s):
corosync-keygen # needs entropy — bang the keyboard, or generate load
# (e.g. `dd if=/dev/urandom of=/tmp/x bs=1M count=100`
# in a loop) rather than downloading a file to /dev/null
scp /etc/corosync/authkey /etc/default/corosync /etc/corosync/corosync.conf \
user@peer:~
Old crm shell equivalents of the modern pcs commands above¶
crm configure property stonith-enabled=false
crm configure property no-quorum-policy=ignore
crm_verify -L
crm configure primitive ClusterIP ocf:heartbeat:IPaddr2 \
params ip=172.16.100.69 cidr_netmask=32 \
op monitor interval=30s
crm configure show # view running config
crm_mon # live status (like `pcs status`)
crm resource migrate <resource> <node> # like `pcs resource move`
crm configure rsc_defaults resource-stickiness=100
Deleting a stale node from an old cluster:
crm_node -l # list nodes incl. stale/"lost" entries
crm_node -R <nodeid> --force # remove a stale one
# or, simpler:
crm
crm(live)# node
crm(live)node# delete <node_name>
DRBD for shared/replicated storage (deduplicated from 3 near-identical copies)¶
Install:
aptitude install drbd8-utils # Ubuntu / old Debian-family
yum install drbd84-utils kmod-drbd84 # CentOS/RHEL (needs ELRepo)
Resource definition (/etc/drbd.d/diskN.res), one block per DRBD device:
resource disk0 {
protocol C;
syncer { rate 100M; }
startup { wfc-timeout 15; degr-wfc-timeout 60; }
net {
cram-hmac-alg sha1;
shared-secret "some-shared-secret";
}
on node1 { device /dev/drbd0; disk /dev/vdb2; address 10.0.1.2:7788; meta-disk /dev/vdb1[0]; }
on node2 { device /dev/drbd0; disk /dev/vdb2; address 10.0.1.3:7788; meta-disk /dev/vdb1[0]; }
}
Copy the same .res files to every node, then on all nodes:
drbdadm create-md disk0
drbdadm up disk0
On the node that should start as primary:
drbdadm -- --overwrite-data-of-peer primary all
watch -n1 cat /proc/drbd # watch the initial sync complete
mkfs.ext4 /dev/drbd0
Fixing DRBD split-brain (the same recovery procedure appeared three times verbatim in the source notes — this is the deduplicated version):
- Confirm it's split-brain:
bash cat /proc/drbd grep -i split-brain /var/log/messagesBoth sides showingPrimary(or oneStandAlone) is the tell. - Pick a "victim" — whichever side has data you're willing to discard. On the
victim:
bash drbdadm disconnect <resource> drbdadm secondary <resource> drbdadm -- --discard-my-data connect <resource> # pre-8.4 syntax # drbdadm connect --discard-my-data <resource> # 8.4+ syntax - On the survivor:
bash drbdadm connect <resource> - If status shows
Diskless/UpToDateafterward, also run on the victim:bash drbdadm attach <resource> crm node online <node_name> # (or `pcs cluster unstandby` on modern stacks) drbdadm connect <resource> - If status shows
WFConnection/Outdated/DUnknownon both sides, follow the same disconnect → discard-my-data → connect sequence — the underlying fix is identical regardless of which specific symptom string DRBD reports.
Managing DRBD as a Pacemaker resource (old ms/master-slave syntax — modern pcs
uses pcs resource master against an ocf:linbit:drbd primitive, same idea):
primitive mysql-drbd ocf:linbit:drbd \
params drbd_resource="disk0" \
op monitor interval="15s" start interval="0" timeout="240" stop interval="0" timeout="100"
ms ms_mysql_drbd mysql-drbd \
meta master-max="1" master-node-max="1" clone-max="2" clone-node-max="1" notify="true"
Removing a resource from cluster management temporarily¶
crm resource unmanage <resource> # old
pcs resource unmanage <resource> # modern equivalent
# ... do manual maintenance ...
crm resource manage <resource>
pcs resource manage <resource>
← Part 06: OpenStack Example · Part 08: Building from Source & References →