Andrew Mercer
on this page

9. Legacy appendix — Heartbeat / openais / crm shell (historical reference only)

This section covers the older Ubuntu 10.04–12.04 stack (crm shell, openais compatibility mode, DRBD for shared storage) from the OSCAR-cluster notes. It's obsolete for new builds — use §2–§6 instead — but preserved here because the DRBD split-brain recovery steps in particular are still useful if you ever touch one of these old boxes again.

Old-style corosync.conf (openais compatibility, multicast)

totem {
    version: 2
    secauth: off
    interface {
        ringnumber: 0
        bindnetaddr: 10.0.1.0
        mcastaddr: 226.94.1.1
        mcastport: 5405
    }
}
service {
    ver:  1
    name: pacemaker
}

Generate the shared authkey once, on the primary node only, then copy it to the peer(s):

corosync-keygen           # needs entropy — bang the keyboard, or generate load
                            # (e.g. `dd if=/dev/urandom of=/tmp/x bs=1M count=100`
                            # in a loop) rather than downloading a file to /dev/null
scp /etc/corosync/authkey /etc/default/corosync /etc/corosync/corosync.conf \
    user@peer:~

Old crm shell equivalents of the modern pcs commands above

crm configure property stonith-enabled=false
crm configure property no-quorum-policy=ignore
crm_verify -L

crm configure primitive ClusterIP ocf:heartbeat:IPaddr2 \
    params ip=172.16.100.69 cidr_netmask=32 \
    op monitor interval=30s

crm configure show                       # view running config
crm_mon                                  # live status (like `pcs status`)
crm resource migrate <resource> <node>   # like `pcs resource move`
crm configure rsc_defaults resource-stickiness=100

Deleting a stale node from an old cluster:

crm_node -l                              # list nodes incl. stale/"lost" entries
crm_node -R <nodeid> --force              # remove a stale one
# or, simpler:
crm
crm(live)# node
crm(live)node# delete <node_name>

DRBD for shared/replicated storage (deduplicated from 3 near-identical copies)

Install:

aptitude install drbd8-utils         # Ubuntu / old Debian-family
yum install drbd84-utils kmod-drbd84 # CentOS/RHEL (needs ELRepo)

Resource definition (/etc/drbd.d/diskN.res), one block per DRBD device:

resource disk0 {
    protocol C;
    syncer { rate 100M; }
    startup { wfc-timeout 15; degr-wfc-timeout 60; }
    net {
        cram-hmac-alg sha1;
        shared-secret "some-shared-secret";
    }
    on node1 { device /dev/drbd0; disk /dev/vdb2; address 10.0.1.2:7788; meta-disk /dev/vdb1[0]; }
    on node2 { device /dev/drbd0; disk /dev/vdb2; address 10.0.1.3:7788; meta-disk /dev/vdb1[0]; }
}

Copy the same .res files to every node, then on all nodes:

drbdadm create-md disk0
drbdadm up disk0

On the node that should start as primary:

drbdadm -- --overwrite-data-of-peer primary all
watch -n1 cat /proc/drbd     # watch the initial sync complete
mkfs.ext4 /dev/drbd0

Fixing DRBD split-brain (the same recovery procedure appeared three times verbatim in the source notes — this is the deduplicated version):

  1. Confirm it's split-brain: bash cat /proc/drbd grep -i split-brain /var/log/messages Both sides showing Primary (or one StandAlone) is the tell.
  2. Pick a "victim" — whichever side has data you're willing to discard. On the victim: bash drbdadm disconnect <resource> drbdadm secondary <resource> drbdadm -- --discard-my-data connect <resource> # pre-8.4 syntax # drbdadm connect --discard-my-data <resource> # 8.4+ syntax
  3. On the survivor: bash drbdadm connect <resource>
  4. If status shows Diskless/UpToDate afterward, also run on the victim: bash drbdadm attach <resource> crm node online <node_name> # (or `pcs cluster unstandby` on modern stacks) drbdadm connect <resource>
  5. If status shows WFConnection / Outdated/DUnknown on both sides, follow the same disconnect → discard-my-data → connect sequence — the underlying fix is identical regardless of which specific symptom string DRBD reports.

Managing DRBD as a Pacemaker resource (old ms/master-slave syntax — modern pcs uses pcs resource master against an ocf:linbit:drbd primitive, same idea):

primitive mysql-drbd ocf:linbit:drbd \
    params drbd_resource="disk0" \
    op monitor interval="15s" start interval="0" timeout="240" stop interval="0" timeout="100"
ms ms_mysql_drbd mysql-drbd \
    meta master-max="1" master-node-max="1" clone-max="2" clone-node-max="1" notify="true"

Removing a resource from cluster management temporarily

crm resource unmanage <resource>   # old
pcs resource unmanage <resource>   # modern equivalent
# ... do manual maintenance ...
crm resource manage <resource>
pcs resource manage <resource>

← Part 06: OpenStack Example · Part 08: Building from Source & References →