Galera Cluster: Pacemaker HA Integration¶
Pacemaker/Corosync can manage Galera as a multi-state (Master/Slave) resource, automating the bootstrap-node election that you'd otherwise do by hand (see Crash Recovery). This page assumes a Pacemaker cluster of node1, node2, node3 with the Galera resource named galera.
Important terminology note: Pacemaker's Master/Slave resource states are not the same thing as Galera's own Primary/Non-primary concept (see Overview: Quorum and Cluster States). A node can be Pacemaker-promoted to "Master" while Galera itself considers it Non-primary — that's an error condition, not a contradiction to shrug off.
How the Resource Agent Boots the Cluster¶
The resource agent breaks cluster boot into discrete steps, tracked via Pacemaker's Cluster Information Base (CIB):
- When the
galeraresource isStarted, the agent retrieves the local node's lastseqnofrom MariaDB, stores it in the CIB, and moves toSlavestate. No Galera server is running yet at this point. - Once all nodes report into
Slavestate, the agent elects a bootstrap node (the one with the highest seqno), tags it in the CIB, and tells Pacemaker it may promote that node's resource toMaster. - When Pacemaker promotes the bootstrap node, the agent starts the Galera server there, which founds the new cluster. The agent then marks the remaining nodes ready for promotion. The resource on the bootstrap node moves to
Master. - Pacemaker promotes the remaining nodes one by one. Each triggers a Galera server start that syncs via SST — this can take a while. Promotion to
Mastercompletes only once sync finishes.
At the end of this, every node's galera resource shows Master, and the cluster is fully up.
The catch: this entire process depends on the agent being able to retrieve seqno from every node. If one node is unreachable or in a bad state, no bootstrap node can be elected and Pacemaker will not boot the cluster on its own — this is when manual override becomes necessary.
Overriding the Boot Process¶
⚠️ Only do the following if you are certain the node you're forcing is actually up to date. Bootstrapping from a stale node will permanently desynchronize the cluster and lose data — the automated process exists specifically to prevent this mistake.
Scenario 1: Cluster Restart Blocked Because One Node Won't Come Up¶
Say node3 is unavailable (crashed inconsistent state, hardware failure). The agent can't retrieve all seqno values, so it's stuck. To unblock it:
Take Galera out of Pacemaker's control:
pcs resource unmanage galera
Find the highest seqno. If Pacemaker already attempted a restart, it may be in the CIB:
crm_attribute -N node1 -l reboot --name galera-last-committed -Q
Otherwise, recover it directly from MariaDB:
mysqld_safe --wsrep-recover
151002 13:59:50 mysqld_safe WSREP: Recovered position 4c7ba2a8-566a-11e5-8250-1e939ac17c77:9
Once you've identified the node with the highest seqno (assume node1 here), force it into the bootstrap role:
crm_attribute -N node1 -l reboot --name galera-bootstrap -v true
crm_attribute -N node1 -l reboot --name master-galera -v 100
crm_resource --force-promote -r galera -V
--force-promote requires Pacemaker ≥ 1.1.13.
Tell Pacemaker to re-probe resource state and clear stale failure history:
pcs resource refresh galera
(On RHEL 7.4 and earlier, use pcs resource cleanup galera instead — pcs resource cleanup only re-probes resources already showing as failed, whereas refresh always re-probes.)
Hand control back to Pacemaker so the remaining nodes can join automatically:
pcs resource enable galera
pcs resource manage galera
Scenario 2: Multiple Hardware Failures, Keep Service on the Remaining Node¶
If node2 and node3 fail in sequence, leaving only node1, Galera's own rules mean the survivor is inquorate — Non-primary — even though it's the only node left. This is treated as an error by the resource agent, and Pacemaker will stop the Galera server on the remaining node rather than leave it serving traffic alone.
Apply the same steps as Scenario 1 to force node1 to bootstrap (again, only if you're confident it's up to date), then also check whether the Pacemaker layer itself has quorum before handing control back — see Crash Recovery: Multiple Hardware Failures for the corosync-quorumtool steps, since Corosync quorum and Galera quorum are separate checks that both need to pass.
FAILED Master (blocked)¶
Master/Slave Set: galera-master [galera]
galera (ocf::heartbeat:galera): FAILED Master controller0 (blocked)
galera (ocf::heartbeat:galera): FAILED Master controller1 (blocked)
Masters: [ controller2 ]
Reference: access.redhat.com/solutions/4795311.
pcs resource cleanup
clustercheck
If that doesn't clear it, the node is likely genuinely out of sync — see Troubleshooting: Node Consistently Shows Out of Sync for the full manual recovery sequence used in OpenStack-managed environments.
Related¶
- Crash Recovery — the manual, non-Pacemaker version of these same procedures
- Troubleshooting — the OpenStack stray-
mysqld_safescenario referenced above