Andrew Mercer
on this page

Galera Cluster: Crash Recovery

Source: MariaDB Cluster Admin Deep Dive course notes, plus Restarting a Galera Cluster and a symmcom recovery walkthrough covering the same ground. Read Overview: Global Transaction ID first if seqno and grastate.dat aren't already familiar.

Severity Levels

One node crashed — restart it; it rejoins the cluster automatically via IST or SST as needed. No manual intervention required.

Multiple nodes crashed, one survives — no data is lost, but the survivor alone doesn't have quorum, so it can't serve writes until you force the issue (see below).

All nodes crashed — every node needs to be evaluated to find the one with the most complete data, and that node bootstraps the cluster.

Determining the Most Advanced Node

Check grastate.dat on each node — the one with the highest seqno and safe_to_bootstrap=1 should be bootstrapped first:

cat /var/lib/mysql/grastate.dat

If grastate.dat doesn't have a usable seqno (e.g. the node crashed hard enough that it wasn't flushed), recover the position directly from the storage engine:

sudo -u mysql mysqld --wsrep_recover

Look for a line like:

WSREP: Recovered position: [ uuid ]:[ seq_number ]

Run this on every candidate node and compare seq_number — the highest one is your bootstrap node. (Reference: Running mysqld as root.)

You can also check the in-database view of the last committed transaction, if the node is up enough to query:

show status like 'wsrep_last_committed';

Bootstrapping After a Full Crash

On the node identified as most advanced:

systemctl stop mariadb

In galera.cnf on that node only:

wsrep_cluster_address=gcomm://
systemctl start mariadb

Confirm it came up alone and Primary:

show global status where Variable_name IN ('wsrep_ready', 'wsrep_cluster_size', 'wsrep_cluster_status', 'wsrep_connected');
-- expect: cluster_size = 1, cluster_status = Primary

Then bring the remaining nodes up normally, pointed at the full peer list, so they sync against this node via SST/IST.

Scenario: Multiple Hardware Failures, One Node Surviving

If two of three nodes fail unexpectedly, the survivor is inquorate by Galera's own rules — even if Pacemaker considers it fine — and Galera will switch it to Non-primary state, causing Pacemaker to stop it. To force it back up (only if you're confident this node has the latest data — otherwise you risk data loss):

  1. Take the node through the same manual-bootstrap steps as Pacemaker HA: Manual Override.
  2. Check whether the Pacemaker layer itself has quorum:

bash corosync-quorumtool -s

If it reports Quorate: No, temporarily unblock quorum to the number of nodes actually online:

bash corosync-quorumtool -e1

This setting is not permanent — once other nodes rejoin, expected votes revert to the original count automatically.

  1. Once quorate, hand control back to Pacemaker:

bash pcs resource manage galera

Scenario: No Master Nodes (Two Replicas, One Offline)

Manual recovery path when Pacemaker/crm_attribute state isn't available or trusted:

  1. On every node, check the last committed transaction:

bash cat /var/lib/mysql/grastate.dat

Whichever node has the largest seqno wins.

  1. Start MySQL manually on the new primary (bootstrap):

bash cd /usr/libexec && mysqld_safe --wsrep_new_cluster gcomm://

  1. Start MySQL manually on the other two nodes, pointed at the full peer list:

bash cd /usr/libexec && mysqld_safe --wsrep_new_cluster gcomm://[ host1 ],[ host2 ],[ host3 ]

  1. Verify sync in new terminals:

bash clustercheck

or any of the Monitoring status checks.

  1. Once synced, stop all nodes and restart them the normal way (not via manual mysqld_safe):

bash mysqladmin stop ps -ef | grep -i mysql # confirm nothing lingering systemctl start galera

  1. For Pacemaker-managed nodes, hand control back via the resource, not systemd directly:

bash pcs resource enable galera pcs resource cleanup

  • Troubleshooting — config and permission errors that stop short of a full crash
  • Pacemaker HA — how the resource agent automates most of the above when Galera is under Pacemaker