Andrew Mercer
on this page

Galera Cluster: Troubleshooting

Errors encountered across both the Docker and bare metal setups. For failures severe enough that a node or the whole cluster won't come back up, see Crash Recovery instead.

"It may not be safe to bootstrap the cluster from this node"

WSREP: It may not be safe to bootstrap the cluster from this node. It was not the last one to leave the cluster and may not contain all the updates. To force cluster bootstrap with this node, edit the grastate.dat file manually and set safe_to_bootstrap to 1

This is Galera protecting you from bootstrapping off a node that might not have the cluster's latest data. Find the file:

docker inspect galera-node1 | grep Mounts -A 5
sudo vi /var/lib/docker/volumes/galera-node1-data/_data/grastate.dat

Only set safe_to_bootstrap=1 if you're confident this node actually has the most recent data — see Crash Recovery: Determining the Most Advanced Node before doing this. Going forward, make sure the node that originally bootstrapped the cluster is also the last one to leave it (i.e., the last one stopped), so this situation doesn't recur.

TODO: confirm which node actually needs to bootstrap in a given incident — in practice it's not always the one that was used to bootstrap originally.

"Failed to prepare for incremental state transfer: Local state UUID does not match group state UUID"

[ERROR] WSREP: failed to open gcomm backend connection: 131: [some_uuid] last prims not consistent (FATAL)

Make sure port 4444 (the rsync SST daemon) is open on every instance's firewall — this is the most common cause. See Bare Metal Deployment: Firewall Configuration.

systemd Reports a Failed Start, But MariaDB Is Actually Running

Job for mariadb.service failed because a fatal signal was delivered to the control process. See "systemctl status mariadb.service" and "journalctl -xe" for details.

When issuing sudo systemctl start mariadb, systemd may report failure — but sudo systemctl status mariadb or ps -ef | grep -i mysql shows it's actually running. Running sudo systemctl restart mariadb will then report success and status correctly. Treat the first failure report with suspicion and re-check actual process state before assuming the start genuinely failed.

SELinux Is Preventing mysqld from Open Access

setroubleshoot: SELinux is preventing mysqld from open access on the file /tmp/wsrep_recovery.e86zDL.

Investigate the specific denial:

sealert -l <alert-id-from-the-log>

The standard remediation (generate and load a local policy module) is:

ausearch -c 'mysqld' --raw | audit2allow -M my-mysqld
semodule -i my-mysqld.pp

Related reading: - http://galeracluster.com/documentation-webpages/arbitrator.html - https://www.percona.com/files/presentations/percona-live/nyc-2012/PLNY12-galera-cluster-best-practices.pdf - https://access.redhat.com/articles/2332651

Node Consistently Shows Out of Sync (OpenStack-Managed Environments)

If a node keeps reporting out of sync with the rest of the cluster, check for a stray manually-started mysqld_safe process still running underneath — this can happen after a manual recovery attempt that wasn't fully cleaned up:

root       29597  0.0  0.0  11816  1120 ?        S    17:16   0:00 /bin/sh /usr/bin/mysqld_safe --defaults-file=/etc/my.cnf ...
mysql      31577  0.5  0.0 808108 124568 ?       Sl   17:16   0:26 /usr/libexec/mysqld --defaults-file=/etc/my.cnf ...

Full recovery sequence for this case (do this in a maintenance window — it takes MySQL down):

# Disable and unmanage Galera so Pacemaker stops touching it
pcs resource disable galera
pcs resource unmanage galera

# Disable and unmanage Haproxy so it stops trying to write during recovery
pcs resource disable haproxy
pcs resource unmanage haproxy

# Bring the affected node back into the Pacemaker cluster
pcs cluster unstandby controller-2

# Confirm mysqld is actually down everywhere
mysqladmin shutdown
ps -elf | grep mysql

# Re-add the node to the Galera address list
pcs resource update galera wsrep_cluster_address=gcomm:///controller0,controller1,controller2

# Hand Galera back to Pacemaker
pcs resource manage galera
pcs resource enable galera

# Confirm
clustercheck

# Hand Haproxy back to Pacemaker
pcs resource manage haproxy
pcs resource enable haproxy

# Clear any stale failcounts
pcs resource cleanup

Known related issues worth checking if problems persist: a reported memory leak in the OpenVSwitch Python wrapper can require a neutron service restart to fully clear up an unhealthy node — see RHBZ#1693430 and RHBA-2019:0738.

See Also

  • Crash Recovery — for full-node and full-cluster failure, not just config/permission errors
  • Pacemaker HA — FAILED Master (blocked) and other Pacemaker-specific resource states