Andrew Mercer
on this page

Galera Cluster Overview

Galera is a synchronous, multi-master replication plugin for MySQL and MariaDB. Every node in a Galera cluster can accept writes, and every committed write is certified and applied across the whole cluster before the client gets a commit acknowledgement. This is fundamentally different from traditional MySQL/MariaDB replication (async or semi-sync, single-writer), and the operational model — bootstrapping, quorum, node states — reflects that.

This page covers the concepts referenced throughout the rest of this section. The deployment guides (Docker, bare metal) assume you've read this first.

The wsrep API and Provider Library

Galera itself isn't a database — it's a replication plugin that hooks into MySQL/MariaDB via the wsrep API (Write-Set Replication). dlopen() is used at runtime to load the wsrep provider library (libgalera_smm.so) into the database process, which is what wsrep_provider points to in the config.

The Galera Replication Plugin is layered: - Certification layer — prepares write-sets and runs certification checks to ensure they can be applied without conflict - Replication layer — manages the replication protocol and provides total ordering of transactions - Group communication framework — a plugin architecture for the underlying transport (TCP/UDP/multicast); this is also what produces the Global Transaction ID

Global Transaction ID (GTID)

Every committed write advances a cluster-wide position, tracked as two parts: - State UUID — a unique identifier for the cluster's state lineage - Seqno — a 64-bit signed integer marking the transaction's position in the sequence

Together, UUID:seqno identifies exactly where a node's data stands relative to the rest of the cluster. This pair is what you check in grastate.dat or via wsrep_recover when deciding which node should bootstrap after a crash — see Crash Recovery.

Quorum and Cluster States

Galera nodes track membership through a quorum mechanism: - A node is in Primary state if it belongs to a partition of the cluster that has quorum (i.e., is active and able to serve writes) - A node is in Non-primary state if it detects its partition lacks quorum — SQL queries will fail against it in this state

This is a different concept from Pacemaker's Master/Slave resource states (used when Galera is managed by Pacemaker — see Pacemaker HA). Don't conflate the two: a node can be Pacemaker "Master" and Galera "Non-primary" simultaneously if something has gone wrong.

Why a Minimum of Three Nodes

Two-node clusters can't distinguish a real network partition from a node failure without external tie-breaking, so they're vulnerable to split-brain. Three nodes lets the cluster use majority quorum (2 of 3) to decide which partition stays Primary.

Three nodes also matters for availability during State Snapshot Transfer (SST): if one node needs a full state transfer, one of the two others becomes its donor and is blocked while the transfer completes. With three nodes, the third can keep serving clients. With two, you'd have no node free to answer queries during SST.

Cluster Address and Bootstrapping

wsrep_cluster_address is how a node finds (or starts) the cluster, in URI form:

wsrep_cluster_address="gcomm://node1,node2,node3"

This value has a special dual meaning: - gcomm:// (empty) — bootstrap a brand-new cluster, seeded from this node's local state. This node becomes the founding Primary Component. - gcomm://node1,node2,node3 — join an existing cluster by connecting to the listed peers.

In practice: the first node up is started with an empty gcomm:// (or the --wsrep-new-cluster / galera_new_cluster helper) to found the cluster, then reconfigured with the full peer list before subsequent restarts, so a future crash doesn't accidentally re-bootstrap a stale copy of the data. Getting this step wrong — leaving a node pointed at an empty gcomm:// — is the single most common cause of split-brain and data loss in hand-managed Galera clusters.

Ports

Port Purpose
3306 MySQL/MariaDB client connections; also used by mysqldump-based SST
4567 Galera replication traffic (TCP and UDP for multicast replication)
4568 Incremental State Transfer (IST)
4444 All other State Snapshot Transfer methods (e.g. rsync) — also the default Galera Load Balancer control port, see Load Balancing; don't confuse the two when auditing firewall rules

Deployment Topologies

The notes this section is built from distinguish a few usage patterns, worth keeping in mind when designing around Galera rather than just standing it up:

  • Primary-replica style access — even though every Galera node is writable, many deployments still route all writes through one node (via a load balancer) and reads elsewhere, purely to avoid write-write conflict certification failures under contention. This is an application/topology choice, not a Galera requirement — unlike standard MariaDB primary-replica replication, there's no technical primary here.
  • Write-scalable cluster — spreading writes across nodes for throughput; only the write-set (the change), not the full original transaction, is replicated, which is why this is cheaper than it sounds.
  • Disaster recovery cluster — geographically distributing nodes so a site outage doesn't take out the cluster. If pairing with Galera Load Balancer, give the DR node a lower weight so it's not preferred under normal conditions.

Where to Go Next