Pacemaker + Corosync Cluster Guide (Personal Reference)¶
Part 1 of an 8-part series consolidated from the ClusterLabs notes (Cluster
Information.md, Cluster Management.md, Cluster Setup.md, Cluster
Troubleshooting.md, Create Cluster Resources.md, Highly Available OSCAR
Cluster.md, Really old OSCAR HA wiki document.md, Linux Cluster Dump 1.md, Red
Hat Cluster Dump.md, Sample configuration for an OpenStack controller.md).
Duplicated command blocks across the OSCAR/Ubuntu/CentOS write-ups have been merged
into one canonical version each; the old Heartbeat/openais/crm shell material lives
in its own legacy page (part 7) rather than being deleted, since some of it (DRBD
split-brain recovery in particular) is still genuinely useful if an old box is ever
touched again.
The current stack for anything new is pcs + pacemaker + corosync on
CentOS/RHEL/Fedora — that's what this page and parts 2–6 cover. The legacy
crm/openais stack (Ubuntu 10.04–12.04 era) is obsolete for new builds and is
covered separately in part 7.
Series: 1. Concepts & Setup (this page) · 2. Cluster Management · 3. Fencing & STONITH · 4. Creating Resources · 5. Troubleshooting · 6. OpenStack Example · 7. Legacy Appendix · 8. Building from Source & References
1. Quick concept primer¶
- Corosync — membership, messaging, quorum (the plumbing).
- Pacemaker — resource manager sitting on top of Corosync (decides what runs where).
- pcs — the modern CLI/daemon (
pcsd) for managing both. Replaces the oldcrmshell. - STONITH/fencing — forcibly powers off or isolates a node the cluster can't otherwise be sure is dead. Without it, "no quorum" protection alone doesn't prevent split-brain.
- Quorum — the minimum number of votes needed before Pacemaker is allowed to run resources. Below quorum, resource management is disabled by default.
2. Prerequisites & Installation¶
Run on all nodes unless noted.
yum install -y pcs pacemaker corosync fence-agents-all resource-agents
Hostname resolution¶
Nodes must resolve each other's hostnames — via /etc/hosts or real DNS. Get this
right before anything else; half the "nodes only see themselves" issues below trace
back to this.
Firewall¶
firewall-cmd --permanent --add-service=high-availability
firewall-cmd --reload
Manual port reference if not using the high-availability firewalld service:
| Port | Needed for |
|---|---|
| TCP 2224 | pcsd web UI + node-to-node pcs communication. Must be reachable from every node to every node, including itself. Also needed on Booth arbitrators / quorum-device hosts. |
| TCP 3121 | Pacemaker Remote nodes (crmd → pacemaker_remoted). |
| TCP 5403 | Quorum device host (corosync-qnetd). Configurable via -p. |
| UDP 5404/5405 | Corosync Totem traffic (multicast/unicast). |
| TCP 21064 | DLM-dependent resources (clvm, GFS2). |
| TCP/UDP 9929 | Booth ticket manager (multi-site clusters). |
SELinux¶
Not usually required, but if you hit weird permission denials:
setsebool -P daemons_enable_cluster_mode 1
hacluster user¶
The hacluster system user is created by the pcs/pacemaker packages. Set the
same password on every node:
passwd hacluster
Known gotcha: if pcs cluster auth fails with "Username and/or password is
incorrect" even though the password is definitely right, check that hacluster is
actually a member of the haclient group on every node:
id hacluster # groups=... should include haclient
usermod -a -G haclient hacluster # on every node if it's missing, then retry auth
This is the single most time-wasting failure mode in the whole stack — check it early.
Enable pcsd¶
systemctl start pcsd.service
systemctl enable pcsd.service
3. Building the cluster¶
Run these from one node only.
# Older pcs (<0.10): "pcs cluster auth"
pcs cluster auth node1 node2 node3
# Username: hacluster / Password: (the one set above)
# Newer pcs (0.10+ / RHEL8+): "pcs host auth"
pcs host auth node1 node2 node3 -u hacluster -p <password>
pcs cluster setup --name mycluster node1 node2 node3 --start --enable
This generates /etc/corosync/corosync.conf on every node for you (expected_votes
is derived automatically from the node count) and starts + enables both corosync
and pacemaker.
If you're rebuilding a cluster and get "node is already in a cluster", add
--force:
pcs cluster setup --name mycluster node1 node2 node3 --start --force
Verify¶
pcs status
pcs cluster status
pcs status corosync
pcs quorum status
Healthy output shows all nodes Online and partition with quorum.
Enable at boot¶
Decide deliberately: leaving the cluster disabled at boot means a rebooted node
won't rejoin (and start resources) until you pcs cluster start it — useful if you
want to investigate a problem before it rejoins automatically.
pcs cluster enable --all # enable auto-start at boot on every node
# — or, if you leave it disabled —
pcs cluster start --all # manual start after a reboot