Andrew Mercer
on this page

Pacemaker + Corosync Cluster Guide (Personal Reference)

Part 1 of an 8-part series consolidated from the ClusterLabs notes (Cluster Information.md, Cluster Management.md, Cluster Setup.md, Cluster Troubleshooting.md, Create Cluster Resources.md, Highly Available OSCAR Cluster.md, Really old OSCAR HA wiki document.md, Linux Cluster Dump 1.md, Red Hat Cluster Dump.md, Sample configuration for an OpenStack controller.md). Duplicated command blocks across the OSCAR/Ubuntu/CentOS write-ups have been merged into one canonical version each; the old Heartbeat/openais/crm shell material lives in its own legacy page (part 7) rather than being deleted, since some of it (DRBD split-brain recovery in particular) is still genuinely useful if an old box is ever touched again.

The current stack for anything new is pcs + pacemaker + corosync on CentOS/RHEL/Fedora — that's what this page and parts 2–6 cover. The legacy crm/openais stack (Ubuntu 10.04–12.04 era) is obsolete for new builds and is covered separately in part 7.

Series: 1. Concepts & Setup (this page) · 2. Cluster Management · 3. Fencing & STONITH · 4. Creating Resources · 5. Troubleshooting · 6. OpenStack Example · 7. Legacy Appendix · 8. Building from Source & References


1. Quick concept primer

  • Corosync — membership, messaging, quorum (the plumbing).
  • Pacemaker — resource manager sitting on top of Corosync (decides what runs where).
  • pcs — the modern CLI/daemon (pcsd) for managing both. Replaces the old crm shell.
  • STONITH/fencing — forcibly powers off or isolates a node the cluster can't otherwise be sure is dead. Without it, "no quorum" protection alone doesn't prevent split-brain.
  • Quorum — the minimum number of votes needed before Pacemaker is allowed to run resources. Below quorum, resource management is disabled by default.

2. Prerequisites & Installation

Run on all nodes unless noted.

yum install -y pcs pacemaker corosync fence-agents-all resource-agents

Hostname resolution

Nodes must resolve each other's hostnames — via /etc/hosts or real DNS. Get this right before anything else; half the "nodes only see themselves" issues below trace back to this.

Firewall

firewall-cmd --permanent --add-service=high-availability
firewall-cmd --reload

Manual port reference if not using the high-availability firewalld service:

Port Needed for
TCP 2224 pcsd web UI + node-to-node pcs communication. Must be reachable from every node to every node, including itself. Also needed on Booth arbitrators / quorum-device hosts.
TCP 3121 Pacemaker Remote nodes (crmd → pacemaker_remoted).
TCP 5403 Quorum device host (corosync-qnetd). Configurable via -p.
UDP 5404/5405 Corosync Totem traffic (multicast/unicast).
TCP 21064 DLM-dependent resources (clvm, GFS2).
TCP/UDP 9929 Booth ticket manager (multi-site clusters).

SELinux

Not usually required, but if you hit weird permission denials:

setsebool -P daemons_enable_cluster_mode 1

hacluster user

The hacluster system user is created by the pcs/pacemaker packages. Set the same password on every node:

passwd hacluster

Known gotcha: if pcs cluster auth fails with "Username and/or password is incorrect" even though the password is definitely right, check that hacluster is actually a member of the haclient group on every node:

id hacluster        # groups=... should include haclient
usermod -a -G haclient hacluster   # on every node if it's missing, then retry auth

This is the single most time-wasting failure mode in the whole stack — check it early.

Enable pcsd

systemctl start pcsd.service
systemctl enable pcsd.service

3. Building the cluster

Run these from one node only.

# Older pcs (<0.10): "pcs cluster auth"
pcs cluster auth node1 node2 node3
# Username: hacluster / Password: (the one set above)

# Newer pcs (0.10+ / RHEL8+): "pcs host auth"
pcs host auth node1 node2 node3 -u hacluster -p <password>
pcs cluster setup --name mycluster node1 node2 node3 --start --enable

This generates /etc/corosync/corosync.conf on every node for you (expected_votes is derived automatically from the node count) and starts + enables both corosync and pacemaker.

If you're rebuilding a cluster and get "node is already in a cluster", add --force:

pcs cluster setup --name mycluster node1 node2 node3 --start --force

Verify

pcs status
pcs cluster status
pcs status corosync
pcs quorum status

Healthy output shows all nodes Online and partition with quorum.

Enable at boot

Decide deliberately: leaving the cluster disabled at boot means a rebooted node won't rejoin (and start resources) until you pcs cluster start it — useful if you want to investigate a problem before it rejoins automatically.

pcs cluster enable --all     # enable auto-start at boot on every node
# — or, if you leave it disabled —
pcs cluster start --all      # manual start after a reboot

Part 02: Cluster Management →