Andrew Mercer
on this page

7. Troubleshooting

Nodes only see themselves online

Checklist, in the order that actually finds the problem fastest: 1. Were the resources created in a --group? A missing group is sometimes mistaken for a membership problem. 2. systemctl restart pacemaker corosync on each node — clears a surprising number of stuck states. 3. Check pcs status output for partition WITHOUT quorum — if every node reports only itself online, this is a Totem/network problem, not a Pacemaker one. 4. udpu transport packet rejection ([TOTEM ] Packet rejected from x.x.x.x): historically fixed (temporarily) by switching transport: udpu to transport: tcp in corosync.conf, but the real fix was almost always the KVM host's kernel/bridge networking being misconfigured (bridged NIC, multicast on the bridge, net.bridge.bridge-nf-call-* sysctls) — fix that first rather than chasing the transport setting. 5. Totem is unable to form a cluster because of an operating system or network fault → almost always the local firewall (see §2) or, on virtualized labs, the hypervisor's bridge/kernel networking.

Auth fails with correct password

See §2 — hacluster not in the haclient group is the classic cause.

"Unable to communicate with cluster" during auth

Usually a typo'd hostname in the auth command rather than a real communication problem — re-run in --debug and read the actual DNS/resolve error:

pcs cluster auth node1 node2 node3 --debug

Resources stuck in "stopped"

Try disabling STONITH temporarily to isolate whether fencing is the blocker, then configure it properly rather than leaving it off:

pcs property set stonith-enabled=false

If that's not it, check pcs status for the actual failure reason under Failed Resource Actions — the two recurring causes in these notes were: - IP resource configured with a CIDR baked into the ip= value (see §6). - The resource genuinely isn't installed/configured yet on that node (e.g. MySQL never initialized its data directory before Pacemaker tried to start it).

Resources won't migrate

Check constraints — a location or colocation constraint pinning the resource is almost always the reason a manual pcs resource move doesn't stick.

Enable verbose Corosync logging

# /etc/corosync/corosync.conf
logging {
    to_logfile: yes
    logfile: /var/log/cluster/corosync.log   # or /var/log/corosync/corosync.log
    to_syslog: yes
    debug: on
    timestamp: on
}

Clear a stuck fail-count

pcs resource cleanup <resource>

Taking over an unfamiliar cluster

Find the VIP a cluster was configured with:

pcs resource show | grep -i "params ip"

Corosync logs live in /var/log/corosync/ (or /var/log/cluster/corosync.log on older layouts).


← Part 04: Creating Resources · Part 06: OpenStack Example →