7. Troubleshooting¶
Nodes only see themselves online¶
Checklist, in the order that actually finds the problem fastest:
1. Were the resources created in a --group? A missing group is sometimes mistaken
for a membership problem.
2. systemctl restart pacemaker corosync on each node — clears a surprising number
of stuck states.
3. Check pcs status output for partition WITHOUT quorum — if every node reports
only itself online, this is a Totem/network problem, not a Pacemaker one.
4. udpu transport packet rejection ([TOTEM ] Packet rejected from x.x.x.x):
historically fixed (temporarily) by switching transport: udpu to transport: tcp
in corosync.conf, but the real fix was almost always the KVM host's kernel/bridge
networking being misconfigured (bridged NIC, multicast on the bridge,
net.bridge.bridge-nf-call-* sysctls) — fix that first rather than chasing the
transport setting.
5. Totem is unable to form a cluster because of an operating system or network
fault → almost always the local firewall (see §2) or, on virtualized labs, the
hypervisor's bridge/kernel networking.
Auth fails with correct password¶
See §2 — hacluster not in the haclient group is the classic cause.
"Unable to communicate with cluster" during auth¶
Usually a typo'd hostname in the auth command rather than a real communication
problem — re-run in --debug and read the actual DNS/resolve error:
pcs cluster auth node1 node2 node3 --debug
Resources stuck in "stopped"¶
Try disabling STONITH temporarily to isolate whether fencing is the blocker, then configure it properly rather than leaving it off:
pcs property set stonith-enabled=false
If that's not it, check pcs status for the actual failure reason under
Failed Resource Actions — the two recurring causes in these notes were:
- IP resource configured with a CIDR baked into the ip= value (see §6).
- The resource genuinely isn't installed/configured yet on that node (e.g. MySQL
never initialized its data directory before Pacemaker tried to start it).
Resources won't migrate¶
Check constraints — a location or colocation constraint pinning the resource is
almost always the reason a manual pcs resource move doesn't stick.
Enable verbose Corosync logging¶
# /etc/corosync/corosync.conf
logging {
to_logfile: yes
logfile: /var/log/cluster/corosync.log # or /var/log/corosync/corosync.log
to_syslog: yes
debug: on
timestamp: on
}
Clear a stuck fail-count¶
pcs resource cleanup <resource>
Taking over an unfamiliar cluster¶
Find the VIP a cluster was configured with:
pcs resource show | grep -i "params ip"
Corosync logs live in /var/log/corosync/ (or /var/log/cluster/corosync.log on
older layouts).
← Part 04: Creating Resources · Part 06: OpenStack Example →