Andrew Mercer
on this page

Cluster health

GET _cluster/health
GET _cluster/health?level=indices&filter_path=indices.*.status
GET _health_report
Status Meaning
green All primary and replica shards are allocated
yellow All primaries are allocated, some replicas aren't. Data is complete, redundancy is reduced. Expected on a single node with replicas > 0
red At least one primary is unallocated. Some data can't be searched or written

_health_report reports why the cluster is in that state, broken down by indicator (disk, shards availability, ILM, SLM, and others).

Why is a shard unassigned?

GET _cat/shards?v&h=index,shard,prirep,state,unassigned.reason&s=state
GET _cluster/allocation/explain

Called without a body, allocation/explain explains the first unassigned shard it finds. Common answers: disk watermark exceeded, no node with the required tier or attribute, or more replicas than nodes.

Node resources

GET _cat/nodes?v&h=name,node.role,heap.percent,ram.percent,cpu,load_1m,disk.used_percent&s=name
GET _nodes/stats/jvm,os,fs,process?filter_path=nodes.*.name,nodes.*.jvm.mem.heap_used_percent,nodes.*.os.cpu.percent,nodes.*.fs.total
  • Heap: sustained use above ~75% with frequent old-generation GCs means memory pressure. Look at jvm.gc.collectors.old.
  • ram.percent is normally high, because the filesystem cache is doing its job. It isn't something to alert on.

Thread pools: rejections are the real signal

GET _cat/thread_pool/write,search?v&h=node_name,name,active,queue,rejected,completed

A growing rejected count on write means indexing is outpacing the cluster, and clients see 429 errors. On search, it means queries are queueing. Alert on the rate of change of rejected.

Indexing and search rates

The stats APIs return cumulative counters. To get a rate, take two samples and divide the difference by the interval:

q() { curl -s -H "Authorization: ApiKey $ES_API_KEY" \
  "$ES_URL/_stats/indexing,search?filter_path=_all.total.indexing.index_total,_all.total.search.query_total"; }

a=$(q); sleep 60; b=$(q)
jq -n --argjson a "$a" --argjson b "$b" '{
  index_per_sec: (($b._all.total.indexing.index_total - $a._all.total.indexing.index_total) / 60),
  query_per_sec: (($b._all.total.search.query_total  - $a._all.total.search.query_total)  / 60)
}'

Latency works the same way: divide the change in index_time_in_millis by the change in index_total, and the change in query_time_in_millis by the change in query_total.

Something is slow right now

GET _nodes/hot_threads
GET _tasks?detailed=true&actions=*search*
GET _cat/pending_tasks?v

What to alert on

Signal API
Status is not green for more than N minutes _cluster/health
Disk use is climbing fast _cat/allocation (see disk space)
Rejections are increasing _cat/thread_pool
Heap stays above 85% _nodes/stats/jvm
ILM errors */_ilm/explain?only_errors=true (see ILM)