Cluster health¶
GET _cluster/health
GET _cluster/health?level=indices&filter_path=indices.*.status
GET _health_report
| Status | Meaning |
|---|---|
green |
All primary and replica shards are allocated |
yellow |
All primaries are allocated, some replicas aren't. Data is complete, redundancy is reduced. Expected on a single node with replicas > 0 |
red |
At least one primary is unallocated. Some data can't be searched or written |
_health_report reports why the cluster is in that state, broken down by indicator (disk, shards availability, ILM, SLM, and others).
Why is a shard unassigned?¶
GET _cat/shards?v&h=index,shard,prirep,state,unassigned.reason&s=state
GET _cluster/allocation/explain
Called without a body, allocation/explain explains the first unassigned shard it finds. Common answers: disk watermark exceeded, no node with the required tier or attribute, or more replicas than nodes.
Node resources¶
GET _cat/nodes?v&h=name,node.role,heap.percent,ram.percent,cpu,load_1m,disk.used_percent&s=name
GET _nodes/stats/jvm,os,fs,process?filter_path=nodes.*.name,nodes.*.jvm.mem.heap_used_percent,nodes.*.os.cpu.percent,nodes.*.fs.total
- Heap: sustained use above ~75% with frequent old-generation GCs means memory pressure. Look at
jvm.gc.collectors.old. ram.percentis normally high, because the filesystem cache is doing its job. It isn't something to alert on.
Thread pools: rejections are the real signal¶
GET _cat/thread_pool/write,search?v&h=node_name,name,active,queue,rejected,completed
A growing rejected count on write means indexing is outpacing the cluster, and clients see 429 errors. On search, it means queries are queueing. Alert on the rate of change of rejected.
Indexing and search rates¶
The stats APIs return cumulative counters. To get a rate, take two samples and divide the difference by the interval:
q() { curl -s -H "Authorization: ApiKey $ES_API_KEY" \
"$ES_URL/_stats/indexing,search?filter_path=_all.total.indexing.index_total,_all.total.search.query_total"; }
a=$(q); sleep 60; b=$(q)
jq -n --argjson a "$a" --argjson b "$b" '{
index_per_sec: (($b._all.total.indexing.index_total - $a._all.total.indexing.index_total) / 60),
query_per_sec: (($b._all.total.search.query_total - $a._all.total.search.query_total) / 60)
}'
Latency works the same way: divide the change in index_time_in_millis by the change in index_total, and the change in query_time_in_millis by the change in query_total.
Something is slow right now¶
GET _nodes/hot_threads
GET _tasks?detailed=true&actions=*search*
GET _cat/pending_tasks?v
What to alert on¶
| Signal | API |
|---|---|
| Status is not green for more than N minutes | _cluster/health |
| Disk use is climbing fast | _cat/allocation (see disk space) |
| Rejections are increasing | _cat/thread_pool |
| Heap stays above 85% | _nodes/stats/jvm |
| ILM errors | */_ilm/explain?only_errors=true (see ILM) |