Think in shards, not indices¶
Shards are what actually sit on disk, get allocated, get merged and trip watermarks. An index is just a name for a group of them. Most disk problems are allocation and lifecycle problems that show up as disk pressure, so adding disk usually treats the symptom.
Shard sizing¶
Elastic's guidance: 10–50 GB per shard and fewer than ~200M documents per shard. Time-series data usually does well at 30–50 GB.
- Too many small shards: every shard carries fixed overhead (segment metadata, cluster state, heap). By default a node accepts at most 1,000 non-frozen shards (
cluster.max_shards_per_node). - Too few huge shards: merges get expensive, and recovering or relocating a shard after a node loss or watermark breach moves a lot of data at once.
The old "20 shards per GB of heap" rule no longer appears in current docs. Per-field mapping overhead is what drives heap use now (see field count).
Find outliers:
GET _cat/shards?v&s=store:desc&h=index,shard,prirep,store,docs
Shards well outside 10–50 GB point to a primary shard count or rollover condition that needs fixing.
Rollover: size, not just age¶
max_age: 1d by itself only works if ingest volume never changes. Roll over on primary shard size, and keep age as a backstop for low-volume streams:
"rollover": {
"max_primary_shard_size": "50gb",
"max_age": "30d"
}
max_size measures the whole index, so it depends on shard count. max_primary_shard_size targets the number that matters. Lifecycle phases are covered in ILM.
Disk watermarks¶
| Setting | Default | Effect |
|---|---|---|
cluster.routing.allocation.disk.watermark.low |
85% | No new shards are allocated to the node |
cluster.routing.allocation.disk.watermark.high |
90% | Shards are actively moved off the node |
cluster.routing.allocation.disk.watermark.flood_stage |
95% | Every index with a shard on the node gets index.blocks.read_only_allow_delete, and writes fail |
cluster.routing.allocation.disk.watermark.flood_stage.frozen |
95% | Same, for the frozen tier's shared cache disk |
Headroom on large disks. On 8.x and later, each watermark also has a max_headroom (low 200 GB, high 150 GB, flood 100 GB, frozen 20 GB). It caps how much free space a percentage can require, so a 20 TB disk doesn't reserve 3 TB at 85%. The caps only apply while the percentage is at its default. If you set a percentage explicitly, set max_headroom too, or use absolute values such as "high": "150gb".
Flood stage releases itself. Since 7.4, the read_only_allow_delete block clears automatically once usage drops below the high watermark. If writes still fail after you've freed space, check whether usage is actually below high, then clear the block by hand if you need to:
PUT */_settings?expand_wildcards=all
{ "index.blocks.read_only_allow_delete": null }
Single-node clusters. Older 7.x single-node clusters ignored watermarks (enable_for_single_data_node defaulted to false) and could fill the disk to 100%. From 8.0 watermarks always apply and the setting is gone. If you're on an old homelab or dev cluster, check this.
Force merge¶
Merging read-only indices down to one segment cuts per-segment overhead and improves compression, but:
- It temporarily needs up to the index's size again in free space, because new segments are written before old ones are deleted. Running force merges across many indices at once during a cleanup has caused flood-stage events. Do one index at a time.
- Only force merge indices that are no longer written to. In ILM, that means after rollover (warm phase or later).
Mapping and field-count cost¶
Every mapped field costs disk (doc values, index structures) and heap, even when hardly any documents use it. When dynamic mapping is on and different services log slightly different JSON shapes, an index can end up with thousands of sparse fields.
GET [ index ]/_mapping
GET [ index ]/_field_caps?fields=*
- Set
index.mapping.total_fields.limiton purpose rather than letting the 1,000 default act as an accidental circuit breaker. - Use
"dynamic": false(keep the field in_source, but don't map it) on free-form metadata objects. - Use ECS names, so the same data doesn't get mapped under five different names.
Compression, _source and replicas¶
best_compression(DEFLATE instead of LZ4) typically saves 15–30% at some CPU cost.index.codecis a static setting, so apply it with ILM'sforcemergeaction ("index_codec": "best_compression") in the warm phase rather than on a live index.- Disabling
_sourcesaves space but breaks reindex, update and update-by-query, and makes the data impossible to migrate later. It's almost never worth it for logs. (Synthetic_sourceis the modern alternative where your license allows it.) number_of_replicas: 0halves storage for data you can restore from snapshots, but only if you have a tested restore path. Replicas are also your protection against node loss.
Snapshots: the same disk, a different problem¶
A snapshot repository on the same filesystem as the data nodes (common in small and homelab setups) draws from the same disk budget. Check that SLM retention actually prunes:
GET _slm/policy
GET _slm/stats
GET _cat/snapshots/[ repo ]?v&s=end_epoch
Confirm each policy has a retention block (expire_after, min_count, max_count) and that old snapshots really are disappearing.
Emergency playbook: disk is full now¶
In order of how fast each step gives relief:
- Delete or close the oldest, least valuable data:
DELETE [ old-index-pattern ]*. For data streams, delete old backing indices, or the whole stream if it's disposable. - Drop replicas temporarily on large, non-critical indices:
PUT [ index ]/_settings { "index.number_of_replicas": 0 }. - Clean up leftovers: orphaned
.kibana_*indices from old upgrades, old.ml-*results, failed-rollover fragments. - Raise watermarks temporarily to buy time, and put them back afterwards:
json
PUT _cluster/settings
{ "persistent": { "cluster.routing.allocation.disk.watermark.high": "93%",
"cluster.routing.allocation.disk.watermark.flood_stage": "97%" } }
Undo it by setting both back to null. Use persistent, because transient cluster settings are deprecated.
5. Confirm the flood-stage block has cleared (see above).
6. Then fix the root cause (rollover, retention, field explosion) before adding disk.
Catch it early¶
Alert on the rate of change as well as the level: a node at 70% climbing 5% a day is more urgent than one sitting flat at 80%.
GET _cat/allocation?v # per-node disk % and shard count
GET _cat/indices?v&s=pri.store.size:desc # biggest indices
GET _nodes/stats/fs # raw filesystem stats
GET */_ilm/explain?only_errors=true # ILM stuck on errors
An ILM policy can be attached and still be stuck, for example waiting on a shrink that can't allocate. The index then sits in the hot phase indefinitely while it keeps growing. See finding stuck indices.