Andrew Mercer
on this page

ILM moves indices through phases as they age. Typically: roll over when big enough, compact once writes stop, move to cheaper hardware, then delete.

Phases

Phase Typical actions Notes
Hot rollover Actively written. Avoid force merging here: it competes with indexing I/O
Warm forcemerge (1 segment, index_codec: best_compression), shrink, allocate No longer written. Shrink indices that were over-sharded for their final size
Cold readonly, allocate to cold-tier nodes, optionally searchable_snapshot Rarely queried. Denser, cheaper disk
Frozen searchable_snapshot (partially mounted) Data lives in object storage (S3, Azure Blob, GCS) and only a local cache uses disk. Often the biggest disk saving for long-retention logs. Needs an Enterprise licence
Delete delete (optionally wait_for_snapshot) End of retention

min_age for each phase is measured from rollover, not from index creation.

Example policy

PUT _ilm/policy/logs-90d
{
  "policy": {
    "phases": {
      "hot": {
        "actions": {
          "rollover": { "max_primary_shard_size": "50gb", "max_age": "7d" }
        }
      },
      "warm": {
        "min_age": "7d",
        "actions": {
          "forcemerge": { "max_num_segments": 1, "index_codec": "best_compression" },
          "set_priority": { "priority": 50 }
        }
      },
      "cold": {
        "min_age": "30d",
        "actions": { "readonly": {}, "set_priority": { "priority": 0 } }
      },
      "delete": {
        "min_age": "90d",
        "actions": { "delete": {} }
      }
    }
  }
}

Attach the policy through the index template (index.lifecycle.name). See index templates.

Find unmanaged and stuck indices

Unmanaged indices were created before a policy existed, or by a process that didn't set index.lifecycle.name. They stay in hot forever: no rollover, no merge, no deletion.

GET */_ilm/explain?filter_path=indices.*.managed,indices.*.index
GET _cat/indices?v&h=index,pri.store.size&s=pri.store.size:desc

Any sizeable index with "managed": false is quietly leaking disk.

Stuck indices are managed, but not progressing:

GET */_ilm/explain?only_errors=true
GET */_ilm/explain?only_managed=true&filter_path=indices.*.phase,indices.*.action,indices.*.step,indices.*.age

Compare each index's age with the next phase's min_age. An index that has sat well past it in the same phase is usually stuck: often a shrink that can't allocate, or a rollover with no write alias. This is often the earliest warning of disk trouble.

Retry after fixing the cause:

POST [ index ]/_ilm/retry

Sizing a policy

See steady state for estimating how much disk a policy will settle at.

in this section
* steady-state disk usage