ILM moves indices through phases as they age. Typically: roll over when big enough, compact once writes stop, move to cheaper hardware, then delete.
Phases¶
| Phase | Typical actions | Notes |
|---|---|---|
| Hot | rollover |
Actively written. Avoid force merging here: it competes with indexing I/O |
| Warm | forcemerge (1 segment, index_codec: best_compression), shrink, allocate |
No longer written. Shrink indices that were over-sharded for their final size |
| Cold | readonly, allocate to cold-tier nodes, optionally searchable_snapshot |
Rarely queried. Denser, cheaper disk |
| Frozen | searchable_snapshot (partially mounted) |
Data lives in object storage (S3, Azure Blob, GCS) and only a local cache uses disk. Often the biggest disk saving for long-retention logs. Needs an Enterprise licence |
| Delete | delete (optionally wait_for_snapshot) |
End of retention |
min_age for each phase is measured from rollover, not from index creation.
Example policy¶
PUT _ilm/policy/logs-90d
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": { "max_primary_shard_size": "50gb", "max_age": "7d" }
}
},
"warm": {
"min_age": "7d",
"actions": {
"forcemerge": { "max_num_segments": 1, "index_codec": "best_compression" },
"set_priority": { "priority": 50 }
}
},
"cold": {
"min_age": "30d",
"actions": { "readonly": {}, "set_priority": { "priority": 0 } }
},
"delete": {
"min_age": "90d",
"actions": { "delete": {} }
}
}
}
}
Attach the policy through the index template (index.lifecycle.name). See index templates.
Find unmanaged and stuck indices¶
Unmanaged indices were created before a policy existed, or by a process that didn't set index.lifecycle.name. They stay in hot forever: no rollover, no merge, no deletion.
GET */_ilm/explain?filter_path=indices.*.managed,indices.*.index
GET _cat/indices?v&h=index,pri.store.size&s=pri.store.size:desc
Any sizeable index with "managed": false is quietly leaking disk.
Stuck indices are managed, but not progressing:
GET */_ilm/explain?only_errors=true
GET */_ilm/explain?only_managed=true&filter_path=indices.*.phase,indices.*.action,indices.*.step,indices.*.age
Compare each index's age with the next phase's min_age. An index that has sat well past it in the same phase is usually stuck: often a shrink that can't allocate, or a rollover with no write alias. This is often the earliest warning of disk trouble.
Retry after fixing the cause:
POST [ index ]/_ilm/retry
Sizing a policy¶
See steady state for estimating how much disk a policy will settle at.