Andrew Mercer
on this page

A single _search?scroll=1m&size=1000 request returns only the first 1,000 hits. Keep calling _search/scroll until no hits come back. The script below does that and writes one _source document per line (NDJSON).

Script

#!/usr/bin/env bash
# export-index.sh <index> [output-file]
set -euo pipefail

: "${ES_URL:?set ES_URL}"
: "${ES_API_KEY:?set ES_API_KEY}"
index="${1:?usage: $0 <index> [output-file]}"
out="${2:-${index}.ndjson}"

es() {
  curl -sS --fail -H "Authorization: ApiKey ${ES_API_KEY}" \
       -H 'Content-Type: application/json' "$@"
}

resp=$(es "${ES_URL}/${index}/_search?scroll=2m" \
          -d '{"size": 1000, "sort": ["_doc"], "query": {"match_all": {}}}')
: > "$out"

while :; do
  scroll_id=$(jq -r '._scroll_id' <<<"$resp")
  [[ $(jq '.hits.hits | length' <<<"$resp") -eq 0 ]] && break
  jq -c '.hits.hits[]._source' <<<"$resp" >> "$out"
  resp=$(es "${ES_URL}/_search/scroll" \
            -d "{\"scroll\": \"2m\", \"scroll_id\": \"${scroll_id}\"}")
done

# Release the scroll context
es -X DELETE "${ES_URL}/_search/scroll" -d "{\"scroll_id\": \"${scroll_id}\"}" >/dev/null

echo "Wrote $(wc -l < "$out") documents to ${out}"
  • "sort": ["_doc"] is the cheapest order for a scroll, so use it when you don't care about order.
  • Replace {"match_all": {}} with any query to export a subset, such as a time range.
  • To keep _id and _index as well, change the jq filter to .hits.hits[] | {_index, _id, _source}.

Several indices

for index in [ index1 ] [ index2 ]; do
  ./export-index.sh "$index"
done

Getting it back in

Convert the file to bulk format and post it:

jq -c '{"index": {}}, .' [ index ].ndjson > bulk.ndjson
curl -sS -H "Authorization: ApiKey $ES_API_KEY" -H 'Content-Type: application/x-ndjson' \
  -X POST "$ES_URL/[ target_index ]/_bulk" --data-binary @bulk.ndjson | jq '.errors'

Split large files first (split -l 10000). A single bulk request should stay in the tens of MB.

elasticdump

elasticdump (npm) does the same job, but last time I tried it, npm audit reported critical vulnerabilities in its dependency tree and it didn't work against my cluster. For reference only:

npx elasticdump \
  --headers='{"Authorization": "ApiKey [ api_key ]"}' \
  --input="$ES_URL/[ index_name ]" \
  --output=/tmp/[ index_name ].json \
  --type=data

The scroll script needs only curl and jq and has no dependencies to audit.