A single _search?scroll=1m&size=1000 request returns only the first 1,000 hits. Keep calling _search/scroll until no hits come back. The script below does that and writes one _source document per line (NDJSON).
Script¶
#!/usr/bin/env bash
# export-index.sh <index> [output-file]
set -euo pipefail
: "${ES_URL:?set ES_URL}"
: "${ES_API_KEY:?set ES_API_KEY}"
index="${1:?usage: $0 <index> [output-file]}"
out="${2:-${index}.ndjson}"
es() {
curl -sS --fail -H "Authorization: ApiKey ${ES_API_KEY}" \
-H 'Content-Type: application/json' "$@"
}
resp=$(es "${ES_URL}/${index}/_search?scroll=2m" \
-d '{"size": 1000, "sort": ["_doc"], "query": {"match_all": {}}}')
: > "$out"
while :; do
scroll_id=$(jq -r '._scroll_id' <<<"$resp")
[[ $(jq '.hits.hits | length' <<<"$resp") -eq 0 ]] && break
jq -c '.hits.hits[]._source' <<<"$resp" >> "$out"
resp=$(es "${ES_URL}/_search/scroll" \
-d "{\"scroll\": \"2m\", \"scroll_id\": \"${scroll_id}\"}")
done
# Release the scroll context
es -X DELETE "${ES_URL}/_search/scroll" -d "{\"scroll_id\": \"${scroll_id}\"}" >/dev/null
echo "Wrote $(wc -l < "$out") documents to ${out}"
"sort": ["_doc"]is the cheapest order for a scroll, so use it when you don't care about order.- Replace
{"match_all": {}}with any query to export a subset, such as a time range. - To keep
_idand_indexas well, change thejqfilter to.hits.hits[] | {_index, _id, _source}.
Several indices¶
for index in [ index1 ] [ index2 ]; do
./export-index.sh "$index"
done
Getting it back in¶
Convert the file to bulk format and post it:
jq -c '{"index": {}}, .' [ index ].ndjson > bulk.ndjson
curl -sS -H "Authorization: ApiKey $ES_API_KEY" -H 'Content-Type: application/x-ndjson' \
-X POST "$ES_URL/[ target_index ]/_bulk" --data-binary @bulk.ndjson | jq '.errors'
Split large files first (split -l 10000). A single bulk request should stay in the tens of MB.
elasticdump¶
elasticdump (npm) does the same job, but last time I tried it, npm audit reported critical vulnerabilities in its dependency tree and it didn't work against my cluster. For reference only:
npx elasticdump \
--headers='{"Authorization": "ApiKey [ api_key ]"}' \
--input="$ES_URL/[ index_name ]" \
--output=/tmp/[ index_name ].json \
--type=data
The scroll script needs only curl and jq and has no dependencies to audit.