skills/ elastic/agent-skills

elasticsearch-reindex

Guide Elasticsearch reindex for performance: local and remote, slicing, throttling, task API. Use when copying or migrating indices, changing mappings, or transforming during reindex.

0
Installs
—
Rating
—
Success rate
5
Files scanned
Scan passedbackend
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

5 files scannedscanner v1.2.0Oct 10, 2026

Content sha256 7ac33666ebdc7a18… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Elasticsearch Reindex

Copy documents from source indices or data streams to a destination using POST /_reindex. An expert reindex workflow prepares the destination explicitly, chooses local versus remote execution, filters at the source when only a subset is needed, runs long copies asynchronously, tracks the task to completion, and verifies the destination document count before reporting results.

Environment Configuration

This skill executes Elasticsearch operations through the elastic CLI. If the elastic CLI is not installed, tell the user what it is needed for. Do not guess credentials, call the HTTP API directly, or attempt other workarounds.

This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping, GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API directly.

Process

  1. Confirm connectivity and deployment type. Call GET /. Read build_flavor and version.number to know whether shard, replica, and cluster-settings APIs are available (Serverless manages shards/replicas internally and blocks most _cluster/* APIs). The decision: continue only when the cluster is reachable. If the call fails, stop — do not guess endpoints or credentials.

  2. Decide local versus remote reindex. Compare where the source and destination live.

    • Same cluster — use local reindex: source.index and dest.index only. Do not add source.remote when both indices are on the cluster you are connected to.
    • Different cluster — use reindex from remote: add source.remote with the remote cluster URL and credentials. Remote reindex does not support slicing; compensate with query-based partitioning (date ranges, term filters) across parallel requests. Confirm the remote host is allowlisted on Self-Managed / ECH (reindex.remote.whitelist in cluster config); Serverless manages allowlisting internally (ECH remotes only, Tech Preview).

    Data needed: source index name(s), destination index name, and whether they share a cluster.

  3. Inspect the source — never guess field names or counts. Call GET /{source}/_mapping to ground field names and types. Call GET /{source}/_count (or GET /_cat/count/{source}?h=count on Self-Managed / ECH) to learn how many documents exist.

    The decision: full copy versus filtered subset.

    • Full copy — omit source.query (match-all behavior).
    • Filtered subset — add source.query with Query DSL. For time ranges, use a range filter on the timestamp field (commonly @timestamp), e.g. "gte": "2025-01-01", "lt": "2025-02-01" for January 2025. Do not run a full-index copy when the user asked for a date range or other filter.

    Data needed: the user's filter criteria and the mapping-confirmed field names.

  4. Prepare the destination index before copying. _reindex does not copy mappings, shard counts, or analyzers. Create the destination with explicit settings and mappings derived from the source mapping via PUT /{dest}.

    • On Self-Managed / ECH: set number_of_replicas: 0 and refresh_interval: "-1" on the destination during the copy for write throughput; restore production values afterward with PUT /{dest}/_settings.
    • On Serverless: omit number_of_shards and number_of_replicas (managed by Elastic); you may set refresh_interval: "-1" during the copy.
    • For data stream destinations: ensure an index template with data_stream: {} exists, create the data stream, and set dest.op_type to "create" (append-only).

    The decision: create/prepare the target rather than relying on auto-creation with dynamic mapping. Wrong or missing mappings cause partial failures or silent type coercion.

    Data needed: destination name, corrected or compatible mappings, and deployment-specific settings constraints.

  5. Build and submit the reindex request. Call POST /_reindex?wait_for_completion=false for any copy that may take more than a few seconds or when the user says the index is large — the response returns a task id immediately instead of blocking.

    Request body essentials:

    • source.index — source index or data stream (correct name, not reversed with dest.index).
    • dest.index — prepared destination from step 4.
    • source.query — include only when step 3 chose a filtered subset.
    • conflicts: "proceed" — when retrying a partially complete reindex.
    • Optional tuning: source.size (batch size), requests_per_second (throttle), slices=auto on local reindex only (parallelize per primary shard — never for remote), scroll (increase keep-alive on slow clusters), max_docs (test runs), script (transform), dest.pipeline (ingest enrichment).

    Example filtered subset (January 2025 only):

    {
      "source": {
        "index": "eval-reindex-src",
        "query": {
          "range": {
            "@timestamp": { "gte": "2025-01-01", "lt": "2025-02-01" }
          }
        }
      },
      "dest": { "index": "eval-reindex-jan" }
    }
    

    Do not reach for _split, _shrink, or snapshot/restore when the task is a filtered subset copy or a straight document migration — those APIs solve different problems.

  6. Track the task to completion. Store the task id from the reindex response. Poll GET /_tasks/{task_id} until completed is true. Read status.total, status.created, and response.failures. On Self-Managed / ECH you may also list active reindex tasks with GET /_tasks?actions=*reindex&detailed; on Serverless, query by task id only (list/cancel are not available). Adjust throttling mid-flight with POST /_reindex/{task_id}/_rethrottle?requests_per_second=N without canceling.

  7. Verify and report the destination count. Call GET /{dest}/_count (works on all deployment types). On Self-Managed / ECH you may also use GET /_cat/count/{dest}?h=count. Compare source filter expectations to the destination count. Report the exact count from the destination — do not estimate or guess.

    After a successful full copy, restore production settings on the destination with PUT /{dest}/_settings (replicas and refresh interval on Self-Managed / ECH; refresh interval only on Serverless).

Deployment constraints

CapabilitySelf-Managed / ECHServerless
Local reindexFull supportFull support
Reindex from remoteFull supportTech Preview — ECH remotes only
number_of_shards/replicasUser-configurableManaged — omit on index creation
slices=auto (local only)SupportedSupported for local reindex
GET /_cat/count/{index}SupportedNot available — use GET /{index}/_count
GET /_tasks (list/cancel)FullGet by task id only
PUT /_cluster/settingsSupportedBlocked
_split / _shrinkSupportedNot available

Consider alternatives first

  • Runtime fields — fix field-type mismatches or add computed fields without reindexing when stored values need not change.
  • Aliases — redirect queries transparently; combine with reindex for zero-downtime mapping changes.
  • Snapshot and restore (Self-Managed / ECH) — faster whole-index transfer when no transformation is needed.

See the decision tree in references/patterns.md.

Reference material

Examples

"Copy logs-2024 into a new index with a corrected mapping" — create the destination first, then reindex:

POST /_reindex
{ "source": { "index": "logs-2024" }, "dest": { "index": "logs-2024-v2" } }

"Reindex a large index in parallel and throttle it" — slice automatically and cap the request rate:

POST /_reindex?slices=auto&requests_per_second=2000
{ "source": { "index": "events" }, "dest": { "index": "events-v2" } }

"Migrate only recent documents" — filter the source with a query:

POST /_reindex
{
  "source": { "index": "metrics", "query": { "range": { "@timestamp": { "gte": "now-30d" } } } },
  "dest": { "index": "metrics-recent" }
}

Guidelines

  • Confirm deployment type first. Call GET / and read build_flavor; shard, replica, cluster-settings, and task APIs differ between Self-Managed / ECH and Serverless (see Deployment constraints).
  • Prefer an alternative when it fits. Runtime fields, aliases, or snapshot-and-restore often avoid a full reindex.
  • Tune the destination for the copy. On Self-Managed / ECH set number_of_replicas: 0 and refresh_interval: "-1" during the copy, then restore production settings afterward; on Serverless these are managed.
  • Parallelize large copies. Use slices=auto for local reindex and throttle with requests_per_second to protect the cluster.
  • Run big jobs asynchronously. Submit with wait_for_completion=false and poll the task instead of blocking.
  • Verify by count. Compare the source filter expectation to the exact destination GET /{dest}/_count — never estimate.

Operations

HTTP API (shorthand)elastic CLI command
GET /elastic es info
GET /{index}/_mappingelastic es indices get-mapping --index '<index>'
GET /{index}/_countelastic es count --index '<index>'
GET /_cat/count/{index}?h=countelastic es cat count --index '<index>' --h count
PUT /{index}elastic es indices create --index '<index>' --mappings '<json>' --settings '<json>'
PUT /{index}/_settingselastic es indices put-settings --index '<index>' --settings '<json>'
POST /_reindex?wait_for_completion=falseelastic es reindex --wait-for-completion false --source '<json>' --dest '<json>'
GET /_tasks/{task_id}elastic es tasks get --task-id '<task_id>'
GET /_tasks?actions=*reindex&detailedelastic es tasks list --actions '*reindex' --detailed
POST /_tasks/{task_id}/_cancelelastic es tasks cancel --task-id '<task_id>'
POST /_reindex/{task_id}/_rethrottle?requests_per_second=Nelastic es reindex-rethrottle --task-id '<task_id>' --requests-per-second <N>

Files

5
47.5 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from elastic/agent-skills8

cloud-onboarding

Onboard an Elastic Cloud organization: configure the `elastic` CLI's Cloud context and API key, establish a default region, then invite users, assign predefined or custom Serverless project roles, and create or revoke Cloud API keys. Use when setting up Cloud authentication or when granting, modifyi

Scan passed 0
cloud-provisioning

Provision and operate Elastic Cloud infrastructure: create, connect to, update, and delete Serverless projects (Elasticsearch, Observability, Security); manage traffic filters (IP and AWS PrivateLink network security); and manage the lifecycle of Elastic Cloud Hosted deployments. Use when creating o

Scan passed 0
elasticsearch-anomaly-detection

Create and manage Elastic ML anomaly detection jobs via the API. Use when setting up jobs on an index or data stream, configuring jobs and datafeeds, or opening, starting, or stopping them.

Scan passed 0
elasticsearch-anomaly-detection-explainer

Explain Elasticsearch ML anomaly detection scores, model behavior, and result interpretation. Use when the user asks why a score is high or low, how the model learns, what the numbers mean, or how to troubleshoot unexpected anomaly scores.

Scan passed 0
elasticsearch-cluster-health

Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Use when an operator reports yellow or red status, unassigned shards, allocation failures, or wants read-only triage before deeper investigation. Teaches replica-vs-primary impact, allocation decide

Scan passed 0
elasticsearch-esql

Execute ES|QL (Elasticsearch Query Language) queries, use when the user wants to query Elasticsearch data, analyze logs, aggregate metrics, explore data, or create charts and dashboards from ES|QL results.

Scan passed 0
elasticsearch-index-design

Design and review Elasticsearch index mappings for stated access patterns: correct field types, text+keyword multi-fields, doc_values tuning, mapping-explosion avoidance, and explicit shard settings. Use when creating a new index, reviewing a mapping for storage or query performance, fixing wrong fi

Scan passed 0
elasticsearch-ingest

Load CSV and JSON files into Elasticsearch indices using the bulk API and explicit mappings when field types matter. Use when batch-importing local files, converting CSV rows or JSON arrays to NDJSON bulk format, or verifying document counts and mappings after ingest — not for Logstash pipelines, Be

Scan passed 0

Related backend skillsscan passed

server-setup

Initialize tRPC with initTRPC.create(), define routers with t.router(), create procedures with .query()/.mutation()/.subscription(), configure context with createContext(), export AppRouter type, merge routers with t.mergeRouters(), lazy-load routers with lazy().

Scan passed 0
api-and-interface-design

Guides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.

Scan passed 0
ck

Persistent per-project memory for Claude Code (Context Keeper) driven by deterministic Node.js /ck commands: init, save, resume, info, list, forget, and v1-to-v2 migrate, plus a SessionStart hook that injects a compact project brief. Use when context must survive across sessions, saving session stat

Scan passed 0
upgrade-stripe

Guide for upgrading Stripe API versions, webhook endpoints, server-side SDKs, Stripe.js, and mobile SDKs

Scan passed 0
deploying-custom-domain-rest-api

Deploys a Regional REST API with a custom domain name, a Lambda backend function, and a request-based Lambda authorizer using AWS CLI. Covers ACM certificate provisioning, API Gateway REST API creation, Lambda function deployment, request authorizer setup, custom domain configuration, base path mapp

Scan passed 0
harness-writing

Designs and improves fuzzing harnesses for C/C++ and Rust. Covers mapping raw bytes onto a target API, generating structured inputs, avoiding non-determinism and false crashes, and deciding what to fuzz together. Use when writing a first LLVMFuzzerTestOneInput or fuzz_target! harness, when a campaig

Scan passed 0