elasticsearch-cluster-health
Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Use when an operator reports yellow or red status, unassigned shards, allocation failures, or wants read-only triage before deeper investigation. Teaches replica-vs-primary impact, allocation decide
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 d5a3c381574a0489… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Diagnose Cluster Health
Triage a non-green Elasticsearch cluster read-only: localize the problem, classify the allocation decider, and report the single most likely cause with remediation. Never mutate cluster state — surface findings and let the operator act.
Environment Configuration
This skill executes Elasticsearch operations through the elastic CLI. If the
elastic CLI is not installed, tell the user what it is needed for. Do
not guess credentials, call the HTTP API directly, or attempt other workarounds.
This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping,
GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document
maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API
directly.
Process
-
Read the overall status. Call
GET /_cluster/health. Thestatusfield is the verdict:green— every primary and replica is assigned. Report healthy and stop.yellow— every primary is assigned but at least one replica is not. Data remains readable; redundancy is degraded. This is not data loss.red— at least one primary is unassigned. Data for that shard is unavailable; treat as urgent.
Also read
unassigned_shards,initializing_shards, andrelocating_shards. The decision: continue only when status is yellow or red. Ifinitializing_shards > 0andunassigned_shards == 0, the cluster is recovering on its own — callGET /_cat/recoveryto confirm progress, wait, and re-checkGET /_cluster/healthbefore escalating.Data needed: cluster-wide
statusand shard counters. -
Localize the problem to one index. Call
GET /_cluster/health?level=indicesand pick the index that drives the cluster-wide status:- Any red index outranks every yellow index.
- Among reds or yellows, prefer the index with the most
unassigned_shards. - A red system index (
.security,.kibana*,.fleet-*) outranks application indices because the rest of the stack depends on it.
Optionally call
GET /_cat/shards/{index}?h=index,shard,prirep,state,unassigned.reasonto list every unassigned shard on that index and see whether failures are primaries (prirep=p) or replicas (prirep=r).The decision: focus the next steps on exactly one index — the one whose recovery unblocks the cluster.
Data needed: per-index
statusandunassigned_shards; shard role (primary vs replica) when available. -
Separate trigger from root cause. Call
POST /_cluster/allocation/explainwith no body so Elasticsearch selects an unassigned shard, or target the worst shard explicitly:{ "index": "<index>", "shard": <id>, "primary": <true|false> }Read these fields in order:
primary—falsemeans a replica is unassigned (typical yellow);truemeans a primary is unassigned (typical red).can_allocate— top-level allocation verdict (no,yes,throttled,no_valid_shard_copy, …).unassigned_info.reason— what triggered reassignment (e.g.NODE_LEFT,INDEX_CREATED). This is not the root cause whencan_allocateisno; it only explains why the shard became unassigned.allocate_explanation— human-readable summary; quote it verbatim in the report.node_allocation_decisions[].deciders[]— per-node decider results. Find deciders withdecision: "NO"; the decider name (e.g.disk_threshold,filter,awareness) is the root cause class.
The decision:
- Yellow +
primary: false— impact is limited to replica redundancy; no data loss. Continue to step 4 to name the blocking decider (do not stop atNODE_LEFT). - Red +
primary: true— data for that shard is missing. Continue to step 4; ifcan_allocateisno_valid_shard_copy, treat as potential data loss immediately.
Data needed: allocation-explain response for one representative unassigned shard on the chosen index.
-
Classify the decider. Map the blocking signal to a cause class. Prefer the decider with
decision: "NO"over theunassigned_info.reasontrigger.Signal Cause class Typical remediation (operator applies) decider: disk_threshold,decision: NODisk high/low watermark exceeded Free disk on the named node, add data-node capacity, or adjust cluster.routing.allocation.disk.watermark.*after confirming usage viaGET /_cat/allocationdecider: filterordecider: awareness,decision: NOAllocation filtering or zone awareness Add a node that satisfies index.routing.allocation.*/ awareness attributes, or adjust index/cluster allocation settingsdecider: throttlingor recovery in progressTransient recovery Wait; monitor GET /_cat/recoveryand re-checkGET /_cluster/healthcan_allocate: no_valid_shard_copy(often with emptynode_allocation_decisions)No surviving shard copy See step 5 — data loss scenario can_allocate: yesbut shard still unassignedDelayed allocation or cluster state catch-up Check unassigned_info.atdelay; wait and re-checkFor disk pressure (common yellow scenario after
NODE_LEFT): replicas relocate to remaining nodes; if a survivor is above the high watermark (cluster.routing.allocation.disk.watermark.high, default 90%), thedisk_thresholddecider blocks replica allocation even though primaries stay assigned. The fix is disk capacity or watermark relief — not deleting the index or forcing an empty primary.Data needed: decider name,
explanationtext, and affected node names fromnode_allocation_decisions. -
Recommend remediation — read-only triage ends here. Report the single most likely cause (decider class + verbatim
allocate_explanation) and one primary remediation path. Match urgency to color and shard role.Yellow / replica unassigned (no data loss):
- State clearly: all primaries are assigned; only replicas are missing; no data loss.
- Name the real decider (e.g. disk high watermark on
es-node-2), not merely “a node left”. - Recommend: free disk space, expand storage, add data nodes, or adjust disk watermarks after reviewing
GET /_cat/allocation. - Do not recommend: deleting the index,
allocate_empty_primary, force-allocating over a healthy primary, or restarting the entire cluster without evidence.
Red / primary unassigned with
no_valid_shard_copy(data loss risk):- State clearly: a primary shard is unassigned; queries/routing for that shard fail; treat as urgent and localized to the named index.
- Explain: the only copy was on the departed node; Elasticsearch cannot allocate a primary because no valid copy
exists on any remaining node (
can_allocate: no_valid_shard_copy). - Recovery paths in order:
- Bring the departed node back if its data directory is intact — the shard copy returns.
- Restore from snapshot into the index (or a new index followed by reindex) when snapshots exist.
- Last resort only:
POST /_cluster/reroutewithallocate_empty_primary— this creates an empty primary and permanently loses all documents on that shard. State data loss explicitly; never present this as the first or casual fix.
- Do not recommend: deleting the index without discussing data loss, or
allocate_empty_primarywithout the data-loss warning.
Self-healing in progress:
- When deciders show throttling or active peer recovery, recommend waiting and re-checking read-only APIs above.
Do not execute reroutes, snapshot restores, or settings changes — surface cause and remediation only.
Guidelines
- Read-only: Use only GET/POST explain APIs for triage. Remediation is advice; the operator performs writes.
- Trigger ≠ cause:
unassigned_info.reason: NODE_LEFTexplains the event;node_allocation_decisionsdeciders explain why allocation still fails. - Replica vs primary: Yellow +
primary: false= redundancy gap, not data loss. Red +primary: true= missing data for that shard. - One index, one cause: Pick the highest-impact index and the strongest NO decider; avoid listing every shard.
- Cat helpers: Use
GET /_cat/allocationfor disk percentages per node andGET /_cat/recoveryfor ongoing recoveries when the decider class is unclear or recovery is in progress.
Examples
Yellow — disk watermark after node departure. Health shows yellow with unassigned replicas on logs-2025-07.
Allocation explain returns primary: false, unassigned_info.reason: NODE_LEFT, but disk_threshold decider NO on
es-node-2 (“above the high watermark … 90%”). Report: no data loss; root cause is disk pressure on the receiving node;
remediate disk/watermark — not “node left” alone.
Red — primary with no valid copy. Health shows red on orders-2025 with one unassigned shard. Explain returns
primary: true, can_allocate: no_valid_shard_copy, last_allocation_status: no_valid_shard_copy. Report: urgent;
primary data missing; restore node or snapshot; mention allocate_empty_primary only as last resort with explicit data
loss.
Operations
| HTTP API (shorthand) | elastic CLI command |
|---|---|
GET /_cluster/health | elastic es cluster health |
GET /_cluster/health?level=indices | elastic es cluster health --level indices |
POST /_cluster/allocation/explain | elastic es cluster allocation-explain |
POST /_cluster/allocation/explain (specific shard) | elastic es cluster allocation-explain --index '<index>' --shard <id> --primary true (replica: false) |
GET /_cat/allocation | elastic es cat allocation |
GET /_cat/recovery | elastic es cat recovery |
GET /_cat/shards/{index}?h=index,shard,prirep,state,unassigned.reason | elastic es cat shards --index '<index>' --h index,shard,prirep,state,unassigned.reason |
POST /_cluster/reroute (last-resort empty primary — operator only) | elastic es cluster reroute --commands '<json>' |
Files
1- SKILL.md
c2e16fcc6c13.0 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from elastic/agent-skills8
Onboard an Elastic Cloud organization: configure the `elastic` CLI's Cloud context and API key, establish a default region, then invite users, assign predefined or custom Serverless project roles, and create or revoke Cloud API keys. Use when setting up Cloud authentication or when granting, modifyi
Provision and operate Elastic Cloud infrastructure: create, connect to, update, and delete Serverless projects (Elasticsearch, Observability, Security); manage traffic filters (IP and AWS PrivateLink network security); and manage the lifecycle of Elastic Cloud Hosted deployments. Use when creating o
Create and manage Elastic ML anomaly detection jobs via the API. Use when setting up jobs on an index or data stream, configuring jobs and datafeeds, or opening, starting, or stopping them.
Explain Elasticsearch ML anomaly detection scores, model behavior, and result interpretation. Use when the user asks why a score is high or low, how the model learns, what the numbers mean, or how to troubleshoot unexpected anomaly scores.
Execute ES|QL (Elasticsearch Query Language) queries, use when the user wants to query Elasticsearch data, analyze logs, aggregate metrics, explore data, or create charts and dashboards from ES|QL results.
Design and review Elasticsearch index mappings for stated access patterns: correct field types, text+keyword multi-fields, doc_values tuning, mapping-explosion avoidance, and explicit shard settings. Use when creating a new index, reviewing a mapping for storage or query performance, fixing wrong fi
Load CSV and JSON files into Elasticsearch indices using the bulk API and explicit mappings when field types matter. Use when batch-importing local files, converting CSV rows or JSON arrays to NDJSON bulk format, or verifying document counts and mappings after ingest — not for Logstash pipelines, Be
Help developers new to Elasticsearch get from zero to a working search experience. Guide them through understanding their intent, mapping their data, and building a search experience with best practices baked in. Use this when the user shows intent to build search-related functionality, asks about E
Related knowledge skillsscan passed
PostHog logs for Python
Configure the tRPC client link chain: httpLink, httpBatchLink, httpBatchStreamLink, splitLink, loggerLink, wsLink, createWSClient, httpSubscriptionLink, unstable_localLink, retryLink. Choose the right terminating link. Route subscriptions via splitLink. Build custom links for SOA routing. Link optio
Delivers changes incrementally in thin, verifiable slices. Use when implementing any feature or change that touches more than one file, or when picking up the next task from a plan. Use when rolling a change out behind a feature flag, when you're about to write a large amount of code at once, or whe
Quick reference for ponytail levels, skills and commands. One-shot display. Use for /ponytail-help, "ponytail help", "how do I use ponytail".
Expert guidance for developing with the tinystruct Java framework. Use when working on the tinystruct codebase or any project built on tinystruct — including generating the bin/dispatcher and bin/dispatcher.cmd launcher scripts when a project lacks them, creating Application classes, @Action-mapped