elasticsearch-anomaly-detection-explainer
Explain Elasticsearch ML anomaly detection scores, model behavior, and result interpretation. Use when the user asks why a score is high or low, how the model learns, what the numbers mean, or how to troubleshoot unexpected anomaly scores.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 2
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 8157cfcafd093eed… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Anomaly Detection Score Explainer
Explain anomaly scores, model behavior, and why results look the way they do. Use the ML REST API for job config and
the standard _search API against .ml-anomalies-* for results — no ES|QL, fully compatible with Elastic
Serverless. For job lifecycle (create, start, stop), use the elasticsearch-anomaly-detection skill.
Environment Configuration
This skill executes Elasticsearch operations through the elastic CLI. If the
elastic CLI is not installed, tell the user what it is needed for. Do
not guess credentials, call the HTTP API directly, or attempt other workarounds.
This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping,
GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document
maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API
directly.
Prerequisite: ML anomaly detection requires a Platinum-equivalent license on self-managed clusters. Serverless projects include ML. The caller needs
monitor_mlto read job config and anomaly results.Serverless note: The
_ml/.../results/*REST endpoints return HTTP 410 in Elastic Serverless. Always usePOST /.ml-anomalies-*/_searchfor result queries instead — fully supported everywhere this skill runs.
Process
-
Decide whether to fetch data or interpret what the user supplied. If the user embeds an anomaly record (or job config) in the prompt, interpret it directly using the domain knowledge below — do not call APIs to re-fetch fields already present. If the job ID, time range, or record is missing, retrieve it from the cluster.
The decision: proceed with judgment-only explanation when the record contains
record_score,initial_record_score,actual,typical, andfunction; otherwise fetch the missing pieces before explaining. -
Verify connectivity when calling the cluster. Call
GET /. If the call fails, stop and surface the connection error — do not guess endpoints or credentials. -
Resolve the job ID and load config. When the job ID is unknown, call
GET /_ml/anomaly_detectorsto list candidates. CallGET /_ml/anomaly_detectors/{job_id}for fullanalysis_config(bucket_span, detectors, custom_rules, use_null, model_plot_config) andGET /_ml/anomaly_detectors/{job_id}/_statsforstate,model_size_stats.memory_status, and data counts.The decision: confirm detector function and direction match the user's question before interpreting scores. A
low_countjob legitimately fires on drops; ahigh_countjob does not. -
Retrieve anomaly records for the time range. Call
POST /.ml-anomalies-*/_searchwithresult_type: record, the job ID, a timestamp range, and optionalrecord_scorefilter. Readinitial_record_score,record_score,actual,typical,function,multi_bucket_impact, andanomaly_score_explanation.Always show both
initial_record_scoreandrecord_score. The gap is the renormalization story. -
Classify the score pattern before speculating on causes.
initial_record_score>>record_score— Renormalization. A later, more extreme anomaly rescale this record downward. This is expected, healthy model behavior — not a broken model or reason to distrust the detection. Useinitial_record_scorefor alerting severity; show both scores and explain the gap explicitly.initial_record_score==record_score— No renormalization has occurred since detection.actual<<typicalwithlow_count,count, orlow_mean— Absence / drop anomaly. A high score is legitimate — the job detected an outage, pipeline stall, or service failure. This is not a false positive. Recommend incident investigation, not score tuning.actual>>typicalwithhigh_countorhigh_mean— Spike anomaly; confirm withsingle_bucket_impact.
Only cite
anomaly_score_explanationfactors present in the record. Ifhigh_variance_penaltyisfalse, do not blame variance. If a factor is absent, note that it was not returned — do not invent it. -
Quantify renormalization across the job (optional). Re-query
POST /.ml-anomalies-*/_searchfor records in the time range sorted bytimestampascending. Computescore_drift = initial_record_score − record_scoreper record and filter to|score_drift| ≥ 20. Large negative drift (initial >> record) confirms renormalization after a more extreme anomaly appeared later. -
Add context when the user asks "what caused this?" or "why so low/high?"
- Model bounds — If
model_plot_config.enabledis true, callPOST /.ml-anomalies-*/_searchwithresult_type: model_plotfor the same job and time range. Compareactualtomodel_lower/model_upper. - Influencers — Call
POST /.ml-anomalies-*/_searchwithresult_type: influencerfor the bucket time range; sort byinfluencer_scoredescending. - Categorization jobs — Call
POST /.ml-anomalies-*/_searchwithresult_type: category_definitionto list learned log patterns (terms,regex,examplespercategory_id).
For aggregations, cross-job queries, bucket-level results, or custom filters beyond score and time, see references/explainer-reference.md.
- Model bounds — If
Common multi-step workflows
| Task | Steps (in order) |
|---|---|
| Explain a specific anomaly | job config → records (job + exact time) → show initial_record_score vs record_score + score factors. |
| Why is my score low? | job config → records → renormalization check → model plot (if enabled) → explain score factors. |
| Why is my score high? | job config → records → check function direction, insufficient history, use_null, cardinality. |
| Renormalization drift | records (timestamp sort) → compute score_drift → list records where initial >> record. |
| Which entities contributed? | influencers (job + time range) → sort by influencer_score. |
| Visualize model bounds | model plot (job + time range) → compare model_lower/model_upper vs actual. |
| Categorization job patterns | category_definition (job_id) → terms, regex, examples per category. |
Critical principles
- Retrieve the record first (or use the one the user supplied). Never explain scores without
initial_record_score,record_score,actual,typical, andfunction. - Renormalization is healthy. When
initial_record_score >> record_score, a more extreme anomaly appeared later and lowered this score — expected behavior, not a model failure. - Direction matters.
low_countfires when values drop;high_countfires on spikes. A high score on a traffic stop withlow_countis correct detection, not a false positive. - Explain factors before speculating. Read
anomaly_score_explanationfrom the record. Only address factors that are present and relevant. - Job config is essential.
bucket_span, detector function,custom_rules,use_null, and memory status all affect scores. Inspect job config when a score is surprising. - Model plot is the most visual explanation. When enabled, show model bounds to illustrate where the actual value falls relative to the expected range.
- For job health ("missing documents", "memory limit", "datafeed not running") use the
elasticsearch-anomaly-detectionskill.
Domain knowledge
Score types
| Term | Meaning |
|---|---|
| record_score | Normalized 0–100 for a single anomaly record; updated by renormalization. >75 critical. |
| initial_record_score | Score assigned at detection time, before renormalization. Use for alerting. |
| anomaly_score | Bucket-level severity aggregated across all detectors in a job. |
| influencer_score | How unusual a specific entity (host, user, service) is in a bucket; high = likely cause. |
| multi_bucket_impact | 0–5; how much sustained, multi-bucket behavior raised the score. ≥3 = behavioral shift. |
anomaly_score_explanation factors
The anomaly_score_explanation field on each record breaks the score into components:
| Factor | Direction | Meaning |
|---|---|---|
| anomaly_length | Raises | Number of consecutive buckets the anomaly spans. Longer → higher score. |
| single_bucket_impact | Raises | Extremity of this single bucket. Lower probability → higher impact. |
| multi_bucket_impact | Raises | Contribution of sustained multi-bucket pattern. |
| anomaly_characteristics_impact | Raises | Whether the anomaly is a mean shift vs. variance change. |
| high_variance_penalty | Lowers | Noisy data or early training → wide confidence bounds → score reduced. |
| incomplete_bucket_penalty | Lowers | Bucket had less data than expected (delayed data, sparse events). |
Why a score might be unexpectedly low
- high_variance_penalty: The metric is historically noisy — wide confidence bounds absorb the spike.
- Renormalization: A more extreme anomaly appeared later and pushed this score down (
initial_record_score>>record_score). - Insufficient training history: Need ≥3 weeks for weekly seasonality, ≥2 full cycles for any detected period.
- bucket_span too large: Short-duration spikes get smoothed. Use a smaller
bucket_spanfor high-frequency events. - Detector function mismatch:
meanvshigh_mean,countvshigh_count— only one direction fires. - incomplete_bucket_penalty: Bucket received less data than expected (ingest latency or gaps).
- custom_rules: A detector filter may be suppressing the anomaly.
Why a score might be unexpectedly high
- Insufficient history: Model hasn't learned the normal pattern yet — early anomalies are unreliable.
- Model split thin: High-cardinality
partition_fieldorby_field→ very few points per entity → unreliable probabilities. - use_null: If
use_null: true, missing entities produce "null" anomalies that may not be meaningful. - Absence / drop detection: With
low_countorlow_mean,actual << typicalproduces a legitimately high score — treat as a real incident, not a false positive.
Model behavior concepts
| Concept | Meaning |
|---|---|
| actual | Observed value. typical is what the model expected. The direction matters. |
| Absence anomaly | actual << typical with count, low_count, or low_mean → outage, pipeline stop, service failure. |
| by_field | Independent baseline per entity (e.g., per host). Each entity compared to its own history. |
| over_field | Population analysis — entity compared to its peer group in the same bucket, not its own history. |
| partition_field | Fully independent sub-models with separate score normalization per partition. |
Model plot and categories
- Model plot: Shows the model's learned upper and lower bounds at each time point. If
actualis within bounds, no anomaly; if outside, the score depends on the distance from bounds. Only available whenmodel_plot_configis enabled on the job. Query viaPOST /.ml-anomalies-*/_searchwithresult_type: model_plot. - Categories: For jobs with a
categorization_field_name, queryresult_type: category_definitionto show log message patterns (terms, regex, examples percategory_id). Anomaly records useby_field_value = <category_id>.
Score troubleshooting protocol
- List jobs — Call
GET /_ml/anomaly_detectorswhen the job ID is unknown. - Get job config and stats — Call
GET /_ml/anomaly_detectors/{job_id}andGET /_ml/anomaly_detectors/{job_id}/_stats. Verifybucket_span, detector function,custom_rules,use_null, job state (opened/closed/failed), andmodel_size_stats.memory_status. - Retrieve the record — Call
POST /.ml-anomalies-*/_searchwithresult_type: record, the job ID, time range, and optional minimumrecord_score. Inspectinitial_record_score,record_score,actual,typical,function,multi_bucket_impact, andanomaly_score_explanation. - Check renormalization — Compare
initial_record_scorevsrecord_score. If initial >> record, re-query records sorted by timestamp and computescore_driftto quantify renormalization across the job. - Visualize model bounds — If
model_plot_configis enabled, queryresult_type: model_plotand show where the actual value fell relative tomodel_lowerandmodel_upper. - Influencers — Query
result_type: influencerfor the anomaly bucket time range; sort byinfluencer_score. - Explain factors — From the record's
anomaly_score_explanation, address each present relevant factor:high_variance_penalty,incomplete_bucket_penalty,anomaly_length,single_bucket_impact,multi_bucket_impact. Do not cite factors absent from the record.
Examples
- "Why is my anomaly score only 15 when the spike looks huge?" → Check renormalization:
initial_record_score(~92) >>record_score(~15). The spike was real; the current score was rescaled down. Use the initial score for alerting. - "Traffic stopped and I got a HIGH score — false positive?" → No.
low_countwithactualfar belowtypicalis legitimate absence detection. Investigate the outage. - "Which entities contributed most to the anomalies in job X last night?" → Query influencers for the time range.
- "Show me the model bounds for this job." → Query model plot when
model_plot_configis enabled. - "List records where the score was renormalized down a lot." → Records sorted by timestamp; filter large
initial_record_score − record_score.
Guidelines
- Report only what the API or the user-supplied record contains; do not invent scores, timestamps, entity values, or explanation factors.
- Always show both
initial_record_scoreandrecord_scorewhen explaining a record; state explicitly whether renormalization occurred. - When a score factor is missing from the record, do not assert it; note that the field was not returned.
- Do not attribute low scores to
high_variance_penaltyorincomplete_bucket_penaltywhen those flags arefalseor absent in the record. - For investigation ("what caused this?", "which service is responsible?") query influencers or construct cross-job searches via references/explainer-reference.md.
- For job health use the
elasticsearch-anomaly-detectionskill.
Operations
| HTTP API (shorthand) | elastic CLI command |
|---|---|
GET / | elastic es info |
GET /_ml/anomaly_detectors | elastic es ml get-jobs |
GET /_ml/anomaly_detectors/{job_id} | elastic es ml get-jobs --job-id '<job_id>' |
GET /_ml/anomaly_detectors/{job_id}/_stats | elastic es ml get-job-stats --job-id '<job_id>' |
POST /.ml-anomalies-*/_search | elastic es search --index '.ml-anomalies-*' --input-file '<search-body.json>' |
Query body shapes for each result_type (record, influencer, model_plot, category_definition) are documented in
references/explainer-reference.md.
Files
2- SKILL.md
b970d67ea117.5 KB - references/explainer-reference.md
e3ea120b106.7 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from elastic/agent-skills8
Onboard an Elastic Cloud organization: configure the `elastic` CLI's Cloud context and API key, establish a default region, then invite users, assign predefined or custom Serverless project roles, and create or revoke Cloud API keys. Use when setting up Cloud authentication or when granting, modifyi
Provision and operate Elastic Cloud infrastructure: create, connect to, update, and delete Serverless projects (Elasticsearch, Observability, Security); manage traffic filters (IP and AWS PrivateLink network security); and manage the lifecycle of Elastic Cloud Hosted deployments. Use when creating o
Create and manage Elastic ML anomaly detection jobs via the API. Use when setting up jobs on an index or data stream, configuring jobs and datafeeds, or opening, starting, or stopping them.
Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Use when an operator reports yellow or red status, unassigned shards, allocation failures, or wants read-only triage before deeper investigation. Teaches replica-vs-primary impact, allocation decide
Execute ES|QL (Elasticsearch Query Language) queries, use when the user wants to query Elasticsearch data, analyze logs, aggregate metrics, explore data, or create charts and dashboards from ES|QL results.
Design and review Elasticsearch index mappings for stated access patterns: correct field types, text+keyword multi-fields, doc_values tuning, mapping-explosion avoidance, and explicit shard settings. Use when creating a new index, reviewing a mapping for storage or query performance, fixing wrong fi
Load CSV and JSON files into Elasticsearch indices using the bulk API and explicit mappings when field types matter. Use when batch-importing local files, converting CSV rows or JSON arrays to NDJSON bulk format, or verifying document counts and mappings after ingest — not for Logstash pipelines, Be
Help developers new to Elasticsearch get from zero to a working search experience. Guide them through understanding their intent, mapping their data, and building a search experience with best practices baked in. Use this when the user shows intent to build search-related functionality, asks about E
Related ai-ml skillsscan passed
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
Neural search via Exa MCP for web, code, and company research. Use when the user needs web search, code examples, company intel, people lookup, or AI-powered deep research with Exa's neural search engine.
MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v
Generates python code that evaluates SageMaker models. Supports two evaluation types: LLM-as-Judge and Custom Scorer. Use when the user says "evaluate my model", "run a benchmark", "test model performance", "how did my model perform", "compare models", or other similar requests.
Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. Use when the user asks to "build a voice agent", "create a LiveKit agent", "add voice AI to my app", "implement handoffs", "structure an agent workflow", "my agent is slow / too chatty", "it says it booked but nothing was saved",
Use Neo4j GenAI Plugin ai.text.* functions and procedures for in-Cypher