elasticsearch-ingest
Load CSV and JSON files into Elasticsearch indices using the bulk API and explicit mappings when field types matter. Use when batch-importing local files, converting CSV rows or JSON arrays to NDJSON bulk format, or verifying document counts and mappings after ingest — not for Logstash pipelines, Be
- 0
- Installs
- —
- Rating
- —
- Success rate
- 4
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 3994229283c6696b… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Elasticsearch File Ingest
Load local data files into Elasticsearch by converting them to bulk NDJSON, creating an index with the right mappings when types matter, bulk-indexing documents, and verifying the outcome.
Environment Configuration
This skill executes Elasticsearch operations through the elastic CLI. If the
elastic CLI is not installed, tell the user what it is needed for. Do
not guess credentials, call the HTTP API directly, or attempt other workarounds.
This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping,
GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document
maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API
directly.
Scope
This skill covers file → index loading through POST /_bulk. It does not use Logstash, Filebeat, Elastic Agent,
Node.js ingest tools, or other sidecar pipelines. For copying documents between existing indices, use index-to-index
reindex instead of re-parsing source files.
Supported source shapes:
| Source shape | Example | Bulk requirement |
|---|---|---|
| CSV with header row | id,name,age,... then data rows | Parse header into field names; emit one action line + one JSON object per data row |
| JSON array file | [{"a":1},{"a":2}] | Split into per-document lines — never bulk-load the raw array as a single document |
| NDJSON / JSON Lines | one JSON object per line | Optionally add action lines if missing; otherwise ready for bulk |
Parquet, Arrow, and other binary columnar formats are out of scope unless the user converts them to CSV or JSON first.
Process
-
Confirm connectivity. Call
GET /. If the call fails, stop and resolve CLI configuration before reading files or mutating cluster state. -
Inspect the source file and classify its shape. Open the file (or sample the first lines) and decide:
- CSV — first line is a comma-separated header; subsequent lines are records. Count data rows (exclude the header) — you will report this count after load.
- JSON array — file starts with
[and contains an array of objects. Count array elements — each element becomes one indexed document, not one. - NDJSON — one JSON value per line; lines alternate action metadata and document source, or each line is a document that still needs a preceding action line.
The decision: pick the conversion path from NDJSON Bulk Format. Never send raw CSV text or a raw JSON array body to
POST /_bulk. -
Choose the target index name. Use the name the user supplied, or propose a lowercase name derived from the file. Index names must be lowercase, cannot contain spaces or
/, and should not start with-,_, or+. -
Decide whether an explicit mapping is required. Call
GET /{index}/_mappingif the index may already exist.Create an explicit mapping before bulk loading when:
- CSV columns include numbers, dates, or booleans that must be queryable as typed fields (not plain text).
- The user asks for usable column types or aggregation-friendly fields.
- A prior load indexed everything as
text/keywordstrings and must be corrected.
When every field can remain string-like and the user did not specify types, dynamic mapping on first bulk ingest may suffice — but prefer explicit mappings for CSV unless the user explicitly accepts all-string typing.
Read Mapping Design for Ingest for type choices. When the index exists with wrong types, ask the user before calling
DELETE /{index}and recreating it. -
Create the index when needed. When step 4 requires explicit types (or the index does not exist), call
PUT /{index}with amappingsblock before bulk loading. Do not rely on dynamic mapping to inferlong,date, orbooleanfrom CSV string cells — dynamic mapping often maps ambiguous strings totextwith a.keywordsub-field. -
Convert the file to bulk NDJSON. Write a temporary NDJSON file where each document occupies two lines:
- Line 1 — action metadata, e.g.
{"index":{"_index":"<index>"}}(add"_id"only when the user requires stable IDs). - Line 2 — document JSON with correctly typed values (numbers as JSON numbers, booleans as
true/false, dates as ISO-8601 strings such as2023-01-15).
For CSV, map the header row to JSON field names and convert cell values to the JSON types that match the mapping from step 5. For JSON arrays, iterate each array element and emit the action line + object line pair. See worked examples in NDJSON Bulk Format.
- Line 1 — action metadata, e.g.
-
Bulk index the documents. Call
POST /_bulkwith the NDJSON file produced in step 6. Inspect the response: iferrorsistrue, read per-itemerrorobjects, fix mapping or document issues, and retry failed items after remediation. Do not assume success from a zero exit code alone. -
Verify the outcome. Always confirm the load — never report counts from file inspection alone.
- Call
GET /{index}/_countand compare to the expected row/element count from step 2. - When typed columns matter, call
GET /{index}/_mappingand confirm fields such asageare numeric (long/integer), dates aredate, and booleans areboolean— nottext.
Report the verified document count and, when relevant, the confirmed field types. If count or mapping checks fail, see Troubleshooting.
- Call
Guidelines
- Bulk only. All file loads go through
POST /_bulkwith NDJSON action lines — not single-documentPUTloops for batch files, not ingest pipelines as a substitute for client-side CSV parsing, and not posting the untouched source file. - JSON arrays must be split. A four-element array bulk-loaded as one document yields count
1; the correct load yields count4. - CSV header is schema. The first CSV row names fields; each remaining row is one document. A file with one header
plus five data rows must report count
5after ingest. - Type coercion happens in the document JSON. CSV cells arrive as strings; when mappings declare
long,date, orboolean, emit JSON numbers, ISO date strings, and boolean literals in the bulk body — do not rely on Elasticsearch to infer types from quoted CSV strings after dynamic mapping chosetext. - Prefer explicit mappings for typed CSV. Creating the index with
PUT /{index}first prevents silent all-text indexing that breaks range queries and aggregations. - Idempotent re-loads. When reloading into an existing index, ask the user before deleting data. Duplicate bulk
indexactions append new documents unless_idis specified.
Examples
CSV with typed columns
Source (users.csv — header + 5 data rows):
id,name,age,signup_date,active
1,Ada Lovelace,36,2023-01-15,true
Create the index with explicit types, convert rows to NDJSON (five action+document pairs for five data rows), bulk load,
then verify count 5 and mapping types. Full walkthrough:
Mapping Design for Ingest and
NDJSON Bulk Format.
JSON array file
Source (events.json):
[
{ "event_id": "e-1", "type": "login", "user_id": 1, "value": 12.5 },
{ "event_id": "e-2", "type": "logout", "user_id": 1, "value": 0.0 }
]
Convert to four bulk line pairs for four array elements (not one pair for the whole array). Verify GET /{index}/_count
returns 4. See NDJSON Bulk Format.
NDJSON already prepared
When the file alternates action lines and document lines, validate the format and pass it directly to POST /_bulk
after confirming the target index and mappings.
When Not to Use
- Continuous or streaming ingestion — use Elastic Agent or Beats to tail logs and metrics.
- Complex enrichment pipelines — design server-side ingest pipelines separately; this skill still converts files to bulk NDJSON client-side before load.
- Index-to-index copy or mapping migration — reindex between indices instead of exporting to files.
- Very large binary columnar files — convert to CSV or JSON offline first, then follow this skill.
References
- NDJSON Bulk Format — CSV and JSON-array conversion, action-line syntax, batch sizing
- Mapping Design for Ingest — explicit mappings for CSV types, eval-style schemas
- Troubleshooting — wrong counts, text-typed numerics, bulk item errors
Operations
| HTTP API (shorthand) | elastic CLI command |
|---|---|
GET / | elastic es info |
PUT /{index} | elastic es indices create --index '<index>' --mappings '<json>' |
DELETE /{index} | elastic es indices delete --index '<index>' |
POST /_bulk | elastic es bulk --index '<index>' --input-file '<ndjson-path>' |
GET /{index}/_count | elastic es count --index '<index>' |
GET /{index}/_mapping | elastic es indices get-mapping --index '<index>' |
Files
4- SKILL.md
fd03c9f5a310.6 KB - references/mapping-design.md
6ed21aa9e04.0 KB - references/ndjson-bulk-format.md
9658e4b0703.7 KB - references/troubleshooting.md
649ca9db952.9 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from elastic/agent-skills8
Onboard an Elastic Cloud organization: configure the `elastic` CLI's Cloud context and API key, establish a default region, then invite users, assign predefined or custom Serverless project roles, and create or revoke Cloud API keys. Use when setting up Cloud authentication or when granting, modifyi
Provision and operate Elastic Cloud infrastructure: create, connect to, update, and delete Serverless projects (Elasticsearch, Observability, Security); manage traffic filters (IP and AWS PrivateLink network security); and manage the lifecycle of Elastic Cloud Hosted deployments. Use when creating o
Create and manage Elastic ML anomaly detection jobs via the API. Use when setting up jobs on an index or data stream, configuring jobs and datafeeds, or opening, starting, or stopping them.
Explain Elasticsearch ML anomaly detection scores, model behavior, and result interpretation. Use when the user asks why a score is high or low, how the model learns, what the numbers mean, or how to troubleshoot unexpected anomaly scores.
Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Use when an operator reports yellow or red status, unassigned shards, allocation failures, or wants read-only triage before deeper investigation. Teaches replica-vs-primary impact, allocation decide
Execute ES|QL (Elasticsearch Query Language) queries, use when the user wants to query Elasticsearch data, analyze logs, aggregate metrics, explore data, or create charts and dashboards from ES|QL results.
Design and review Elasticsearch index mappings for stated access patterns: correct field types, text+keyword multi-fields, doc_values tuning, mapping-explosion avoidance, and explicit shard settings. Use when creating a new index, reviewing a mapping for storage or query performance, fixing wrong fi
Help developers new to Elasticsearch get from zero to a working search experience. Guide them through understanding their intent, mapping their data, and building a search experience with best practices baked in. Use this when the user shows intent to build search-related functionality, asks about E
Related backend skillsscan passed
PostHog error tracking for Node.js
Mount tRPC as a Fastify plugin with fastifyTRPCPlugin from @trpc/server/adapters/fastify. Configure prefix, trpcOptions (router, createContext, onError). Enable WebSocket subscriptions with useWSS and @fastify/websocket. Set routerOptions.maxParamLength for batch requests. Requires Fastify v5+. Fast
Guides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.
Persistent per-project memory for Claude Code (Context Keeper) driven by deterministic Node.js /ck commands: init, save, resume, info, list, forget, and v1-to-v2 migrate, plus a SessionStart hook that injects a compact project brief. Use when context must survive across sessions, saving session stat
Guide for upgrading Stripe API versions, webhook endpoints, server-side SDKs, Stripe.js, and mobile SDKs
Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management. Sets up table bucket, namespace, table, schema, Glue catalog registration, partitioning, IAM access control. Triggers on: create table, data lake table, analytics table, s