neo4j-graphrag-skill
Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python
- 0
- Installs
- —
- Rating
- —
- Success rate
- 7
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 be0953408d294b6c… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Neo4j GraphRAG Skill
When to Use
- Building GraphRAG retrieval pipelines with
neo4j-graphragPython package - Choosing between VectorRetriever, HybridRetriever, VectorCypherRetriever, HybridCypherRetriever
- Writing
retrieval_queryCypher fragments for graph-augmented context - Wiring retriever + LLM into a
GraphRAGpipeline - Using LLM-routed multi-retriever with
ToolsRetriever - Debugging low retrieval quality
- Integrating Neo4j with LangChain, LlamaIndex, or Haystack
When NOT to Use
- KG construction from documents →
neo4j-document-import-skill - Plain vector/semantic search without graph traversal →
neo4j-vector-index-skill - Hybrid search that combines vector with fulltext or other ranked sources →
neo4j-vector-index-skill - GDS algorithms (PageRank, Louvain, node embeddings) →
neo4j-gds-skill - Agent long-term memory →
neo4j-agent-memory-skill - Writing raw Cypher queries →
neo4j-cypher-skill
Retriever Selection
Has fulltext index?
YES → Hybrid variants (HybridRetriever / HybridCypherRetriever)
NO → Vector variants (VectorRetriever / VectorCypherRetriever)
Need graph traversal after vector lookup?
YES → Cypher variants (VectorCypherRetriever / HybridCypherRetriever)
NO → plain variants
Natural-language-to-Cypher? → Text2CypherRetriever (no embedder needed)
LLM should route between retrievers? → ToolsRetriever
Vectors stored in external DB? → WeaviateNeo4jRetriever / PineconeNeo4jRetriever / QdrantNeo4jRetriever
| Retriever | Vector | Fulltext | Graph | Best For |
|---|---|---|---|---|
VectorRetriever | ✓ | — | — | Baseline semantic search |
HybridRetriever | ✓ | ✓ | — | Better recall, no graph expansion |
VectorCypherRetriever | ✓ | — | ✓ | GraphRAG without fulltext |
HybridCypherRetriever | ✓ | ✓ | ✓ | Production GraphRAG — default |
Text2CypherRetriever | — | — | ✓ | NL→Cypher, no embedder |
ToolsRetriever | varies | varies | varies | LLM-routed multi-retriever |
WeaviateNeo4jRetriever | ✓ | — | ✓ | Vectors in Weaviate |
PineconeNeo4jRetriever | ✓ | — | ✓ | Vectors in Pinecone |
QdrantNeo4jRetriever | ✓ | — | ✓ | Vectors in Qdrant |
Install
pip install neo4j-graphrag[openai] # OpenAI LLM + embeddings
pip install neo4j-graphrag[anthropic] # Anthropic Claude
pip install neo4j-graphrag[google] # Vertex AI / Gemini
pip install neo4j-graphrag[bedrock] # Amazon Bedrock (boto3)
pip install neo4j-graphrag[cohere] # Cohere
pip install neo4j-graphrag[mistralai] # MistralAI
pip install neo4j-graphrag[ollama] # Ollama (local)
pip install neo4j-graphrag[weaviate] # Weaviate external retriever
pip install neo4j-graphrag[pinecone] # Pinecone external retriever
pip install neo4j-graphrag[qdrant] # Qdrant external retriever
Requires: Python >= 3.10, neo4j >= 5.17.0 (driver 6.x supported).
Step 2 — Choose Retriever
Pick from Retriever Selection table above.
For custom Cypher hybrid search outside the neo4j-graphrag retriever APIs, use neo4j-vector-index-skill.
Vector backend selection [v1.16+, auto]: on Neo4j 2026.01+ all four vector/hybrid retrievers auto-route through the Cypher 25 SEARCH ... WHERE clause when filters are SEARCH-compatible (simple AND comparisons) and all filter props are declared in the index WITH [n.prop] list. $or, $in, $like, or undeclared props → automatic fallback to db.index.vector.queryNodes() procedure path (with warning log). Declare filterable properties via filterable_properties=[...] on create_vector_index().
Step 3 — Create Indexes (run once)
// Vector index (all retrievers need this)
CREATE VECTOR INDEX chunk_embedding IF NOT EXISTS
FOR (c:Chunk) ON (c.embedding)
OPTIONS { indexConfig: {
`vector.dimensions`: 1536,
`vector.similarity_function`: 'cosine'
} };
// Fulltext index (Hybrid retrievers only)
CREATE FULLTEXT INDEX chunk_fulltext IF NOT EXISTS
FOR (c:Chunk) ON EACH [c.text];
// Confirm ONLINE before ingesting:
SHOW INDEXES YIELD name, state
WHERE name IN ['chunk_embedding', 'chunk_fulltext']
RETURN name, state;
// Both must show state = 'ONLINE'
If index not ONLINE: wait, poll every 5s. Do NOT start ingestion until ONLINE.
Step 4 — Core Pattern (HybridCypherRetriever)
from neo4j import GraphDatabase
from neo4j_graphrag.embeddings import OpenAIEmbeddings
from neo4j_graphrag.generation import GraphRAG
from neo4j_graphrag.llm import OpenAILLM
from neo4j_graphrag.retrievers import HybridCypherRetriever
driver = GraphDatabase.driver(NEO4J_URI, auth=(NEO4J_USERNAME, NEO4J_PASSWORD))
embedder = OpenAIEmbeddings(model="text-embedding-3-large") # OPENAI_API_KEY from env
# retrieval_query: Cypher fragment executed after the vector/fulltext lookup.
# Auto-injected variables: node (matched node) score (similarity float)
# MUST include a RETURN clause. score must appear in RETURN.
retrieval_query = """
MATCH (node)<-[:HAS_CHUNK]-(article:Article)
OPTIONAL MATCH (article)-[:MENTIONS]->(org:Organization)
RETURN node.text AS chunk_text,
article.title AS article_title,
collect(DISTINCT org.name) AS mentioned_organizations,
score
"""
retriever = HybridCypherRetriever(
driver=driver,
vector_index_name="chunk_embedding",
fulltext_index_name="chunk_fulltext",
retrieval_query=retrieval_query,
embedder=embedder,
)
llm = OpenAILLM(model_name="gpt-4.1", model_params={"temperature": 0})
rag = GraphRAG(
retriever=retriever,
llm=llm,
)
response = rag.search(
query_text="Who does Alice work for?",
retriever_config={"top_k": 5},
)
print(response.answer)
driver.close()
VectorCypherRetriever
from neo4j_graphrag.retrievers import VectorCypherRetriever
retriever = VectorCypherRetriever(
driver=driver,
index_name="chunk_embedding",
retrieval_query=retrieval_query,
embedder=embedder,
)
response = rag.search(
query_text="What happened at Apple?",
retriever_config={"top_k": 10},
)
Text2CypherRetriever
Translates natural language to Cypher using an LLM. No embedder required.
Security (v1.16.0+): Every LLM-generated Cypher is run through
EXPLAINfirst. Any statement classified as write/destructive raisesText2CypherRetrievalErrorinstead of executing — prevents prompt-injection attacks.
from neo4j_graphrag.retrievers import Text2CypherRetriever
retriever = Text2CypherRetriever(
driver=driver,
llm=OpenAILLM(model_name="gpt-4.1"),
neo4j_schema=None, # None = auto-fetch schema from DB; pass string to trim
examples=[
"Q: Who works at Neo4j? A: MATCH (p:Person)-[:WORKS_AT]->(c:Company {name:'Neo4j'}) RETURN p.name"
],
)
results = retriever.search(query_text="Which people work at Neo4j?")
ToolsRetriever (LLM-routed multi-retriever)
from neo4j_graphrag.retrievers import ToolsRetriever
tools_retriever = ToolsRetriever(
llm=llm,
retrievers=[vector_retriever, text2cypher_retriever],
)
# LLM decides which retriever(s) to invoke per query
# Convert any retriever to a standalone Tool:
tool = vector_retriever.convert_to_tool()
Filters (pre-filter before vector search)
results = retriever.search(
query_text="quarterly earnings",
top_k=5,
filters={
"date": {"$gte": "2024-01-01"},
"source": {"$eq": "10-K"},
},
)
# Operators: $eq $ne $lt $lte $gt $gte $between $in $like $ilike
query_params (parameterized retrieval_query)
retrieval_query = """
MATCH (node)<-[:HAS_CHUNK]-(a:Article)-[:MENTIONS]->(org:Organization {name: $entity_name})
RETURN node.text, a.title, score
"""
# Pass via retriever.search directly:
results = retriever.search(
query_text="What happened at Apple?",
top_k=10,
query_params={"entity_name": "Apple"},
)
# Or via GraphRAG.search:
response = rag.search(
query_text="What happened at Apple?",
retriever_config={"top_k": 10, "query_params": {"entity_name": "Apple"}},
)
Cypher 25 SEARCH Clause (v1.16.0, Neo4j 2026.x+)
# Enable SEARCH clause syntax in vector/hybrid retrievers (requires Neo4j 2026+)
retriever = VectorRetriever(
driver=driver,
index_name="chunk_embedding",
embedder=embedder,
use_search_clause=True,
)
Since v1.19, the vector and vector-Cypher retrievers automatically prefix SEARCH queries
with CYPHER 25 and fall back to the procedure-based vector search when SEARCH is
unsupported or fails.
Component Imports (v1.19 — breaking, preparing 2.0)
All components moved out of the experimental namespace. Old imports still work but emit a
DeprecationWarning and will be removed in 2.0:
# v1.19+ — preferred
from neo4j_graphrag.components.text_splitters.fixed_size_splitter import FixedSizeSplitter
# deprecated (removed in 2.0)
from neo4j_graphrag.experimental.components.text_splitters.fixed_size_splitter import FixedSizeSplitter
# SimpleKGPipeline did NOT move — only valid path:
from neo4j_graphrag.experimental.pipeline.kg_builder import SimpleKGPipeline
neo4j_graphrag.pipeline (v1.21+) = new lazy dataflow DSL (Pipeline, LocalInterpreter, Sink), unrelated to the experimental task-graph Pipeline behind SimpleKGPipeline. neo4j_graphrag.pipeline.kg_builder does not exist → ModuleNotFoundError.
Also since v1.19: Component / RunContext / TaskProgressNotifierProtocol live in
neo4j_graphrag.components.base, and malformed components raise ComponentDefinitionError
(no longer PipelineDefinitionError).
Dataflow Pipeline DSL + Observers (v1.21)
Lazy neo4j_graphrag.pipeline.Pipeline + LocalInterpreter(observers=[...]); TextSplitter.iter_chunks(). Full API → references/pipeline-dsl.md.
ORDER BY on Cypher Retrievers (v1.16.0)
results = retriever.search(
query_text="...",
top_k=10,
order_by="score DESC",
)
If neo4j_schema=None: retriever fetches schema automatically. For large schemas, pass a trimmed string to reduce LLM prompt size.
Destructive-query guard [v1.16+]: Text2CypherRetriever runs EXPLAIN on the generated Cypher before execution and rejects queries that produce writes (CREATE, MERGE, DELETE, SET, REMOVE, etc.). LLM-generated writes are never executed against the graph.
Custom Prompt Template
from neo4j_graphrag.generation.prompts import RagTemplate
template = RagTemplate(
template="""Answer using ONLY the context below.
Context: {context}
Question: {query_text}
Answer:""",
expected_inputs=["context", "query_text"],
)
rag = GraphRAG(retriever=retriever, llm=llm, prompt_template=template)
return_context and response_fallback
response = rag.search(
query_text="...",
retriever_config={"top_k": 5},
return_context=True, # include raw retrieved chunks
response_fallback="No relevant context.", # skip LLM call if retriever returns nothing
)
print(response.answer)
print(response.retriever_result) # RawSearchResult when return_context=True
Message History (multi-turn)
from neo4j_graphrag.message_history import InMemoryMessageHistory
history = InMemoryMessageHistory()
r1 = rag.search(query_text="Who is Alice?", message_history=history)
r2 = rag.search(query_text="Where does she work?", message_history=history)
External Retrievers, LLM + Embedder Providers
- Weaviate / Pinecone / Qdrant constructors → references/retrievers.md
- LLM classes (
OpenAILLM,AnthropicLLM,GeminiLLM,VertexAILLM,BedrockLLM, …),base_url, token usage,close(); embedder classes + dims → references/providers.md
Index Setup
from neo4j_graphrag.indexes import create_vector_index
# Vector index — adjust dimensions to match your embedding model
create_vector_index(
driver,
name="chunk_embedding",
label="Chunk",
embedding_property="embedding",
dimensions=1536,
similarity_fn="cosine", # or "euclidean"
)
# Fulltext index (run as Cypher)
# CREATE FULLTEXT INDEX chunk_fulltext IF NOT EXISTS
# FOR (c:Chunk) ON EACH [c.text]
Schema Inspection
from neo4j_graphrag.schema import get_schema, get_structured_schema
schema_str = get_schema(driver, sample=1000) # human-readable string
schema_dict = get_structured_schema(driver, sample=1000) # dict with labels/rels/props
Common Errors
| Error | Cause | Fix |
|---|---|---|
ModuleNotFoundError: neo4j_genai | Old package name | pip uninstall neo4j-genai && pip install neo4j-graphrag |
retrieval_query returns 0 rows | Missing MATCH or wrong rel direction | EXPLAIN the fragment; check CALL db.schema.visualization() |
KeyError: 'score' in results | retrieval_query RETURN missing score | Add score to every retrieval_query RETURN clause |
score variable not found | score re-declared in retrieval_query | Do not re-declare score — it is auto-injected |
Text2CypherRetrievalError | LLM generated a write statement | Expected security behavior (v1.16.0+); refine prompt or schema |
TypeError: coroutine | Missing await / asyncio.run() | Wrap async calls: asyncio.run(pipeline.run_async(...)) |
| Empty results from HybridRetriever | Fulltext index not ONLINE | SHOW INDEXES YIELD name, state WHERE state <> 'ONLINE' |
VectorCypherRetriever + filters raises "requires: node_label, embedding_node_property, …" (v1.19 name-mismatch bug) | _fetch_index_infos sets self._embedding_node_property, constructor sets _node_embedding_property | Shim after construction: r._node_embedding_property = r._embedding_node_property or "embedding"; append filter props to r._filterable_properties |
| Embedding dimension mismatch | Index dims ≠ model dims | Recreate index with correct dimensions= value |
Verification Checklist
-
neo4j-graphrag(notneo4j-genai) installed;neo4j >= 5.17.0driver - Vector index ONLINE before ingesting embeddings or running retriever
- Fulltext index ONLINE if using Hybrid variants
- Embedding dims in
create_vector_indexmatch the embedder output -
retrieval_queryreturnsnodeandscorein RETURN (not re-declared) -
query_paramspassed viaretriever_configonrag.search()(not on retriever constructor) - API keys in env vars; never hardcoded
-
llm.close()called when done to release resources
References
Load on demand:
- references/retrievers.md — per-retriever constructor params, external vector DB retrievers,
result_formatter, pre-filter operators - references/providers.md — LLM + embedder provider classes, extras, defaults
- references/pipeline-dsl.md — dataflow
PipelineDSL,Source/Sink, stage observers [v1.21] - references/kg-builder.md —
SimpleKGPipelineconstructor, schema, splitters - references/knowledge-graph-construction.md — advanced KG pipeline customization
- neo4j-graphrag package docs
- RAG & GraphRAG user guide
- KG Builder user guide
- GitHub — neo4j-graphrag-python
- Examples folder
Files
7- README.md
4fc120cbae1.5 KB - SKILL.md
59c40acb6f16.6 KB - references/kg-builder.md
5b0f2701e44.4 KB - references/knowledge-graph-construction.md
2228f3b9d36.6 KB - references/pipeline-dsl.md
ecd5b26c7b1.2 KB - references/providers.md
0d6c8a54543.6 KB - references/retrievers.md
1d4b05fc174.8 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from neo4j-contrib/neo4j-skills8
Authoritative reference for the neo4j-agent-memory Python package — a graph-native memory system for AI agents built on Neo4j — and for the hosted service (NAMS) at memory.neo4jlabs.com. Use this skill whenever the user mentions neo4j-agent-memory, agent memory with Neo4j, context graphs, the POLE+O
Manages Neo4j Aura Agents via the v2beta1 REST API — create, list, get, update, delete,
Serverless Aura Graph Analytics (AGA) GDS Sessions — covers GdsSessions,
Provisions and manages Neo4j Aura instances via CLI (aura-cli v1.7+) or REST API.
Use when working with Neo4j command-line tools — neo4j-cli (modern unified
Generates, optimizes, and validates Cypher 25 queries for Neo4j 2025.x and 2026.x.
Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.
Neo4j .NET Driver v6 — IDriver lifecycle, DI registration (singleton), ExecutableQuery
Related knowledge skillsscan passed
Show ponytail's measured savings (code, cost, speed) from the benchmark. One-shot display. Use for /ponytail-gain, "what does ponytail save", "ponytail impact".
Convene a four-voice council for ambiguous decisions, tradeoffs, and go/no-go calls. Use when multiple valid paths exist and you need structured disagreement before choosing.