neo4j-gds-skill
Neo4j Graph Data Science (GDS) embedded plugin via Python client or Cypher —
- 0
- Installs
- —
- Rating
- —
- Success rate
- 4
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 392986eb8242c760… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
When to Use
- Running GDS algorithms against embedded GDS plugin through Python client (
graphdatascience) - Running GDS algorithms through
CALL gds.*Cypher procedures - Aura Pro, self-managed Neo4j, local Neo4j, or offline DBMS with GDS plugin installed
- Projecting named in-memory graphs, running centrality/community/similarity/path/embedding algorithms
- Chaining algorithms via
mutatemode; building FastRP → KNN pipelines - Writing node embeddings for Neo4j vector indexes / structural similarity search
- Memory estimation before large graph operations
When NOT to Use
- Aura Graph Analytics Sessions / AGA /
GdsSessions/AuraGraphDataScience→neo4j-aura-graph-analytics-skill - AuraDB Cypher API with
{ memory: ... }or{ sessionId: ... }→neo4j-aura-graph-analytics-skill - Cypher query authoring →
neo4j-cypher-skill - Driver/connection setup →
neo4j-driver-python-skill - GraphRAG retrieval →
neo4j-graphrag-skill - Creating/querying vector indexes over written embeddings →
neo4j-vector-index-skill
| Context | Use |
|---|---|
| Aura Pro with GDS plugin | This skill |
| Self-managed/local/offline Neo4j with GDS plugin | This skill |
| AuraDB serverless analytics session | neo4j-aura-graph-analytics-skill |
| Self-managed Neo4j attached to AGA session | neo4j-aura-graph-analytics-skill |
| Non-Neo4j data source | neo4j-aura-graph-analytics-skill |
Pre-flight
Use only with embedded GDS plugin.
from graphdatascience import GraphDataScience
gds = GraphDataScience("neo4j+s://xxx.databases.neo4j.io", auth=("neo4j", "pw")) # AuraDS; aura_ds auto-derived
gds = GraphDataScience("bolt://localhost:7687", auth=("neo4j", "password"))
print(gds.server_version())
RETURN gds.version() AS gds_version
GDS plugin unavailable: client raises GdsNotFound at construction; Cypher raises Unknown function 'gds.version'. AuraDB serverless analytics → neo4j-aura-graph-analytics-skill. Self-managed/local → install or enable GDS plugin.
pip install "graphdatascience>=2.1" # 2.1 required for GDS 2026.09
pip install "graphdatascience[rust-ext]" # optional: faster serialization
Compatibility: graphdatascience 2.1 — GDS >= 2.13 and < 2.28 / < 2026.10; 2.0 — GDS < 2026.9. Both: Python >= 3.10 and < 3.15, Neo4j Python driver >= 5.26 and < 7.0, pandas 2–3, pyarrow 21–25. GDS server < 2.13 → DeprecationWarning at construction; pin graphdatascience<2 (client 1.22) there.
Client 1.x fallback (GDS server < 2.13)
| 2.0 | 1.x client |
|---|---|
gds.page_rank, gds.louvain, … — no v2 prefix | gds.v2.page_rank, gds.v2.louvain, … |
gds.graph.project.native(...) | gds.v2.graph.project(...) |
gds.graph.project.cypher(query) | gds.graph.cypher.project(query, database=...) |
Graph / Model | GraphV2 / ModelV2 |
gds.graph.drop(...) → list[GraphInfo] | gds.v2.graph.drop(...) → single GraphInfo |
Graph.drop(fail_if_missing=) | Graph.drop(failIfMissing=) |
run_cypher(...) — always retries | run_cypher(..., retryable=...) |
Migration guide: Neo4j GDS Python client 2.0 migration
GDS plugin releases track the server: 2026.09.0 requires Neo4j 2026.09 — check the GDS compatibility table before upgrading either side.
GDS plugin 2026.07.0 removed CALL gds.userLog() — read hints and warnings from driver result summary notifications or the Neo4j debug log; track task progress with CALL gds.listProgress().
2.0 client rules:
- Plain endpoints, no
v2prefix — untyped 1.x endpoints andgds.v2.*are gone - snake_case parameters:
page_rank,fast_rp,mutate_property,write_property - Typed result attributes:
result.write_millis, notresult["writeMillis"](streamstill returns DataFrame) - Server version via
gds.server_version()— nogds.version()client method (CypherRETURN gds.version()still valid) - Procedure aliases:
gds.betweenness≡gds.betweenness_centrality; alsogds.closeness,gds.degree,gds.eigenvector,gds.harmonic,gds.kcore - Pipelines:
gds.pipeline.node_classification/link_prediction/node_regression— the only API in 2.0 gds.run_cypher(query, auto_commit=True)[2.1] forCALL { … } IN TRANSACTIONS— 2.0 default retryable transaction rejects it; on 2.0 use the Neo4j driver directlygds.db_driver()[2.1] → underlyingneo4j.Driverfor custom sessions/transactions; closed bygds.close()only if client created itmode="READ"/"WRITE"strings accepted whereverQueryModeis [2.1]- No async/job-handle API on the plugin surface —
compute(),*_async,gds.jobsare AGA Sessions only
Graph Catalog Operations
Native Projection
CALL gds.graph.project(
'myGraph',
['Person', 'City'],
{ KNOWS: { orientation: 'UNDIRECTED' }, LIVES_IN: {} }
)
YIELD graphName, nodeCount, relationshipCount
G, result = gds.graph.project.native("myGraph", "Person", "KNOWS")
print(result.node_count, result.relationship_count)
G, result = gds.graph.project.native(
"myGraph",
{"Person": {"properties": ["age", "score"]}, "City": {}},
{"KNOWS": {"orientation": "UNDIRECTED"}, "LIVES_IN": {"properties": ["since"]}},
overwrite=True, # drop same-named graph first
)
Native projection: plugin/simple Python-client workflow only. AGA Sessions → neo4j-aura-graph-analytics-skill.
1.x fallback: gds.v2.graph.project(...).
Cypher Projection (use for new Cypher workflows, filters, transforms)
G, result = gds.graph.project.cypher(
"""
MATCH (source:Person)-[r:KNOWS]->(target:Person)
WHERE source.active = true
RETURN gds.graph.project($graph_name, source, target,
{ sourceNodeProperties: source { .score }, relationshipType: 'KNOWS' })
""",
graph_name="activeGraph",
)
gds.graph.project.cypher(query) takes no database= — set GraphDataScience(..., database=...) at construction or call gds.set_database(...) before projecting.
Query must end with exactly one RETURN gds.graph.project(...). If validation fails: use gds.run_cypher(...), then gds.graph.get("graphName").
1.x fallback: gds.graph.cypher.project(query, database=...).
AGA Sessions → neo4j-aura-graph-analytics-skill; never use plugin Cypher projection.
Undirected Projection
Native projection: set orientation: 'UNDIRECTED' per relationship type.
Plugin Cypher projection: set undirectedRelationshipTypes: ['*'] in fifth gds.graph.project(...) config argument.
Leiden is defined for directed and undirected graphs. Project undirected relationships when community structure is naturally symmetric.
Inspect and Drop
G.node_count() # 12_043
G.relationship_count() # 87_211
G.node_properties() # projected + mutated properties by label
G.relationship_properties() # projected + mutated properties by type
G.size_in_bytes()
gds.graph.drop(G) # frees JVM heap; returns list[GraphInfo]
G = gds.graph.get("myGraph") # re-attach to existing projection
gds.graph.list()
Memory Estimation — run before large projections and algorithms
CALL gds.graph.project.estimate(['Person'], 'KNOWS')
YIELD requiredMemory, bytesMin, bytesMax, nodeCount, relationshipCount
est = gds.graph.project.estimate(node_projection=["Person"], relationship_projection=["KNOWS"])
print(est.required_memory)
G, project_result = gds.graph.project.native("myGraph", "Person", "KNOWS")
print(project_result.node_count)
# Algorithm estimation:
est = gds.page_rank.estimate(G, damping_factor=0.85)
print(est.required_memory)
Execution Modes
| Mode | Side effect | Returns | Use when |
|---|---|---|---|
stream | None | Row per node/pair | Inspect results; top-N |
stats | None | Single aggregate row | Summary/convergence check |
mutate | Adds node property or relationship type/property to in-memory graph only | Stats row | Chain algorithms |
write | Persists node property or relationship to Neo4j DB | Stats row | Final step — make queryable |
Pattern: stream to verify → mutate to chain → write to persist.
mutate_property must not exist in the in-memory graph. Relationship algorithms such as KNN also require mutate_relationship_type.
After write, re-project to use written properties in subsequent GDS calls (in-memory graph does not see DB writes).
gds.util.asNode() — Enrich Stream Results
stream mode yields nodeId (internal GDS integer). gds.util.asNode(nodeId) translates it back to the DB node so you can access properties.
// Single property
CALL gds.pageRank.stream('myGraph', {})
YIELD nodeId, score
RETURN gds.util.asNode(nodeId).name AS name, score
ORDER BY score DESC LIMIT 10
// Multiple properties — convert once with WITH
CALL gds.pageRank.stream('myGraph', {})
YIELD nodeId, score
WITH gds.util.asNode(nodeId) AS node, score
RETURN node.name AS name, node.born AS born, score
ORDER BY score DESC LIMIT 10
Not needed for write, mutate, or stats modes — those don't return per-node data.
Core Algorithms
PageRank (centrality)
CALL gds.pageRank.stream('myGraph', { dampingFactor: 0.85, maxIterations: 20 })
YIELD nodeId, score
RETURN gds.util.asNode(nodeId).name AS name, score ORDER BY score DESC LIMIT 10
// score: relative influence — not absolute. Compare within same run only.
// didConverge: true means score stabilized; if false, increase maxIterations.
CALL gds.pageRank.write('myGraph', { writeProperty: 'pagerank', dampingFactor: 0.85 })
YIELD nodePropertiesWritten, ranIterations, didConverge
pr_df = gds.page_rank.stream(G, damping_factor=0.85)
mutate_result = gds.page_rank.mutate(G, mutate_property="pagerank", damping_factor=0.85)
write_result = gds.page_rank.write(G, write_property="pagerank", damping_factor=0.85)
print(write_result.write_millis)
Louvain (community detection)
CALL gds.louvain.stream('myGraph', { relationshipWeightProperty: 'weight' })
YIELD nodeId, communityId
CALL gds.louvain.write('myGraph', { writeProperty: 'community' })
YIELD communityCount, modularity
louvain_df = gds.louvain.stream(G)
write_result = gds.louvain.write(G, write_property="community")
print(write_result.community_count)
Leiden is a refinement of Louvain avoiding poorly connected communities — use when community quality > raw speed.
modularity in stats result: range -0.5 to 1.0. [field] Values > 0.3 often indicate meaningful community structure; > 0.7 is strong.
Leiden is defined for directed and undirected graphs. Project undirected relationships when community structure is naturally symmetric.
WCC — Weakly Connected Components
Run WCC first to understand graph structure; partition disconnected graphs before expensive algorithms.
CALL gds.wcc.stream('myGraph', { minComponentSize: 10 })
YIELD nodeId, componentId
CALL gds.wcc.write('myGraph', { writeProperty: 'componentId' })
YIELD nodePropertiesWritten, componentCount
wcc_df = gds.wcc.stream(G)
write_result = gds.wcc.write(G, write_property="componentId")
print(write_result.node_properties_written)
Betweenness Centrality
gds.betweenness.stream(G) # alias of betweenness_centrality; identifies bottleneck/bridge nodes
gds.betweenness.write(G, write_property="betweenness")
Node Similarity
Jaccard similarity from common neighbors — no node properties required.
gds.node_similarity.stream(G, similarity_cutoff=0.1, top_k=10)
gds.node_similarity.write(G, write_relationship_type="SIMILAR", write_property="score",
similarity_cutoff=0.1, top_k=10)
FastRP (node embeddings)
Fast, scalable, production ML pipelines. Set randomSeed for reproducibility.
CALL gds.fastRP.mutate('myGraph', {
embeddingDimension: 256,
iterationWeights: [0.0, 1.0, 1.0],
featureProperties: ['score'],
propertyRatio: 0.5,
normalizationStrength: -0.5,
randomSeed: 42,
mutateProperty: 'embedding'
})
YIELD nodePropertiesWritten
gds.fast_rp.mutate(G, embedding_dimension=256, iteration_weights=[0.0, 1.0, 1.0],
random_seed=42, mutate_property="embedding")
write_result = gds.fast_rp.write(G, embedding_dimension=256, write_property="embedding",
random_seed=42)
print(write_result.write_millis)
For ANN search over structural embeddings, after write, create a Neo4j vector index over the written property. Use neo4j-vector-index-skill.
KNN — K-Nearest Neighbors
Finds k most similar nodes per node based on node properties (typically embeddings).
CALL gds.knn.stream('myGraph', {
nodeProperties: ['embedding'], topK: 10,
sampleRate: 0.5, similarityCutoff: 0.7
})
YIELD node1, node2, similarity
CALL gds.knn.write('myGraph', {
nodeProperties: ['embedding'], topK: 10,
writeRelationshipType: 'SIMILAR', writeProperty: 'score'
})
YIELD relationshipsWritten
knn_df = gds.knn.stream(G, node_properties=["embedding"], top_k=10)
gds.knn.write(G, node_properties=["embedding"], top_k=10,
write_relationship_type="SIMILAR", write_property="score")
FastRP → KNN Pipeline (recommendation)
# 1. Project
G, _ = gds.graph.project.native("myGraph", "Product",
{"BOUGHT_TOGETHER": {"orientation": "UNDIRECTED"}})
# 2. Estimate memory
print(gds.fast_rp.estimate(G, embedding_dimension=128).required_memory)
# 3. Embed
gds.fast_rp.mutate(G, embedding_dimension=128, random_seed=42, mutate_property="emb")
# 4. Similarity
gds.knn.write(G, node_properties=["emb"], top_k=10,
write_relationship_type="SIMILAR", write_property="score")
# 5. Cleanup
gds.graph.drop(G)
Algorithm Selection
| Goal | Algorithm |
|---|---|
| Influence via network links | PageRank / ArticleRank |
| Bottleneck / bridge nodes | Betweenness Centrality |
| Direct connections | Degree Centrality |
| Community (general, fast) | Louvain |
| Community (higher quality) | Leiden |
| Is graph connected? | WCC (run first) |
| Similarity from embeddings | KNN |
| Similarity from neighbors | Node Similarity |
| Shortest path (positive weights) | Dijkstra / A* |
| k alternative paths | Yen's |
| Fast scalable embeddings | FastRP |
| Feature-rich nodes | GraphSAGE (client: gds.graph_sage; Cypher: gds.beta.graphSage) |
Full algorithm catalog → references/algorithms.md
Common Errors
| Error | Cause | Fix |
|---|---|---|
Unknown function 'gds.version' | Embedded GDS plugin unavailable | AGA → neo4j-aura-graph-analytics-skill; self-managed/local → install plugin |
GdsNotFound at client construction | GDS plugin not installed on target DB | Install/enable GDS plugin; AuraDS endpoint only with GDS |
AttributeError: ... no attribute 'version' | gds.version() does not exist in client 2.0 | Use gds.server_version() |
DeprecationWarning at client construction | GDS server < 2.13 | Upgrade GDS server, or pin graphdatascience<2 and use the 1.x mapping table |
Insufficient heap memory / OOM | Graph too large for available JVM heap | Run gds.graph.project.estimate; increase dbms.memory.heap.max_size |
Procedure not found: gds.leiden | Older or incompatible GDS | Check CALL gds.list() for available procedures; upgrade GDS or use Louvain |
Node property 'X' not found after mutate | Property not projected or wrong graph name | Verify G.node_properties() includes the property; check mutate_property spelling |
Graph 'myGraph' already exists | Leftover projection from failed run | overwrite=True, CALL gds.graph.drop('myGraph'), or gds.graph.drop(G) |
mutate_property already exists | Re-running algorithm on same projection | Drop and re-project, or use different mutate_property name |
No algorithm results | Source/target node not in projection | Verify node labels/rel types match projection; check G.node_count() |
AttributeError: 'list' object ... after gds.graph.drop(...) | 2.0 returns list[GraphInfo] | Index the result; 1.x client returns a single GraphInfo |
A query with 'CALL { ... } IN TRANSACTIONS' can only be executed in an implicit transaction from run_cypher | 2.0 runs every query in a retryable managed transaction | graphdatascience>=2.1 + run_cypher(query, auto_commit=True) |
Full Workflow
- Create
gdswithGraphDataScience(...). - Verify plugin:
gds.server_version()orRETURN gds.version(). - Estimate memory:
gds.graph.project.estimate(...)and algorithm.estimate(...). - Project named graph with
gds.graph.project.native(...)orgds.graph.project.cypher(query). - Run
gds.*.streamfirst; switch tomutate; usewriteonly when satisfied. - Drop graph with
gds.graph.drop(G). - GDS server < 2.13 → client 1.22 via the mapping table above.
Built-in test datasets: gds.graph.datasets.load_cora(), gds.graph.datasets.load_karate_club(), gds.graph.datasets.load_imdb()
MCP Tool Mapping
| Operation | MCP tool |
|---|---|
RETURN gds.version() | read-cypher |
gds.pageRank.stream(...) | read-cypher |
gds.pageRank.write(...) | write-cypher |
gds.graph.drop(...) | write-cypher |
| List available procedures | read-cypher → CALL gds.list() |
Before any write-cypher: show exact Cypher, expected nodes/relationships affected, and ask for confirmation. For algorithm write mode, estimate or run stats first when available.
References
- references/algorithms.md — full algorithm catalog: all procedures, parameters, tiers, Cypher + Python examples
- references/graph-projection.md — projection deep-dive: filtering, heterogeneous graphs, relationship orientation, property types
- GDS Manual
- Python Client Docs
Checklist
- Embedded GDS plugin confirmed with
gds.server_version()orRETURN gds.version() - Graph/algorithm memory estimated before large work
- Python examples use 2.0 endpoints (no
v2prefix), snake_case params, typed result attributes - Client 1.x used only with GDS server < 2.13, via the mapping table
- Projection uses native or plugin Cypher projection; no
gds.graph.project.remote(...) - Named graph dropped after use (
gds.graph.drop(G); 1.x:gds.v2.graph.drop(G)) - Execution mode chosen:
stream(inspect) →mutate(chain) →write(persist) -
write_property/mutate_propertychecked for collision with existing properties -
overwrite=Truewhen re-projecting an existing graph name -
randomSeedset for reproducible embeddings - WCC run first on graphs that may be disconnected
Files
4- README.md
24cf4319821.7 KB - SKILL.md
200b4a8b7419.7 KB - references/algorithms.md
cd1d5ce7308.4 KB - references/graph-projection.md
d35bdc7a687.0 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from neo4j-contrib/neo4j-skills8
Authoritative reference for the neo4j-agent-memory Python package — a graph-native memory system for AI agents built on Neo4j — and for the hosted service (NAMS) at memory.neo4jlabs.com. Use this skill whenever the user mentions neo4j-agent-memory, agent memory with Neo4j, context graphs, the POLE+O
Manages Neo4j Aura Agents via the v2beta1 REST API — create, list, get, update, delete,
Serverless Aura Graph Analytics (AGA) GDS Sessions — covers GdsSessions,
Provisions and manages Neo4j Aura instances via CLI (aura-cli v1.7+) or REST API.
Use when working with Neo4j command-line tools — neo4j-cli (modern unified
Generates, optimizes, and validates Cypher 25 queries for Neo4j 2025.x and 2026.x.
Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.
Neo4j .NET Driver v6 — IDriver lifecycle, DI registration (singleton), ExecutableQuery
Related knowledge skillsscan passed
Configure the tRPC client link chain: httpLink, httpBatchLink, httpBatchStreamLink, splitLink, loggerLink, wsLink, createWSClient, httpSubscriptionLink, unstable_localLink, retryLink. Choose the right terminating link. Route subscriptions via splitLink. Build custom links for SOA routing. Link optio
Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine th
Quick reference for ponytail levels, skills and commands. One-shot display. Use for /ponytail-help, "ponytail help", "how do I use ponytail".
Stop hook that blocks Claude from finishing until quality checks pass. Detects rationalization patterns (surface text heuristics), stale learning logs (filesystem mtime), and low disk space. Complements self-audit by mechanically enforcing learning capture habits. Use when Claude should be mechanica