datalineage-summary
Summarizes data lineage graphs on Google Cloud to help users debug data quality issues and understand data provenance for BigQuery and Cloud Storage. Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report. Don't use for generic Bi
- 0
- Installs
- —
- Rating
- —
- Success rate
- 2
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 f489d811fd92ef69… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Data Lineage Summary
This skill guides the agent in investigating and summarizing the Data Lineage graph for a specific focal asset (Table-Level Lineage) or specific fields (Column-Level Lineage). It provides an intuitive left-to-right walkthrough of how data enters and leaves the asset, abstracting away complex node and link details into plain English.
Prerequisites
This skill relies on the data lineage (Knowledge Catalog) MCP Server on
Google Cloud for graph traversal. Ensure you can run search_lineage queries in
both upstream and downstream directions. For detailed connection configurations
and tool schemas, refer to MCP Usage.
Workflow Logic
1. Get Lineage
Fetch the lineage graph in both directions from the focal point (both upstream
and downstream) by making two separate calls to the MCP tool: one with
"direction": "UPSTREAM" and another with "direction": "DOWNSTREAM".
-
Location Strategy: You MUST use the
read_urltool to fetch the comprehensive list of locations dynamically from the provided Knowledge Catalog Locations link. To ensure cross-regional lineage is not missed, always verify the current list of Google Cloud regions using this link before populating thelocationsarray. You MUST populate thelocationsarray with all supported physical regions fetched from this link. You may optionally additionally determine the asset's specific active region (usingbq showorgcloud storage ls). -
Search Parameters: Use
maxDepth = 10,maxResults = 5000andmaxProcessPerLink = 10as robust defaults when callingsearch_lineage. For example, a DOWNSTREAM call should be formatted like this (expanding thelocationsarray as needed):{ "parent": "projects/project_id/locations/us", "locations": [ "us", "us-central1", "us-east1", "us-west1", "europe-west1", "asia-northeast1" ], "rootCriteria": { "entities": { "entities": [ { "fullyQualifiedName": "bigquery:project.dataset.table" } ] } }, "direction": "DOWNSTREAM", "limits": { "maxDepth": 10, "maxResults": 5000, "maxProcessPerLink": 10 } }Ensure you make a similar call with
"direction": "UPSTREAM"to fetch the upstream lineage. -
Column-Level Lineage (CLL): The
search_lineagetool can find all Column-Level Lineage (CLL) by configuring thefieldarray. If Table-Level Lineage (TLL) is requested, configure the call to get CLL links along with the TLL links by exploiting the"*"wildcard. For example:"rootCriteria": { "entities": { "entities": [ { "fullyQualifiedName": "bigquery:project.dataset.table", "field": [ "*" ] } ] } }If evaluating a specific column, replace
"*"with the specific column name (e.g.,"efficiency_score").
2. Summarize
Generate the summary using the prompt guidelines below.
- Persona: Act as an expert Data Lineage Analyst generating a concise, easy-to-understand left-to-right walkthrough of the data flow.
- Structure & Flow: Start immediately with the summary text, structured as
follows:
- Overall Flow Type: State the inferred workflow type and data domain (e.g., "This appears to be a Feature Engineering workflow...").
- Systems Overview: List the primary systems involved up front. If the request is for Column-Level Lineage, you MUST explicitly declare that the scope of the analysis is limited to the specified field up front.
- Upstream Lineage: Use the exact bold header
**Upstream Lineage:**. Narrative must detail how data arrives at the focal asset, mentioning key source systems, projects, and processing tasks (e.g., Spark on Dataproc). - Downstream Lineage: Use the exact bold header
**Downstream Lineage:**. Detail where data goes from the focal asset to final consumer systems. - Analysis Metadata: Display the parameters used for the API call to
provide transparency on the boundaries of the summary. The output must
contain:
- Locations Searched:
{list_of_locations_queried} - Parent Location:
{parent_path} - Depth Limit:
{maxDepth} - Process per Link Limit:
{maxProcessPerLink} - Tip for User: A prompt suggesting they can ask to rerun with expanded locations (if not all were used) or depth.
- Locations Searched:
- Granularity Constraints:
- Prioritize flows between Systems, Projects, and Datasets over individual files/tables.
- You MUST explicitly list specific asset names (e.g., source tables, intermediate views, consumer tables) if there are fewer than 5. Do not just summarize counts if there are fewer than 5; name them explicitly. Otherwise, if 5 or more, aggregate them by count (e.g., "5 Cloud Storage buckets").
- Only mention counts for ultimate sources, final consumers, and total assets.
- Do not repeat project names redundantly for every dataset if only one project is involved.
- Tone: Avoid jargon and generic phrases like "There are distinct factual points." Be direct and clear. The final output is Markdown.
3. Return the Summary
Return the final summarized output back to the user.
External Documentation
Files
2- SKILL.md
31276db6386.6 KB - references/mcp-usage.md
fb83d86a981.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.