bigquery-slot-cost-optimizer
Analyzes Google Cloud BigQuery slot consumption, query costs, and execution bottlenecks using INFORMATION_SCHEMA. Use when diagnosing slow BigQuery queries, slot starvation, high on-demand query costs, unpartitioned table scans, or join performance issues. Don't use for generic BigQuery administrati
- 0
- Installs
- —
- Rating
- —
- Success rate
- 4
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 3b57866112100191… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
BigQuery slot and cost optimizer
This skill equips AI agents and cloud engineers with procedural heuristics to analyze BigQuery resource consumption, calculate slot hours, identify slot contention and queueing, mitigate Cartesian joins, and optimize unpartitioned table scans.
Trigger conditions and intent mapping
Activate this skill whenever the user asks to:
- "Optimize BigQuery query performance or reduce slot usage"
- "Find the most expensive queries in BigQuery"
- "Diagnose BigQuery slot contention or queueing"
- "Fix slow running BigQuery jobs or memory spillage"
- "Detect Cartesian joins or row count explosions in BigQuery"
- "Identify unpartitioned table scans or missing partition filters"
Prerequisites and environment setup
Before executing this skill, ensure the environment is configured with the necessary SDKs, permissions, and billing:
-
Cloud SDK and client library installation:
-
Install the Google Cloud CLI: Google Cloud SDK installation guide
-
Install the BigQuery Python client:
pip install google-cloud-bigquery
-
-
Project, billing, and regional selection:
-
Set the active project:
gcloud config set project <PROJECT_ID> -
Important: the target Google Cloud project must have an active Cloud Billing account attached.
-
Regional selection: specify the target BigQuery dataset location or execution region, as BigQuery
INFORMATION_SCHEMAviews are strictly region-scoped (for example, multi-regions likeregion-usorregion-eu, or single regions likeregion-us-central1). Querying the wrong region returns empty job telemetry. Pass the matching region via--region(the script automatically normalizes location names likeus-central1toregion-us-central1). For valid location identifiers, see BigQuery locations.
-
-
API enablement:
-
Enable the BigQuery API on the project:
gcloud services enable bigquery.googleapis.com
-
-
Authentication setup:
-
Authenticate the local gcloud environment and configure Application Default Credentials (ADC):
gcloud auth login gcloud auth application-default login
-
-
IAM roles and permissions:
- The executing principal requires the following minimum IAM roles:
roles/bigquery.jobUser: grants permission to run queries and analyze telemetry.roles/bigquery.resourceViewer: grants read-only access to query metadata inINFORMATION_SCHEMA.JOBS_BY_PROJECTand capacity reservations.
- The executing principal requires the following minimum IAM roles:
-
Pricing reference:
- Cost estimates in this skill are for planning purposes. Before running
scripts/slot_analyzer.py, retrieve live BigQuery billing rates at runtime from official Google Cloud BigQuery Pricing (and consult BigQuery editions introduction for edition capabilities) after considering user-specific parameters such as target region, chosen edition (Standard,Enterprise,Enterprise Plus), and commitment tier (Pay-as-you-go,1-year,3-year). Pass these runtime-fetched rates explicitly via--ondemand-rate <USD_PER_TIB>and--slot-hour-rate <USD_PER_SLOT_HOUR>.
- Cost estimates in this skill are for planning purposes. Before running
Diagnostic execution workflow
Execute automated telemetry extraction
Run scripts/slot_analyzer.py to pull and analyze historical query telemetry from INFORMATION_SCHEMA.JOBS_BY_PROJECT, passing the runtime-retrieved pricing rates for your specific region, edition, and commitment tier:
# General analysis passing live regional pricing rates fetched from BigQuery pricing
python3 scripts/slot_analyzer.py --project-id <PROJECT_ID> --days 7 \
--ondemand-rate <USD_PER_TIB> --slot-hour-rate <USD_PER_SLOT_HOUR> --format table
# Output structured JSON for programmatically parsing recommendations
python3 scripts/slot_analyzer.py --project-id <PROJECT_ID> --days 7 \
--ondemand-rate <USD_PER_TIB> --slot-hour-rate <USD_PER_SLOT_HOUR> --format json
# Offline verification mode using synthetic or extracted telemetry
python3 scripts/slot_analyzer.py --mock-data-file path/to/extracted_telemetry.json \
--ondemand-rate <USD_PER_TIB> --slot-hour-rate <USD_PER_SLOT_HOUR> --format table
# Dry-run mode to inspect regional SQL query
python3 scripts/slot_analyzer.py --project-id <PROJECT_ID> --region region-us --dry-run
Run python3 scripts/slot_analyzer.py --help to inspect all supported CLI flags, focus modes (--mode), and required pricing rate arguments (--ondemand-rate per TiB and --slot-hour-rate per slot-hour).
Metric interpretation and decision tree
Evaluate the telemetry output using the following decision rules. CRITICAL MANDATE: After classifying the query issue using the decision tree below, you MUST immediately call view_file on references/remediation_playbooks.md to read and execute the corresponding remediation playbook (Rule SLOT-001, Rule JOIN-001, or Rule PART-001) and include all mandatory diagnostic SQL queries and 4-step checklists in your response.
[Query Telemetry Analyzed]
|
+---> If wait_ratio_avg > 0.40 OR slot_contention == TRUE
| --> Classify as slot contention and queueing (Rule SLOT-001)
| --> MANDATORY: Read Rule SLOT-001 in references/remediation_playbooks.md
|
+---> If shuffle_output_bytes_spilled > 0 OR records_written > 10 * records_read
| --> Classify as Cartesian join (Rule JOIN-001)
| --> MANDATORY: Read Rule JOIN-001 in references/remediation_playbooks.md
|
+---> If total_bytes_billed > 10 GB AND no date/partition filters
| --> Classify as unpartitioned scan (Rule PART-001)
| --> MANDATORY: Read Rule PART-001 in references/remediation_playbooks.md
|
+---> Otherwise
--> Check BI Engine, search indexes, or materialized view opportunities
--> MANDATORY: Read references/optimization_rules.md
Remediation playbooks and architectural reference links
To minimize token consumption in SKILL.md, concrete remediation playbooks (Rule SLOT-001, Rule JOIN-001, Rule PART-001), diagnostic SQL queries, and DDL rewrite patterns are housed in references/:
- Rule SLOT-001 (Slot contention and queueing): Rule SLOT-001 playbook
- Rule JOIN-001 (Cartesian and exploding joins): Rule JOIN-001 playbook
- Rule PART-001 (Unpartitioned scans and partition pruning): Rule PART-001 playbook
- Table partitioning, multi-column clustering, and BI Engine: optimization rules
- BigQuery search indexes and materialized views: optimization rules
Verification and validation protocol
Before finalizing query rewrites:
Dry-run validation
Validate query syntax and calculate estimated bytes scanned without incurring cost:
from google.cloud import bigquery
client = bigquery.Client()
job_config = bigquery.QueryJobConfig(dry_run=True, use_query_cache=False)
query_job = client.query(optimized_sql, job_config=job_config)
print(f"Scanned bytes: {query_job.total_bytes_processed / (1024**3):.2f} GB")
Offline and dry-run validation
-
Offline mock telemetry verification: validate heuristic classification, slot contention detection, Cartesian join identification, and cost estimation offline using synthetic or extracted JSON telemetry payloads (
--mock-data-file):python3 scripts/slot_analyzer.py --mock-data-file path/to/extracted_telemetry.json \ --ondemand-rate <USD_PER_TIB> --slot-hour-rate <USD_PER_SLOT_HOUR> --format table -
CLI dry-run inspection: verify regional SQL query formation and script execution without contacting BigQuery or incurring costs:
python3 scripts/slot_analyzer.py --project-id <PROJECT_ID> --region region-us --dry-run
Files
4- SKILL.md
e7f916cd838.8 KB - references/optimization_rules.md
7749260c6e14.4 KB - references/remediation_playbooks.md
35eb8fed528.7 KB - scripts/slot_analyzer.py
ff0258b4d036.7 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related tooling skillsscan passed
Write growth log entries that extract reusable patterns from completed work — root cause, transferable rule, and a recognizable signal — instead of diary-style event narration, with a 4-8 sentence template and merge-duplicates discipline. Use when capturing what was learned after a complex task, deb
Web performance regression detection. (gstack)
This skill should be used when the user asks to "demonstrate skills", "show skill format", "create a skill template", or discusses skill development patterns. Provides a reference template for creating Claude Code plugin skills.
Helps you build and check a color system for your project. It generates palettes, names semantic tokens, converts between formats and measures contrast.
Creates a new Angular app using the Angular CLI. This skill should be used whenever a user wants to create a new Angular application and contains important guidelines for how to effectively create a modern Angular application.
Audit, diagnose, or optimize website loading and interaction performance, Core Web Vitals, and Lighthouse performance scores.