skills/ google/skills

bigquery-optimization

Provides workflows to optimize BigQuery environments (capacity planning, editions), storage assets (partitioning, clustering, storage lifecycles, billing models), and SQL queries. Use when optimizing cost, modeling Edition migrations, rightsizing reservations, evaluating logical vs. physical storage

0
Installs
—
Rating
—
Success rate
9
Files scanned
Scan passeddatabase
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

9 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 ecfb52760cf89a4e… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

BigQuery Optimization Workflow

Prerequisites & Environment Setup

Before executing optimization analyses, evaluating editions, or applying DDL modifications:

  1. Google Cloud SDK: Ensure the Google Cloud SDK is installed and configured.

  2. Project Selection: Set the active Google Cloud project:

    gcloud config set project {project_id}
    
  3. API Enablement: Ensure BigQuery and BigQuery Reservation APIs are enabled:

    gcloud services enable \
        bigquery.googleapis.com bigqueryreservation.googleapis.com
    
  4. Authentication: Authenticate the environment:

    • CLI tools and bq commands: gcloud auth login
    • SDKs and automation: gcloud auth application-default login
    • Service accounts: Set GOOGLE_APPLICATION_CREDENTIALS="/path/to/key.json"
  5. Billing & IAM Roles:

    • Verify an active Google Cloud Billing account is attached to {project_id}.
    • Ensure appropriate IAM roles:
      • roles/bigquery.admin or roles/bigquery.resourceAdmin: Reservation and capacity commitment management.
      • roles/bigquery.dataEditor or roles/bigquery.admin: Modifying table schemas, partitioning, clustering, and storage billing models.
      • roles/bigquery.jobUser: Running evaluation queries.
  6. Companion Skills Installation: This skill is part of a 3-pillar operations suite (bigquery-observability, bigquery-optimization, bigquery-troubleshooting). If any companion skill is not yet installed in your environment, install the full suite:

    npx skills add google/skills --skill bigquery-observability --skill bigquery-optimization --skill bigquery-troubleshooting
    

    (If bigquery-observability is not installed, use the self-contained baseline formulas and query templates provided directly in the reference sections below).

Workflows

Determine the optimization focus of the user's request and follow the relevant workflow:

  • Telemetry & Observability Baseline: For direct raw usage telemetry, INFORMATION_SCHEMA queries, and baseline metric calculations, consult bigquery-observability (bigquery_observability). If the bigquery-observability companion skill is not available in the active environment, all optimization guidelines, DDL templates, and decision models across this skill and its reference guides are fully self-contained.
  • Capacity & Editions Modeling: Evaluate the cost-efficiency of migrating workloads from On-Demand to Editions, as well as rightsizing active Edition reservations, baseline commitments, and autoscaling caps.
    • Instructions: Read references/capacity_planning_editions.md to provide deep links to BigQuery's built-in recommendation UIs (e.g., Slot Estimator) and guide the user through UI navigation: 1. navigate to the Slot Estimator tab, 2. select 'On-Demand' as the source to analyze historical query volume, and 3. review the Cost-Optimized Recommendations and Slot Usage Chart.
  • Table & Storage Optimization: Optimize storage costs from a billing model, physical layout, and lifecycle perspective.
    • Billing Architecture: Read references/storage_billing_models.md for guidance on evaluating aggregate compression ratios (e.g. >2:1 threshold in US) to recommend Physical vs. Logical billing, noting that the break-even ratio depends on specific regional rates and custom enterprise contracts. When providing TABLE_STORAGE queries, always scope with WHERE table_schema = '{dataset_id}', use the regional dataset view, and warn that 0 rows indicates a region mismatch or lack of native tables rather than zero billable usage.
    • Partitioning & Clustering Strategy: Read references/table_partitioning_clustering.md to generate production DDL templates (CREATE TABLE, CTAS migrations for unpartitioned tables, and modifying clustering specifications), enforce pruning with require_partition_filter = true, and manage partition limits (up to 10,000 partitions/table).
    • Lifecycle Management: Read references/storage_lifecycle_management.md to pinpoint inactive data and define precise Time-to-Live (TTL) partition expirations, dataset expirations, and Time Travel window reductions.
    • Acceleration Structures: Read references/search_indexes.md to decide when to propose a search index for selective lookups over STRING or JSON data, write the CREATE SEARCH INDEX DDL, and tell the user how to verify index coverage and usage. Read references/materialized_views.md to decide when to propose a materialized view for repeated aggregations or joins over large base tables, write the CREATE MATERIALIZED VIEW DDL, and tell the user how to verify that smart tuning uses it. Read references/bi_engine.md to decide when to propose BI Engine vs. Materialized Views for BI dashboard acceleration, diagnose PARTIAL or DISABLED BI Engine fallback reasons (bi_engine_statistics), and combine BI Engine with materialized views that pre-join or pre-aggregate the data.
  • SQL Optimization: Optimize individual SQL queries to reduce slot-time and the amount of data read.
    • Instructions: Follow the instructions in references/sql_optimization.md to provide recommendations to the user on how to rewrite their SQL query to reduce slot-time and the amount of data read.

Execution Guardrails

  • Terminology & Cost Framing: Never promise or guarantee "cost-reduction" or "reducing expenditure." Always frame recommendations using the terminology "optimizing your bill" or "improving cost-efficiency."
  • Explicit Scope Framing & Region Resolution: Always state the target project_id and region at the very top of your response so the user immediately knows the exact scope being evaluated. Follow this 3-tier resolution hierarchy:
    1. Explicit Region: Use the region specified in the user's prompt (e.g., europe-west1).
    2. Contextual Region: Resolve the region from the specific dataset or resource mentioned in the context.
    3. Unspecified Fallback: Default to us / region-us, explicitly state that us was assumed as the default, and instruct the user to substitute their region if their resources reside elsewhere. Region Formatting: In Cloud Console deep links, use the region identifier directly (e.g., region=us, region=europe-west1). In SQL queries against INFORMATION_SCHEMA, use the regional dataset qualifier (e.g., region-us, region-europe-west1).
  • Zero-Row Result Guard: If querying TABLE_STORAGE with WHERE table_schema = '{dataset_id}' returns 0 rows, do not proceed with an empty or zero-usage evaluation. Treat this as an indicator that the dataset may reside in a different region or have no native tables; stop and prompt the user to confirm the dataset's regional location.
  • Populate Concrete Parameters: When generating URLs and SQL queries, always substitute known project_id and region values directly into the code and links. Never leave literal {project_id} or {location} placeholders for the user to manually edit.
  • No Autonomous Purchasing or Financial Mutations: Never provide the user with executable scripts (e.g., gcloud or bq shell commands like bq update --storage_billing_model=...) designed to autonomously purchase annual commitments, alter edition tier bindings, or mutate storage billing models. Always guide the user to execute commitment purchases, reservation changes, and storage billing model updates manually via the Cloud Console UI.

Files

9
79.4 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related database skillsscan passed

prisma-patterns

Prisma ORM patterns for TypeScript backends — schema design, query optimization, transactions, pagination, and critical traps like updateMany returning count not records, $transaction timeouts, migrate dev resetting the DB, @updatedAt skipped on bulk writes, and serverless connection exhaustion. Use

Scan passed 0
stripe-projects

Use when the user wants to provision infrastructure or third-party services using Stripe Projects. Triggers: "I need a database", "set up auth", "add caching", "give me a Postgres", "provision Redis", "I need hosting", "add a vector DB", "get me an API key for X", "get credentials for X", "sign up f

Scan passed 0
cloudflare-one-migrations

Assess and plan migrations from existing VPN, SWG, or SASE platforms to Cloudflare One, including policy mapping, parity gaps, and rollout.

Scan passed 0
deprecation-and-migration

Manages deprecation and migration. Use when removing old systems, APIs, or features. Use when migrating users from one implementation to another. Use when migrating a database schema in production, such as renaming or dropping a column without downtime (expand/contract). Use when deciding whether to

Scan passed 0
firebase-data-connect

Builds and deploys Firebase SQL Connect (aka Firebase Data Connect) backends with PostgreSQL securely. Use when designing schemas with tables and relations, writing authorized queries and mutations, configuring real-time data updates, or generating type-safe SDKs. Use when you need a relational data

Scan passed 0
redshift-guide

Amazon Redshift is NOT PostgreSQL — corrects PostgreSQL-derived LLM mistakes; covers Redshift-specific SQL, DDL, COPY/UNLOAD, system views, metadata discovery, and operational patterns. Applies ONLY when the task is about Redshift itself (cluster, Serverless workgroup, or Redshift SQL). Pushes back

Scan passed 0