skills/ google/skills

gke-cost-analysis

Answer natural language questions and perform analysis on GKE cluster and workload costs using BigQuery billing exports, cost allocation data, and live cluster monitoring metrics. Use when querying GKE costs across projects, namespaces, or workloads, analyzing billing reports in BigQuery (`bq`), che

0
Installs
—
Rating
—
Success rate
2
Files scanned
Scan passeddevops
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

2 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 b7fffc96ce046087… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

GKE Cost Analysis

This skill provides guidance on answering natural language questions about GKE-related costs, billing reports, and utilization analysis.

Overview

When users ask about GKE costs (e.g., "What are my costs across projects?", "What's my most expensive namespace?", "Why is my cluster cost spiking?"), use this skill to provide a structured and expert response using BigQuery billing exports, cost allocation metadata, and live cluster metrics.

Instructions

When handling a cost-related question:

  1. Provide a Direct Answer: Address the specific cost question or analytical request clearly and concisely.
  2. Explain BigQuery Integration: Explain how to query BigQuery for historical cost breakdown. Note that GKE costs originate from the GCP Billing Detailed BigQuery Export (gcp_billing_export_resource_v1_*).
  3. Check & Verify Cost Allocation: Explain that GKE Cost Allocation must be enabled on the cluster (--enable-cost-allocation) for namespace, label, and workload-level billing granularity. If queries return empty labels, provide the gcloud command to enable it.
  4. Analyze Pricing Drivers & Utilization: When diagnosing cost drivers, explain whether the cluster is in Autopilot (billed by requested pod CPU/memory) or Standard mode (billed by underlying VM node size + control plane fees), and compare live utilization (kubectl top) against provisioned requests.
  5. Provide Actionable Commands/Queries: Provide concrete BigQuery CLI (bq query) commands or read-only gcloud/kubectl inspection commands. Prefer bq over BigQuery Studio when available.

Key Points & Pricing Drivers

  • Data Source: GKE costs come from GCP Billing Detailed BigQuery Export. The user must provide the full path to their BigQuery table (dataset name and table name containing the Billing Account ID).
  • Granularity Requirement: GKE Cost Allocation (--enable-cost-allocation) must be enabled on the cluster to populate goog-k8s-cluster-name, k8s-namespace, k8s-workload-name, and k8s-workload-type labels in BigQuery.
  • Autopilot vs. Standard Cost Drivers:
    • Autopilot Pricing: Billed directly on pod resource requests (requests.cpu, requests.memory, ephemeral storage). Over-requested pods drive up billing regardless of whether the pod actively uses those CPU cycles or memory.
    • Standard Pricing: Billed on provisioned node pool VMs (e2, n4, c3, etc.). Idle nodes or multiple low-utilization dev clusters drive excess infrastructure costs.
    • Cluster Management Fee: ~$0.10/hour per cluster applies to BOTH Standard and Autopilot modes. The free tier waives it for one eligible cluster per billing account.
  • Credits & Discounts Impact: When analyzing cost versus cost_before_credits, note that Committed Use Discounts (CUDs) and Spot VMs appear as credits or reduced rate charges in the billing export.
  • Tools & Syntax: BigQuery CLI (bq) is preferred. When writing Standard SQL queries, use a dot (.) instead of a colon (:) to separate the project ID and dataset name ({project_id}.{dataset_name}.{table_name}).
  • Defaults: Assume last 30 days, row limit 10, ordering by cost descending (ORDER BY cost DESC), unless specified otherwise.

Live Cluster & Cost Monitoring

Use read-only CLI commands to inspect current cluster budgets, node utilization, and pod resource consumption vs. requests:

# View billing budgets for an account (requires Cost Management API)
gcloud billing budgets list --billing-account={billing_account} --quiet

# View live node resource utilization across the cluster
kubectl top nodes

# View pod resource usage across namespaces (compare against requested limits to diagnose waste)
kubectl top pods --all-namespaces --containers

Warning — cluster mutation, not read-only: Enabling GKE cost allocation modifies the cluster. Get explicit user confirmation before running it, and note that namespace/workload labels populate in the billing export only from enablement onward (no historical backfill).

gcloud container clusters update {cluster_name} \
    --enable-cost-allocation \
    --region {region}

Applying Cost Optimizations

To apply rightsizing changes based on analysis (such as setting up VPA recommendation mode, adjusting CPU/memory to P95 * 1.2, configuring Spot VMs via nodeSelector or ComputeClass, enforcing ResourceQuotas, or selecting machine types and CUDs), use the gke-cost-optimization skill.

BigQuery Query Templates

Ready-to-adapt bq query templates — single workload cost, per-workload per-cluster breakdown, per-namespace breakdown — with the placeholder policy and defaults (30 days, LIMIT 10, ORDER BY cost DESC) are in references/billing-queries.md. All parameters (dataset, table, project, cluster, etc.) must be replaced with user values.

Note: Checking that the goog-k8s-cluster-name label exists scopes the total billing data specifically to GKE costs.

Files

2
9.3 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related devops skillsscan passed