gke-cost-analysis
Answer natural language questions and perform analysis on GKE cluster and workload costs using BigQuery billing exports, cost allocation data, and live cluster monitoring metrics. Use when querying GKE costs across projects, namespaces, or workloads, analyzing billing reports in BigQuery (`bq`), che
- 0
- Installs
- —
- Rating
- —
- Success rate
- 2
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 b7fffc96ce046087… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
GKE Cost Analysis
This skill provides guidance on answering natural language questions about GKE-related costs, billing reports, and utilization analysis.
Overview
When users ask about GKE costs (e.g., "What are my costs across projects?", "What's my most expensive namespace?", "Why is my cluster cost spiking?"), use this skill to provide a structured and expert response using BigQuery billing exports, cost allocation metadata, and live cluster metrics.
Instructions
When handling a cost-related question:
- Provide a Direct Answer: Address the specific cost question or analytical request clearly and concisely.
- Explain BigQuery Integration: Explain how to query BigQuery for
historical cost breakdown. Note that GKE costs originate from the GCP
Billing Detailed BigQuery Export (
gcp_billing_export_resource_v1_*). - Check & Verify Cost Allocation: Explain that GKE Cost Allocation must be
enabled on the cluster (
--enable-cost-allocation) for namespace, label, and workload-level billing granularity. If queries return empty labels, provide thegcloudcommand to enable it. - Analyze Pricing Drivers & Utilization: When diagnosing cost drivers,
explain whether the cluster is in Autopilot (billed by requested pod
CPU/memory) or Standard mode (billed by underlying VM node size + control
plane fees), and compare live utilization (
kubectl top) against provisioned requests. - Provide Actionable Commands/Queries: Provide concrete BigQuery CLI (
bq query) commands or read-onlygcloud/kubectlinspection commands. Preferbqover BigQuery Studio when available.
Key Points & Pricing Drivers
- Data Source: GKE costs come from GCP Billing Detailed BigQuery Export. The user must provide the full path to their BigQuery table (dataset name and table name containing the Billing Account ID).
- Granularity Requirement: GKE Cost Allocation
(
--enable-cost-allocation) must be enabled on the cluster to populategoog-k8s-cluster-name,k8s-namespace,k8s-workload-name, andk8s-workload-typelabels in BigQuery. - Autopilot vs. Standard Cost Drivers:
- Autopilot Pricing: Billed directly on pod resource requests
(
requests.cpu,requests.memory, ephemeral storage). Over-requested pods drive up billing regardless of whether the pod actively uses those CPU cycles or memory. - Standard Pricing: Billed on provisioned node pool VMs (
e2,n4,c3, etc.). Idle nodes or multiple low-utilization dev clusters drive excess infrastructure costs. - Cluster Management Fee: ~$0.10/hour per cluster applies to BOTH Standard and Autopilot modes. The free tier waives it for one eligible cluster per billing account.
- Autopilot Pricing: Billed directly on pod resource requests
(
- Credits & Discounts Impact: When analyzing
costversuscost_before_credits, note that Committed Use Discounts (CUDs) and Spot VMs appear as credits or reduced rate charges in the billing export. - Tools & Syntax: BigQuery CLI (
bq) is preferred. When writing Standard SQL queries, use a dot (.) instead of a colon (:) to separate the project ID and dataset name ({project_id}.{dataset_name}.{table_name}). - Defaults: Assume last 30 days, row limit 10, ordering by cost descending
(
ORDER BY cost DESC), unless specified otherwise.
Live Cluster & Cost Monitoring
Use read-only CLI commands to inspect current cluster budgets, node utilization, and pod resource consumption vs. requests:
# View billing budgets for an account (requires Cost Management API)
gcloud billing budgets list --billing-account={billing_account} --quiet
# View live node resource utilization across the cluster
kubectl top nodes
# View pod resource usage across namespaces (compare against requested limits to diagnose waste)
kubectl top pods --all-namespaces --containers
Warning — cluster mutation, not read-only: Enabling GKE cost allocation modifies the cluster. Get explicit user confirmation before running it, and note that namespace/workload labels populate in the billing export only from enablement onward (no historical backfill).
gcloud container clusters update {cluster_name} \ --enable-cost-allocation \ --region {region}
Applying Cost Optimizations
To apply rightsizing changes based on analysis (such as setting up VPA
recommendation mode, adjusting CPU/memory to P95 * 1.2, configuring Spot VMs
via nodeSelector or ComputeClass, enforcing ResourceQuotas, or selecting
machine types and CUDs), use the gke-cost-optimization skill.
BigQuery Query Templates
Ready-to-adapt bq query templates — single workload cost, per-workload
per-cluster breakdown, per-namespace breakdown — with the placeholder policy
and defaults (30 days, LIMIT 10, ORDER BY cost DESC) are in
references/billing-queries.md. All parameters
(dataset, table, project, cluster, etc.) must be replaced with user values.
Note: Checking that the goog-k8s-cluster-name label exists scopes the total
billing data specifically to GKE costs.
Files
2- SKILL.md
af70936ecb5.8 KB - references/billing-queries.md
076986c8903.5 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related devops skillsscan passed
Pre-deployment checks for router and switch configuration, including dangerous commands, duplicate addresses, subnet overlaps, stale references, management-plane risk, and IOS-style security hygiene. Use when reviewing a router or switch configuration before deployment.
Post-deploy canary monitoring. (gstack)
Migrate Cloudflare Sandbox apps from stable @cloudflare/sandbox to @cloudflare/sandbox@next (SDK 1.0 preview). Use sandbox-next for apps already on the preview.
Deploy tRPC on WinterCG-compliant edge runtimes with fetchRequestHandler() from @trpc/server/adapters/fetch. Supports Cloudflare Workers, Deno Deploy, Vercel Edge Runtime, Astro, Remix, SolidStart. FetchCreateContextFnOptions provides req (Request) and resHeaders (Headers) for context creation. The
Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil