gke-ai-troubleshooting-tpu-metrics-monitoring
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics fo
- 0
- Installs
- —
- Rating
- —
- Success rate
- 3
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 b5cdcedffc5d7c18… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
GKE TPU Metrics Monitoring Guide
This skill enables the agent to monitor GKE TPU workloads, nodes, and node pools using GKE system metrics. It helps diagnose if workload interruptions or performance issues are caused by underlying infrastructure.
Step 0: Mandatory Context
Independently gather required context (such as cluster details or node pool names) using available GKE and Cloud tools, or use the provided {variable} placeholders:
{project_id}: The GCP Project ID.{cluster_name}: The GKE Cluster Name.{location}: The GKE Cluster Location (region or zone).{node_name}: (Optional) The name of the specific GKE node.{node_pool_name}: (Optional) The name of the GKE node pool.
Diagnostic Steps
Step 1: Verify TPU Runtime Metrics Configuration [Low Risk] [Auto]
Before analyzing runtime metrics, verify that the workload is configured to export them. This ensures the cluster and container environment are set up for automated metric scraping and visibility into accelerator health.
- Action: Verify that the Pod specification and cluster meet the following prerequisites:
containerPort: 8431exposed on the TPU container (required for Prometheus metric scraping).- JAX version
0.4.14or later if using JAX (earlier versions do not export runtime metrics). - GKE version is
1.27.4-gke.900or later (required for TPU runtime metric support). - GKE System Metrics are enabled on the cluster (required for Cloud Monitoring ingestion).
Step 2: Monitor TPU Runtime Metrics [Low Risk] [Auto]
If configured correctly, the following metrics are available in Cloud Monitoring (monitored resources k8s_node and k8s_container):
- Container Metrics:
kubernetes.io/container/accelerator/duty_cycle: Percentage of time over the past sampling period (60 seconds) during which the TensorCores were actively processing on a TPU chip.kubernetes.io/container/accelerator/memory_used: Amount of accelerator memory allocated in bytes.kubernetes.io/container/accelerator/memory_total: Total accelerator memory in bytes.
- Node Metrics:
kubernetes.io/node/accelerator/duty_cyclekubernetes.io/node/accelerator/memory_usedkubernetes.io/node/accelerator/memory_total
Step 3: Check Node Status Condition [Low Risk] [Auto]
Query the status condition of GKE nodes (GKE version 1.32.1-gke.1357001 or later).
- PromQL Query (Check if a specific node is Ready):
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", node_name="{node_name}", condition="Ready", status="True"} - PromQL Query (List nodes with non-Ready conditions that are True):
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition!="Ready", status="True"} - PromQL Query (List nodes that are NOT Ready):
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition="Ready", status="False"} - PromQL Query (Fleet-wide node status):
avg by (condition,status)(avg_over_time(kubernetes_io:node_status_condition{monitored_resource="k8s_node"}[5m]))
Step 4: Check Node Pool Status [Low Risk] [Auto]
Query the status of multi-host TPU node pools.
- PromQL Query (Verify if a specific node pool is Running):
kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}", node_pool_name="{node_pool_name}", status="Running"} - PromQL Query (Monitor node pools grouped by status):
Possible statuses:count by (status)(count_over_time(kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool"}[5m]))Provisioning,Running,Error,Reconciling,Stopping.
Step 5: Check Node Pool Availability [Low Risk] [Auto]
Query if all nodes in a multi-host TPU node pool are available.
- PromQL Query (Check availability over time):
Value:avg by (node_pool_name)(avg_over_time(kubernetes_io:node_pool_multi_host_available{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[5m]))1(True, all nodes available) or0(False, some nodes unavailable).
Step 6: Analyze Node Interruptions [Low Risk] [Auto]
Query the count of interruptions for GKE nodes.
- PromQL Query (Breakdown of interruptions and causes):
Interruption Types:sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node"}[5m]))TerminationEvent,MaintenanceEvent,PreemptionEvent. Interruption Reasons:HostError,Eviction,AutoRepair. - PromQL Query (Filter for Host Maintenance events):
sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", interruption_reason="HW/SW Maintenance"}[5m])) - PromQL Query (Interruption count aggregated by node pool):
sum by (node_pool_name,interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_pool_interruption_count{monitored_resource="k8s_node_pool", interruption_reason="HW/SW Maintenance", node_pool_name="{node_pool_name}"}[5m]))
Step 7: Calculate Recovery and Interruption Metrics [Low Risk] [Auto]
Calculate Mean Time to Recovery (MTTR) and Mean Time Between Interruptions (MTBI) over the last 7 days.
- PromQL Query (MTTR - Mean Time to Recovery):
sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_sum{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_count{monitored_resource="k8s_node_pool",cluster_name="{cluster_name}"}[7d])) - PromQL Query (MTBI - Mean Time Between Interruptions):
sum(count_over_time(kubernetes_io:node_memory_total_bytes{monitored_resource="k8s_node", node_name=~"gke-tpu.*|gk3-tpu.*", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", node_name=~"gke-tpu.*|gk3-tpu.*", cluster_name="{cluster_name}"}[7d]))
Step 8: Monitor TPU Host Metrics [Low Risk] [Auto]
For GKE version 1.28.1-gke.1066000 or later, monitor TPU host performance.
- Container Metrics:
kubernetes.io/container/accelerator/tensorcore_utilization: Current percentage of the TensorCore that is utilized.kubernetes.io/container/accelerator/memory_bandwidth_utilization: Current percentage of the accelerator memory bandwidth that is being used.
- Node Metrics:
kubernetes.io/node/accelerator/tensorcore_utilizationkubernetes.io/node/accelerator/memory_bandwidth_utilization
Files
3- SKILL.md
6e75a9c3797.3 KB - references/failure_signatures.md
c368a9a04c1.4 KB - scripts/validate_queries.sh
23083e779a348 B
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related devops skillsscan passed
Design, configure, troubleshoot, or review Cloudflare One Zero Trust and SASE deployments. Use cloudflare-one-migrations for migration planning from other vendors.
Configure deployment settings for /land-and-deploy.
Deploy tRPC on WinterCG-compliant edge runtimes with fetchRequestHandler() from @trpc/server/adapters/fetch. Supports Cloudflare Workers, Deno Deploy, Vercel Edge Runtime, Astro, Remix, SolidStart. FetchCreateContextFnOptions provides req (Request) and resHeaders (Headers) for context creation. The
Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
Use when managing an Uncloud cluster — deploying services, configuring Caddy ingress, adding static proxy routes for non-cluster devices, publishing ports, scaling, inspecting logs, or managing machines and volumes with the `uc` CLI.
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil