skills/ google/skills

gke-ai-troubleshooting-tpu-performance-degradation

Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system metrics (`kubernetes.io/n

0
Installs
—
Rating
—
Success rate
5
Files scanned
Scan passedai-ml
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

5 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 f128a2a45effb579… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Troubleshoot GKE TPU performance degradation with ML Diagnostics Workload Monitoring

Diagnose and mitigate Cloud TPU training throughput drops and step-time regressions (15%+ drop in TPU duty cycle) on Google Kubernetes Engine (GKE) by correlating ML Diagnostics Workload Monitoring MonitoredEvent analyzer reports with 1-minute Cloud Monitoring system metrics and GKE node topology labels.


Prerequisites

  • Tools: Install the Google Cloud SDK (gcloud with alpha component for gcloud alpha mldiagnostics) and kubectl.
  • Cloud Billing & Project Configuration: Verify an active billing account is linked (gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com, monitoring.googleapis.com, and hypercomputecluster.googleapis.com are enabled.
  • Supported workloads and GKE versions: Google Cloud ML Diagnostics only supports JAX on TPUs (see ML Diagnostics platform). Workload Monitoring is enabled by default, supports the jobset and job GKE job types, and is compatible with GKE versions 1.36.0-gke.4681000 and later, as stated in Configure GKE for ML Diagnostics. If a workload uses another framework (such as PyTorch) or another custom resource type, gcloud alpha mldiagnostics won't list ML runs or monitored events for it. On-demand profiling (Step 4 Path A) additionally requires the cluster setup described in that document.
  • Required IAM Roles:
    • Cluster Director Editor (roles/hypercomputecluster.editor), the role listed in the "IAM permissions" section of ML Diagnostics platform, for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions)
    • Monitoring Viewer (roles/monitoring.viewer) for the PromQL queries in Step 2
    • Kubernetes Engine Viewer (roles/container.viewer) for the kubectl get nodes query in Step 3
    • For remediation ([High Risk] steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
  • Reference Documentation:

Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.

When you recommend a fix, link the doc section that describes it.


Analyzer routing overview

By default, Workload Monitoring treats a 15% drop in TPU duty cycle as a performance degradation, raises a MonitoredEvent (type: PERFORMANCE_DEGRADATION), and runs its analyzers. Consult the Workload Monitoring and analyzers section in the official Cloud TPU documentation for the complete analyzer definitions, detection criteria, and subsections, and route remediation by category:

  • Category A: workload resource bottlenecks (never cordon or replace nodes):
  • Category B: infrastructure and network fabric throttling (culprit nodes):
  • Category C: PERFORMANCE_DEGRADATION event fired, but no analyzer reports detectionState: "DETECTED":

Diagnostic workflow

Step 0: Collect context and set the investigation window [Low Risk]

Collect the target parameters. By default, query a 60-minute window `[T - 30m, T

  • 30m]around{issue_time}`:
  • {project_id}: Google Cloud project ID
  • {location}: Google Cloud region where the ML run and GKE cluster reside (for example, us-central1)
  • {cluster_name}: GKE cluster name
  • {workload_name} / {ml_run_id}: JobSet / workload name or ML Diagnostics run ID
  • {issue_time}: Timestamp when throughput degradation was observed (T, ISO-8601 UTC)
  • {start_time}: T - 30m
  • {end_time}: T + 30m

Step 1: Query ML runs and monitoredEvents [Low Risk]

  1. List active or recent ML runs: Give the user gcloud alpha mldiagnostics machine-learning-run list with a link to List machine learning runs in the ML Diagnostics CLI reference, and locate {ml_run_id} matching {workload_name}.
  2. List performance degradation events: Give the user gcloud alpha mldiagnostics monitored-events list with a link to Monitored-events commands, or the hypercomputecluster.googleapis.com/v1alpha API with a link to Access Workload Monitoring information through the API, to check for PERFORMANCE_DEGRADATION events during [{start_time}, {end_time}].
  3. Describe the MonitoredEvent: Give the user gcloud alpha mldiagnostics monitored-events describe with a link to the same Monitored-events commands section, and inspect the analyzerReports array (analyzer, detectionState, details, and recommendedActions).
    • If a PERFORMANCE_DEGRADATION event fired, run Step 2 to corroborate with the 1-minute system metrics; if none of its analyzerReports entries has detectionState: "DETECTED", follow Path C: Event fired with no DETECTED analyzer.
    • If no PERFORMANCE_DEGRADATION event exists and duty cycle is steady in Step 2, rule out TPU performance degradation by following Path D: Healthy telemetry.

Step 2: Correlate with 1-minute Cloud Monitoring system metrics [Low Risk]

Consult the System Metrics section in the Workload Monitoring documentation for the 1-minute Cloud Monitoring metrics exported for TPU and host devices, and run read-only PromQL queries over [{start_time}, {end_time}] to corroborate the analyzer report:

# 1. Node TPU duty cycle (look for the drop on the affected nodes)
kubernetes_io:node_accelerator_duty_cycle{
  monitored_resource="k8s_node",
  project_id="{project_id}",
  cluster_name="{cluster_name}"
}

# 2. HBM utilization ratio by node (around 0.90 is approaching the limit)
sum by (node_name) (
  kubernetes_io:node_accelerator_memory_used{
    monitored_resource="k8s_node",
    project_id="{project_id}",
    cluster_name="{cluster_name}"
  }
)
/
sum by (node_name) (
  kubernetes_io:node_accelerator_memory_total{
    monitored_resource="k8s_node",
    project_id="{project_id}",
    cluster_name="{cluster_name}"
  }
)

# 3. Host memory and CPU allocatable utilization
kubernetes_io:node_memory_allocatable_utilization{
  monitored_resource="k8s_node",
  project_id="{project_id}",
  cluster_name="{cluster_name}"
}

kubernetes_io:node_cpu_allocatable_utilization{
  monitored_resource="k8s_node",
  project_id="{project_id}",
  cluster_name="{cluster_name}"
}

Step 3: Map culprit instance IDs to GKE nodes and topology [Low Risk]

When an infrastructure analyzer reports culprit numeric Compute Engine instance IDs in details or recommendedActions, map those numeric instance IDs to GKE Node names and physical topology blocks using this read-only kubectl query inspecting container.googleapis.com/instance_id:

kubectl get nodes -l cloud.google.com/gke-tpu-accelerator \
  -o jsonpath='{range .items[*]}{.metadata.name}{"\tinstance_id="}{.metadata.annotations.container\.googleapis\.com/instance_id}{"\tblock="}{.metadata.labels.cloud\.google\.com/gce-topology-block}{"\tsubblock="}{.metadata.labels.cloud\.google\.com/gce-topology-subblock}{"\thost="}{.metadata.labels.cloud\.google\.com/gce-topology-host}{"\n"}{end}'

Step 4: Remediation by analyzer category

Load and follow only the reference file that matches the analyzer category from Step 1:


Guardrails

  1. Never change GKE-managed instance groups or VMs through Compute Engine: Don't run gcloud compute instance-groups managed commands, such as delete, on a node pool's managed instance group. Handle nodes through GKE, as described in Path B: Infrastructure or network fabric throttling.
  2. Never cordon nodes for workload resource saturation: If only the HBM capacity, host memory, or host CPU utilization analyzers detected an issue, don't cordon or replace nodes.

Files

5
19.9 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related ai-ml skillsscan passed

exa-search

Neural search via Exa MCP for web, code, and company research. Use when the user needs web search, code examples, company intel, people lookup, or AI-powered deep research with Exa's neural search engine.

Scan passed 0
pair-agent

Pair a remote AI agent with your browser. (gstack)

Scan passed 0
ce-noslop

Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market

Scan passed 0
superjson

Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi

Scan passed 0
amazon-workspaces-agent-access

Connects AI agents to remote Windows desktop applications on Amazon WorkSpaces Applications (AppStream 2.0) through the managed Agent Access MCP server, and guides reliable desktop automation. Covers connecting an agent to the MCP endpoint (SigV4, streaming URL, and Active Directory SAML/Domain Join

Scan passed 0
use-case-specification

Creates a reusable use case specification file that defines the business problem, stakeholders, and measurable success criteria for model customization, as recommended by the AWS Responsible AI Lens. Use as the default first step in any model customization plan. Skip only if the user explicitly decl

Scan passed 0