skills/ google/skills

gke-ai-troubleshooting-tpu-mxla-hang

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/contain

0
Installs
—
Rating
—
Success rate
6
Files scanned
Scan passedai-ml
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

6 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 04eae5fada7094e9… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Troubleshoot GKE TPU multi-slice hangs with the MXLA Hang Analyzer

Diagnose Cloud TPU multi-slice training hangs on Google Kubernetes Engine (GKE) by correlating Megascale HANG_DETECTED logs and ML Diagnostics Workload Monitoring Megascale XLA (MXLA) Hang Analyzer reports with 1-minute multi-slice latency metrics (kubernetes.io/container/multislice/*) and GKE node topology labels.


Prerequisites

  • Tools: Install the Google Cloud SDK (gcloud with alpha component for gcloud alpha mldiagnostics) and kubectl.
  • Cloud Billing & Project Configuration: Verify an active billing account is linked (gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com, logging.googleapis.com, monitoring.googleapis.com, and hypercomputecluster.googleapis.com are enabled.
  • Supported workloads and versions: Google Cloud ML Diagnostics only supports JAX on TPUs (see ML Diagnostics platform). Workload Monitoring is enabled by default, supports the jobset and job GKE job types, and is compatible with GKE versions 1.36.0-gke.4681000 and later (Configure GKE for ML Diagnostics, which also covers the cluster setup needed for on-demand profiling in Step 4 Path B). If a workload uses another framework (such as PyTorch) or another custom resource type, gcloud alpha mldiagnostics won't list ML runs or monitored events for it. The Megascale XLA hang analyzer and Megascale XLA metrics require LibTPU 0.40.0 or later (see "Get started" in Workload monitoring with ML Diagnostics).
  • Required IAM Roles:
    • Cluster Director Editor (roles/hypercomputecluster.editor), the role listed in the "IAM permissions" section of ML Diagnostics platform, for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions)
    • Monitoring Viewer (roles/monitoring.viewer) for the PromQL queries in Step 2
    • Logs Viewer (roles/logging.viewer) for the Cloud Logging query in Step 1
    • Kubernetes Engine Viewer (roles/container.viewer) for the kubectl get nodes query in Step 3
    • For remediation ([High Risk] steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
  • Reference Documentation:

Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.

When you recommend a fix, link the doc section that describes it.


MXLA Hang Analyzer routing overview

A Megascale hang occurs when a multi-slice worker has waited on a Megascale communication operation for a set timeout period. The TPU logs then show a Megascale HANG_DETECTED message. HANG_DETECTED is a catch-all signal that the workload isn't progressing, and the cause can be in software or in hardware. When a hang occurs, ML Diagnostics runs the Megascale XLA (MXLA) Hang Analyzer, which reports the likely cause as a code. Consult the Megascale XLA (MXLA) Hang Analyzer section for the definition and recommended action of each code, and route by category:

  • Category A: compiler or HLO divergence (never cordon or replace nodes):
    • When the analyzer reports different HLO modules, inconsistent HLO compilation, or an inconsistent launch order across VMs (such as FINGERPRINT_MISMATCH), route to Path A: Compiler or HLO divergence.
  • Category B: host program queueing or data input stall (never cordon or replace nodes):
  • Category C: hardware or network faults on specific instances:
    • When the analyzer attributes the hang to a TPU chip, SparseCore, ICI, or DCN networking issue, or to an unrecoverable error on specific instances, route to Path C: Hardware or network faults.
  • Category D: hang signal or event exists, but the analyzer is NOT_DETECTED or reports UNKNOWN:
    • When HANG_DETECTED logs or a hang monitored event fired, but the analyzer report has detectionState: "NOT_DETECTED" or reports UNKNOWN ("The MXLA hang was detected but a potential cause is not determined"), route to Path D: No DETECTED analyzer or UNKNOWN cause.

Diagnostic workflow

Step 0: Collect context and set the investigation window [Low Risk]

Collect the target parameters. By default, query a 60-minute window `[T - 30m, T

  • 30m]around{issue_time}`:
  • {project_id}: Google Cloud project ID
  • {location}: Google Cloud region where the ML run and GKE cluster reside (for example, us-central1)
  • {cluster_name}: GKE cluster name
  • {namespace} / {workload_name}: Kubernetes namespace and JobSet/Pod prefix
  • {ml_run_id}: ML Diagnostics run ID
  • {issue_time}: Timestamp when the hang occurred (T, ISO-8601 UTC)
  • {start_time}: T - 30m
  • {end_time}: T + 30m

Step 1: Query HANG_DETECTED logs and hang events [Low Risk]

  1. Check GKE container logs for HANG_DETECTED (read-only Cloud Logging LQL): Query k8s_container logs over [{start_time}, {end_time}] to confirm HANG_DETECTED and identify the first stalled pods:
resource.type="k8s_container"
resource.labels.project_id="{project_id}"
resource.labels.cluster_name="{cluster_name}"
"HANG_DETECTED"
timestamp >= "{start_time}" AND timestamp <= "{end_time}"
  1. List active or recent ML runs: Follow the List machine learning runs section of the ML Diagnostics CLI reference (gcloud alpha mldiagnostics machine-learning-run list) to identify {ml_run_id}.
  2. List and describe hang monitored events: Follow the Monitored-events commands section (gcloud alpha mldiagnostics monitored-events list and gcloud alpha mldiagnostics monitored-events describe) or Access Workload Monitoring information through the API to inspect the Megascale XLA (MXLA) Hang Analyzer report (detectionState, details, and recommendedActions).
    • If HANG_DETECTED logs or a hang monitored event fired, run Step 2 to correlate with the 1-minute metrics; if no analyzer reports detectionState: "DETECTED" or the analyzer reports UNKNOWN, follow Path D: No DETECTED analyzer or UNKNOWN cause.
    • If no HANG_DETECTED log or hang monitored event exists and multi-slice latencies are normal in Step 2, rule out an MXLA hang by following Path E: Healthy telemetry.

Step 2: Correlate with 1-minute multi-slice and TPU metrics [Low Risk]

Consult the System Metrics section of the Workload Monitoring guide for the 1-minute multi-slice network (kubernetes.io/container/multislice/network/*), multi-slice accelerator (kubernetes.io/container/multislice/accelerator/*), and node duty-cycle (kubernetes.io/node/accelerator/duty_cycle) metrics, and run read-only PromQL queries over [{start_time}, {end_time}]:

# 1. P95 multi-slice collective end-to-end latency by pod
histogram_quantile(
  0.95,
  sum by (pod_name, le) (
    rate(kubernetes_io:container_multislice_network_collective_end_to_end_latencies_bucket{
      monitored_resource="k8s_container",
      project_id="{project_id}",
      cluster_name="{cluster_name}"
    }[5m])
  )
)

# 2. P95 multi-slice DCN transfer latency by pod
histogram_quantile(
  0.95,
  sum by (pod_name, le) (
    rate(kubernetes_io:container_multislice_network_dcn_transfer_latencies_bucket{
      monitored_resource="k8s_container",
      project_id="{project_id}",
      cluster_name="{cluster_name}"
    }[5m])
  )
)

# 3. P95 host-to-device transfer latency by pod
histogram_quantile(
  0.95,
  sum by (pod_name, le) (
    rate(kubernetes_io:container_multislice_accelerator_host_to_device_transfer_latencies_bucket{
      monitored_resource="k8s_container",
      project_id="{project_id}",
      cluster_name="{cluster_name}"
    }[5m])
  )
)

# 4. Node TPU duty cycle (Workload Monitoring detects a hang as a prolonged
#    period of minimal to no TPU activity)
kubernetes_io:node_accelerator_duty_cycle{
  monitored_resource="k8s_node",
  project_id="{project_id}",
  cluster_name="{cluster_name}"
}

Step 3: Map culprit Compute Engine instance IDs to GKE nodes [Low Risk]

When the Megascale XLA (MXLA) Hang Analyzer reports culprit numeric Compute Engine instance IDs in details or recommendedActions, map those numeric IDs to GKE Node names and physical topology blocks using this read-only kubectl query inspecting container.googleapis.com/instance_id:

kubectl get nodes -l cloud.google.com/gke-tpu-accelerator \
  -o jsonpath='{range .items[*]}{.metadata.name}{"\tinstance_id="}{.metadata.annotations.container\.googleapis\.com/instance_id}{"\tblock="}{.metadata.labels.cloud\.google\.com/gce-topology-block}{"\tsubblock="}{.metadata.labels.cloud\.google\.com/gce-topology-subblock}{"\thost="}{.metadata.labels.cloud\.google\.com/gce-topology-host}{"\n"}{end}'

Step 4: Remediation by root-cause category

Load and follow only the reference file that matches the root-cause category from Step 1:


Guardrails

  1. Never change GKE-managed instance groups or VMs through Compute Engine: Don't run gcloud compute instance-groups managed commands, such as delete, on a node pool's managed instance group. Handle nodes through GKE, as described in Path C: Hardware or network faults.
  2. Never cordon or replace nodes for compiler divergence or input stalls: If the analyzer reports a compiler or HLO divergence, a program queueing issue, or a data input stall, don't cordon or replace TPU nodes. Follow Path A: Compiler or HLO divergence or Path B: Host program queueing or data input stall instead.

Files

6
21.1 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related ai-ml skillsscan passed

exa-search

Neural search via Exa MCP for web, code, and company research. Use when the user needs web search, code examples, company intel, people lookup, or AI-powered deep research with Exa's neural search engine.

Scan passed 0
pair-agent

Pair a remote AI agent with your browser. (gstack)

Scan passed 0
ce-noslop

Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market

Scan passed 0
superjson

Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi

Scan passed 0
amazon-workspaces-agent-access

Connects AI agents to remote Windows desktop applications on Amazon WorkSpaces Applications (AppStream 2.0) through the managed Agent Access MCP server, and guides reliable desktop automation. Covers connecting an agent to the MCP endpoint (SigV4, streaming URL, and Active Directory SAML/Domain Join

Scan passed 0
use-case-specification

Creates a reusable use case specification file that defines the business problem, stakeholders, and measurable success criteria for model customization, as recommended by the AWS Responsible AI Lens. Use as the default first step in any model customization plan. Skip only if the user explicitly decl

Scan passed 0