gke-ai-troubleshooting-tpu-mxla-hang
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/contain
- 0
- Installs
- —
- Rating
- —
- Success rate
- 6
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 04eae5fada7094e9… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Troubleshoot GKE TPU multi-slice hangs with the MXLA Hang Analyzer
Diagnose Cloud TPU multi-slice training hangs on Google Kubernetes Engine (GKE)
by correlating Megascale HANG_DETECTED logs and ML Diagnostics Workload
Monitoring Megascale XLA (MXLA) Hang Analyzer reports with 1-minute
multi-slice latency metrics (kubernetes.io/container/multislice/*) and GKE
node topology labels.
Prerequisites
- Tools: Install the
Google Cloud SDK (
gcloudwithalphacomponent forgcloud alpha mldiagnostics) andkubectl. - Cloud Billing & Project Configuration: Verify an active billing account is
linked (
gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensurecontainer.googleapis.com,logging.googleapis.com,monitoring.googleapis.com, andhypercomputecluster.googleapis.comare enabled. - Supported workloads and versions: Google Cloud ML Diagnostics only
supports JAX on TPUs (see
ML Diagnostics platform).
Workload Monitoring is enabled by default, supports the
jobsetandjobGKE job types, and is compatible with GKE versions1.36.0-gke.4681000and later (Configure GKE for ML Diagnostics, which also covers the cluster setup needed for on-demand profiling in Step 4 Path B). If a workload uses another framework (such as PyTorch) or another custom resource type,gcloud alpha mldiagnosticswon't list ML runs or monitored events for it. The Megascale XLA hang analyzer and Megascale XLA metrics require LibTPU0.40.0or later (see "Get started" in Workload monitoring with ML Diagnostics). - Required IAM Roles:
- Cluster Director Editor (
roles/hypercomputecluster.editor), the role listed in the "IAM permissions" section of ML Diagnostics platform, for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions) - Monitoring Viewer (
roles/monitoring.viewer) for the PromQL queries in Step 2 - Logs Viewer (
roles/logging.viewer) for the Cloud Logging query in Step 1 - Kubernetes Engine Viewer (
roles/container.viewer) for thekubectl get nodesquery in Step 3 - For remediation (
[High Risk]steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
- Cluster Director Editor (
- Reference Documentation:
- ML Diagnostics platform (Sections: "IAM permissions")
- Configure GKE for ML Diagnostics (Sections: "Set up with gcloud CLI, Google Cloud console, or Terraform", "Manual installation", "Connection-operator")
- Workload monitoring with ML Diagnostics (Sections: "Get started", "Megascale XLA (MXLA) Hang Analyzer", "Access Workload Monitoring information through the API", "System Metrics")
- Get started with the ML Diagnostics CLI (Sections: "List machine learning runs", "Monitored-events commands", "List profiler targets", "Capture on-demand profiler sessions")
- Get support (Sections: "Before you contact support")
- Auto-repair nodes (Sections: "Repair criteria", "Verify node auto-repair is enabled for a Standard node pool", "Enable auto-repair for an existing Standard node pool", "Node auto repair in TPU slice nodes")
- Deploy TPU workloads in GKE Standard (Sections: "Configure auto repair for TPU slice nodes")
- Dump HLO Computations (OpenXLA)
Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
MXLA Hang Analyzer routing overview
A Megascale hang occurs when a multi-slice worker has waited on a Megascale
communication operation for a set timeout period. The TPU logs then show a
Megascale HANG_DETECTED message. HANG_DETECTED is a catch-all signal that
the workload isn't progressing, and the cause can be in software or in hardware.
When a hang occurs, ML Diagnostics runs the Megascale XLA (MXLA) Hang Analyzer, which reports the likely cause as a code. Consult the
Megascale XLA (MXLA) Hang Analyzer
section for the definition and recommended action of each code, and route by
category:
- Category A: compiler or HLO divergence (never cordon or replace nodes):
- When the analyzer reports different HLO modules, inconsistent HLO
compilation, or an inconsistent launch order across VMs (such as
FINGERPRINT_MISMATCH), route to Path A: Compiler or HLO divergence.
- When the analyzer reports different HLO modules, inconsistent HLO
compilation, or an inconsistent launch order across VMs (such as
- Category B: host program queueing or data input stall (never cordon or
replace nodes):
- When the analyzer reports that VMs aren't queuing programs to the TPU or
that data input stalled (such as
DATA_INPUT_STALL), route to Path B: Host program queueing or data input stall.
- When the analyzer reports that VMs aren't queuing programs to the TPU or
that data input stalled (such as
- Category C: hardware or network faults on specific instances:
- When the analyzer attributes the hang to a TPU chip, SparseCore, ICI, or DCN networking issue, or to an unrecoverable error on specific instances, route to Path C: Hardware or network faults.
- Category D: hang signal or event exists, but the analyzer is
NOT_DETECTEDor reportsUNKNOWN:- When
HANG_DETECTEDlogs or a hang monitored event fired, but the analyzer report hasdetectionState: "NOT_DETECTED"or reportsUNKNOWN("The MXLA hang was detected but a potential cause is not determined"), route to Path D: No DETECTED analyzer or UNKNOWN cause.
- When
Diagnostic workflow
Step 0: Collect context and set the investigation window [Low Risk]
Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
- 30m]
around{issue_time}`:
{project_id}: Google Cloud project ID{location}: Google Cloud region where the ML run and GKE cluster reside (for example,us-central1){cluster_name}: GKE cluster name{namespace}/{workload_name}: Kubernetes namespace and JobSet/Pod prefix{ml_run_id}: ML Diagnostics run ID{issue_time}: Timestamp when the hang occurred (T, ISO-8601 UTC){start_time}:T - 30m{end_time}:T + 30m
Step 1: Query HANG_DETECTED logs and hang events [Low Risk]
- Check GKE container logs for
HANG_DETECTED(read-only Cloud Logging LQL): Queryk8s_containerlogs over[{start_time}, {end_time}]to confirmHANG_DETECTEDand identify the first stalled pods:
resource.type="k8s_container"
resource.labels.project_id="{project_id}"
resource.labels.cluster_name="{cluster_name}"
"HANG_DETECTED"
timestamp >= "{start_time}" AND timestamp <= "{end_time}"
- List active or recent ML runs: Follow the
List machine learning runs
section of the ML Diagnostics CLI reference (
gcloud alpha mldiagnostics machine-learning-run list) to identify{ml_run_id}. - List and describe hang monitored events: Follow the
Monitored-events commands
section (
gcloud alpha mldiagnostics monitored-events listandgcloud alpha mldiagnostics monitored-events describe) or Access Workload Monitoring information through the API to inspect theMegascale XLA (MXLA) Hang Analyzerreport (detectionState,details, andrecommendedActions).- If
HANG_DETECTEDlogs or a hang monitored event fired, run Step 2 to correlate with the 1-minute metrics; if no analyzer reportsdetectionState: "DETECTED"or the analyzer reportsUNKNOWN, follow Path D: No DETECTED analyzer or UNKNOWN cause. - If no
HANG_DETECTEDlog or hang monitored event exists and multi-slice latencies are normal in Step 2, rule out an MXLA hang by following Path E: Healthy telemetry.
- If
Step 2: Correlate with 1-minute multi-slice and TPU metrics [Low Risk]
Consult the
System Metrics
section of the Workload Monitoring guide for the 1-minute multi-slice network
(kubernetes.io/container/multislice/network/*), multi-slice accelerator
(kubernetes.io/container/multislice/accelerator/*), and node duty-cycle
(kubernetes.io/node/accelerator/duty_cycle) metrics, and run read-only PromQL
queries over [{start_time}, {end_time}]:
# 1. P95 multi-slice collective end-to-end latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_network_collective_end_to_end_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 2. P95 multi-slice DCN transfer latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_network_dcn_transfer_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 3. P95 host-to-device transfer latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_accelerator_host_to_device_transfer_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 4. Node TPU duty cycle (Workload Monitoring detects a hang as a prolonged
# period of minimal to no TPU activity)
kubernetes_io:node_accelerator_duty_cycle{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}
Step 3: Map culprit Compute Engine instance IDs to GKE nodes [Low Risk]
When the Megascale XLA (MXLA) Hang Analyzer reports culprit numeric Compute
Engine instance IDs in details or recommendedActions, map those numeric IDs
to GKE Node names and physical topology blocks using this read-only kubectl
query inspecting container.googleapis.com/instance_id:
kubectl get nodes -l cloud.google.com/gke-tpu-accelerator \
-o jsonpath='{range .items[*]}{.metadata.name}{"\tinstance_id="}{.metadata.annotations.container\.googleapis\.com/instance_id}{"\tblock="}{.metadata.labels.cloud\.google\.com/gce-topology-block}{"\tsubblock="}{.metadata.labels.cloud\.google\.com/gce-topology-subblock}{"\thost="}{.metadata.labels.cloud\.google\.com/gce-topology-host}{"\n"}{end}'
Step 4: Remediation by root-cause category
Load and follow only the reference file that matches the root-cause category from Step 1:
- Path A (compiler or HLO divergence — different HLO modules, inconsistent HLO
compilation, or inconsistent launch order across VMs, such as
FINGERPRINT_MISMATCH): Read Path A: Compiler or HLO divergence. - Path B (host program queueing or data input stall —
PROGRAM_NOT_QUEUEDorDATA_INPUT_STALL): Read Path B: Host program queueing or data input stall. - Path C (hardware, SparseCore, ICI, DCN networking faults, or
UNRECOVERABLE_ERRORon specific instances): Read Path C: Hardware or network faults. - Path D (
HANG_DETECTEDlogs or hang event fired, but the analyzer isNOT_DETECTEDor reportsUNKNOWN): Read Path D: No DETECTED analyzer or UNKNOWN cause. - Path E (no
HANG_DETECTEDlogs or hang events, and steady telemetry): Read Path E: Healthy telemetry.
Guardrails
- Never change GKE-managed instance groups or VMs through Compute Engine:
Don't run
gcloud compute instance-groups managedcommands, such asdelete, on a node pool's managed instance group. Handle nodes through GKE, as described in Path C: Hardware or network faults. - Never cordon or replace nodes for compiler divergence or input stalls: If the analyzer reports a compiler or HLO divergence, a program queueing issue, or a data input stall, don't cordon or replace TPU nodes. Follow Path A: Compiler or HLO divergence or Path B: Host program queueing or data input stall instead.
Files
6- SKILL.md
1a467aa00d14.5 KB - references/path-a-compiler-hlo-divergence.md
3c585d1f9d964 B - references/path-b-program-queueing-input-stall.md
f1e406dec01.4 KB - references/path-c-hardware-network-faults.md
0c36c972272.5 KB - references/path-d-no-detected-or-unknown.md
49bbeef3e61.5 KB - references/path-e-healthy-telemetry.md
e875f2316c272 B
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related ai-ml skillsscan passed
Neural search via Exa MCP for web, code, and company research. Use when the user needs web search, code examples, company intel, people lookup, or AI-powered deep research with Exa's neural search engine.
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
Connects AI agents to remote Windows desktop applications on Amazon WorkSpaces Applications (AppStream 2.0) through the managed Agent Access MCP server, and guides reliable desktop automation. Covers connecting an agent to the MCP endpoint (SigV4, streaming URL, and Active Directory SAML/Domain Join
Creates a reusable use case specification file that defines the business problem, stakeholders, and measurable success criteria for model customization, as recommended by the AWS Responsible AI Lens. Use as the default first step in any model customization plan. Skip only if the user explicitly decl