gke-ai-troubleshooting-tpu-vbar-oom
Diagnoses and prevents vbar_control_agent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or high-frequency metrics polling. Use when troubleshooting vbar_control_agent crashes, memory cgroup OOM
- 0
- Installs
- —
- Rating
- —
- Success rate
- 3
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 70e96b3359076102… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
TPU Connection Failure and VBAR OOM Troubleshooting
Use this skill to systematically diagnose and prevent vbar_control_agent
segfaults and Out-Of-Memory (OOM) errors on TPU v6e nodes.
⚠️ Prerequisites
- Cloud Logging must be enabled for the project.
- Access to the project and cluster via
gcloudor equivalent tool.
🔍 Diagnostic Workflow
Step 0: Context Acquisition & Time Window Definition
Independently gather required context using available GCP/GKE tools or use the
provided {variable} placeholders:
{project_id}: The GCP Project ID (e.g.,customer-ai-project-123).{cluster_name}: The GKE Cluster Name (e.g.,tpu-cluster-prod).{node_name}: The Node Name or Instance ID (e.g.,tpu-node-1).{workload_name}: The Workload Name / JobSet Name (e.g.,my-training-job-456).{namespace}: The Workload Namespace.{issue_time}: The timestamp of the issue (e.g.,2026-04-14T20:00:00Z).
Time Handling & Execution Rules
- Window Calculation: If an issue timestamp
{issue_time}is provided, calculate the query time window as[{issue_time} - 30m]to[{issue_time} + 30m].- Let
{start_time}={issue_time} - 30m - Let
{end_time}={issue_time} + 30m
- Let
- Informational vs. Live Execution: If the user request is informational or query-formulation (e.g. "How can I check...", "How do I determine..."), or if live GCP project resources are not actively targetable, directly output the calculated time window, log names, and Cloud Logging filter templates without attempting live log execution commands.
Step 1: Check for vbar_control_agent OOMs
Look for specific out of memory messages from vbar_control_agent in serial
console logs (serialconsole.googleapis.com%2fserial_port_1_output).
- Tool to use:
query_logs(for live diagnostics) - Filter Templates:
Serial Console Logs (OOMs):
logName="projects/{project_id}/logs/serialconsole.googleapis.com%2fserial_port_1_output"
AND labels."compute.googleapis.com/resource_name"="{node_name}"
AND SEARCH(text_payload, "Memory cgroup out of memory: Killed process .* (vbar_control_ag)")
AND timestamp >= "{start_time}"
AND timestamp <= "{end_time}"
- Logic: Presence of
Memory cgroup out of memorymessages related tovbar_control_agent. Stack traces pointing tolibtpu::tpunetd::VBARControlHelper::MetricsReadFromVBARare a strong indicator. - Automation: Proceed to next step automatically after reporting findings.
- Reference: See
references/failure_signatures.mdfor example log patterns.
Step 2: Investigate tpu-device-plugin Metrics Fetch Failures [Low Risk]
Check if tpu-device-plugin is reporting metric fetch failures.
- Tool to use:
query_logs - Filter Template:
resource.type="k8s_container"
AND resource.labels.project_id="{project_id}"
AND resource.labels.cluster_name="{cluster_name}"
AND resource.labels.container_name="tpu-device-plugin"
AND severity=ERROR
AND textPayload:"metrics fetch failed for .* deviceID and .* device path with error: checksum didn't match with the metrics data. Corrupt data found"
AND timestamp >= "{start_time}"
AND timestamp <= "{end_time}"
- Logic: Errors indicating "metrics fetch failed" with "checksum didn't match" suggest vBAR memory corruption.
- Automation: Proceed to next step automatically after reporting findings.
Step 3: Check for Custom Metrics Collection Usage [Low Risk]
Inspect cluster configurations, workloads, or container specs to determine if custom TPU metrics collection mechanisms are deployed.
-
Action: Check if custom scripts or agents (e.g., using
libtpu.sdk.tpumonitoring) are deployed that frequently queryGetHostMetricsfromvBAR Control Agent. -
Verification Commands:
- Kubectl Search (Inspect workload env/specs):
kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}{"/"}{.metadata.name}{"\t"}{.spec.containers[*].image}{"\n"}{end}'- Log Search Filter (
query_logs):
resource.type="k8s_container" AND resource.labels.project_id="{project_id}" AND resource.labels.cluster_name="{cluster_name}" AND textPayload:"libtpu.sdk.tpumonitoring" AND timestamp >= "{start_time}" AND timestamp <= "{end_time}" -
Logic: Confirmation of custom metrics collection helps confirm the race condition hypothesis.
🛠️ Resolution Workflow
Resolution 1: Temporarily Disable Custom Metrics Collection [High Risk]
If a custom metrics collection agent is identified, recommend disabling it.
- Action: Recommend disabling the custom metrics collector.
- Justification: Prevents reads from vBAR during device resets, stopping crashes and OOMs.
Resolution 2: Await vbar_control_agent Resiliency Update [Low Risk]
Advise that a permanent fix will be available in a future GKE version.
- Action: Recommend upgrading GKE when the fix is available.
- Justification: The updated agent will be resilient to memory corruption and gracefully handle reads from unbound vBARs.
📋 copypaste checklist
- Acquire context and compute
[{start_time}, {end_time}]window. - Check for
vbar_control_agentsegfaults and OOMs usingquery_logs. - Investigate
tpu-device-pluginfailures usingquery_logs. - Inspect for custom metrics collection usage.
- Advise disabling custom metrics collection if applicable.
- Advise awaiting resiliency update.
Files
3- SKILL.md
82131eb1a56.2 KB - references/failure_signatures.md
aad7247e381.4 KB - scripts/validate_queries.sh
599c30f72e993 B
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related ai-ml skillsscan passed
Neural search via Exa MCP for web, code, and company research. Use when the user needs web search, code examples, company intel, people lookup, or AI-powered deep research with Exa's neural search engine.
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
Connects AI agents to remote Windows desktop applications on Amazon WorkSpaces Applications (AppStream 2.0) through the managed Agent Access MCP server, and guides reliable desktop automation. Covers connecting an agent to the MCP endpoint (SigV4, streaming URL, and Active Directory SAML/Domain Join
Creates a reusable use case specification file that defines the business problem, stakeholders, and measurable success criteria for model customization, as recommended by the AWS Responsible AI Lens. Use as the default first step in any model customization plan. Skip only if the user explicitly decl