dynamo-troubleshoot
Diagnose failed or unhealthy Dynamo deployments. Use when pods, model-cache jobs, PVCs, workers, frontend/router health, endpoints, or benchmark jobs fail; use recipe-runner/router-starter before this for normal bring-up.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 6
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 aefaa210fab1b1cd… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Dynamo Troubleshoot
Purpose
Turn a Dynamo failure into a clear problem class, strongest signal, and next action. Start with read-only evidence, avoid secrets, and fix one layer at a time.
Prerequisites
- Python 3.10+ on the operator machine.
kubectlconfigured with read access to the target namespace.- Permission to read pods, events, jobs, PVCs, and
DynamoGraphDeploymentresources (NOT secrets). - Network reachability to the cluster API server.
Instructions
1. Collect A Read-Only Bundle
Run:
python3 scripts/collect_dynamo_debug_bundle.py \
--namespace "${NAMESPACE}"
If the user names a deployment, include it:
python3 scripts/collect_dynamo_debug_bundle.py \
--namespace "${NAMESPACE}" \
--deployment-name <deployment-name>
Do not collect Kubernetes secrets. Do not print Hugging Face tokens.
2. Classify The Failure
Use references/failure-decision-tree.md and classify into one primary bucket:
- cluster/platform
- namespace/secret
- model cache/PVC/download
- image pull/runtime image
- GPU scheduling/resources
- operator/DynamoGraphDeployment reconciliation
- frontend/router
- worker/backend
- endpoint/API
- benchmark/perf job
3. Debug Top Down
Check in this order:
- namespace, storage class, GPU nodes, and HF secret existence
- PVC and model-download job
DynamoGraphDeploymentstatus and events- pod status,
describe pod, and container logs - frontend service and port-forward
/v1/models/v1/chat/completions- benchmark job only after endpoint smoke test passes
4. Fix One Layer At A Time
Prefer the smallest reversible change:
- create missing namespace or HF secret
- patch
storageClassName - patch image tag or image pull secret
- reduce GPU request only if the recipe can still be valid
- switch KV router to approximate mode only if workers do not publish events
- restart failed jobs after fixing the underlying config
After each fix, rerun the relevant readiness check before moving deeper.
Available Scripts
| Script | Purpose | Arguments |
|---|---|---|
scripts/collect_dynamo_debug_bundle.py | Collect a read-only debug bundle (pods, events, jobs, PVCs, CR status) | --namespace, --deployment-name, --output-dir |
Invoke via the agentskills.io run_script() protocol:
run_script("scripts/collect_dynamo_debug_bundle.py", args=["--namespace", "dynamo-demo"])
Examples
Collect everything in a namespace for triage:
python3 scripts/collect_dynamo_debug_bundle.py --namespace dynamo-demo
Scope to a single failing deployment:
python3 scripts/collect_dynamo_debug_bundle.py \
--namespace dynamo-demo \
--deployment-name qwen-vllm-disagg
Equivalent through the agent protocol:
run_script("scripts/collect_dynamo_debug_bundle.py", args=["--namespace", "dynamo-demo", "--deployment-name", "qwen-vllm-disagg"])
Output Contract
Return:
- problem class
- evidence checked
- strongest signal
- likely cause
- exact next command or patch
- what was ruled out
- whether it is safe to continue deployment or benchmarking
Limitations
- Read-only. Never mutates the cluster; remediation commands are returned, not executed.
- Will not collect secrets or print Hugging Face tokens; some failure modes (auth) may need user-side inspection.
- Bundle size grows with deployment size; on very large namespaces, scope with
--deployment-name. - Does not validate disagg transport — use
dynamo-interconnect-checkfor that.
Troubleshooting
| Symptom | Likely cause | Next step |
|---|---|---|
kubectl returns Forbidden on events/pods | Service account lacks read RBAC | Ask operator for read-only role binding on the namespace |
Bundle missing DynamoGraphDeployment status | Operator not installed or different namespace | Verify dynamo-platform operator is installed and watching the namespace |
Model-download job in Pending | PVC unbound or HF secret missing | Fix PVC binding or create the named HF secret, then rerun the job |
Worker pods CrashLoopBackOff | Image/runtime mismatch or GPU not available | Inspect container logs; check nvidia.com/gpu allocatable on nodes |
Benchmark
See BENCHMARK.md for the NVCARPS-EVAL performance report (auto-generated by the NVSkills CI pipeline). To refresh, re-run /nvskills-ci on an upstream PR touching this skill.
References
- Read
references/failure-decision-tree.mdfor bucket-specific checks. - Use
scripts/collect_dynamo_debug_bundle.pyfor read-only bundle collection.
Files
6- BENCHMARK.md
a613e205d83.0 KB - SKILL.md
3f11c76e095.0 KB - evals/evals.json
492c2ac9263.2 KB - references/failure-decision-tree.md
264aa0a38f3.7 KB - scripts/collect_dynamo_debug_bundle.py
1200ebc7f18.3 KB - skill-card.md
10ba93d93d2.8 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related devops skillsscan passed
Post-deploy canary monitoring. (gstack)
Pre-deployment checks for router and switch configuration, including dangerous commands, duplicate addresses, subnet overlaps, stale references, management-plane risk, and IOS-style security hygiene. Use when reviewing a router or switch configuration before deployment.
Migrate Cloudflare Sandbox apps from stable @cloudflare/sandbox to @cloudflare/sandbox@next (SDK 1.0 preview). Use sandbox-next for apps already on the preview.
Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea
Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
Deploys and configures classic Firebase Hosting for static websites, single-page apps (SPAs), and microservices. Use when deploying static sites/SPAs, setting up custom domains, configuring firebase.json hosting settings (redirects, rewrites, headers, multi-site), or managing preview channels. Don't