nemo-mbridge-perf-cuda-graphs
Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 5
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 0b71dfa17ec730dd… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
CUDA Graphs
Stable documentation: @docs/training/cuda-graphs.md Card: @skills/nemo-mbridge-perf-cuda-graphs/card.yaml
What It Is
CUDA graphs capture GPU operations once and replay them with minimal host-driver overhead. Bridge supports two implementations:
cuda_graph_impl | Mechanism | Scope support |
|---|---|---|
"local" | MCore FullCudaGraphWrapper wrapping entire fwd+bwd | full_iteration |
"transformer_engine" | TE make_graphed_callables() per layer | attn, mlp, moe, moe_router, moe_preprocess, mamba |
Quick Decision
Start with TE-scoped graphs for most training workloads, then verify replay timing against eager on the same dispatcher, layout, and container:
- dense models:
attn, then optionallymlp - dropless MoE:
attn moe_router moe_preprocess - VLMs: the same dropless-MoE scope, but only after the real-data path is stable
Use local + full_iteration only when you specifically want full-iteration
capture and can satisfy the tighter constraints.
For recompute-heavy workloads:
- TE-scoped graphs pair naturally with selective recompute
- full recompute usually pushes you toward
localfull-iteration graphs or away from graphs entirely
Related docs:
- @docs/training/cuda-graphs.md
- @docs/training/activation-recomputation.md
Enablement
Local full-iteration graph
cfg.model.cuda_graph_impl = "local"
cfg.model.cuda_graph_scope = ["full_iteration"]
cfg.model.cuda_graph_warmup_steps = 3
cfg.model.use_te_rng_tracker = True
cfg.rng.te_rng_tracker = True
cfg.rerun_state_machine.check_for_nan_in_loss = False
cfg.ddp.check_for_nan_in_grad = False
TE scoped graph (dense model)
cfg.model.cuda_graph_impl = "transformer_engine"
cfg.model.cuda_graph_scope = ["attn"] # or ["attn", "mlp"]
cfg.model.cuda_graph_warmup_steps = 3
cfg.model.use_te_rng_tracker = True
cfg.rng.te_rng_tracker = True
TE scoped graph (MoE model)
cfg.model.cuda_graph_impl = "transformer_engine"
cfg.model.cuda_graph_scope = ["attn", "moe_router", "moe_preprocess"]
cfg.model.cuda_graph_warmup_steps = 3
cfg.model.use_te_rng_tracker = True
cfg.rng.te_rng_tracker = True
Performance harness CLI
uv run python scripts/performance/run_script.py \
-m qwen \
-mr qwen3_30b_a3b \
--task pretrain \
-g h100 \
-c bf16 \
-ng 16 \
--cuda_graph_impl transformer_engine \
--cuda_graph_scope attn,moe_router,moe_preprocess \
...
Valid CLI values live in scripts/performance/argument_parser.py:
VALID_CUDA_GRAPH_IMPLS:["none", "local", "transformer_engine"]VALID_CUDA_GRAPH_SCOPES:["full_iteration", "attn", "mlp", "moe", "moe_router", "moe_preprocess", "mamba"]
The performance harness uses a comma-separated --cuda_graph_scope value and
auto-enables model.use_te_rng_tracker plus rng.te_rng_tracker when
--cuda_graph_impl is not none.
Required constraints
use_te_rng_tracker = True(enforced ingpt_provider.py)full_iterationscope only withcuda_graph_impl = "local"full_iterationscope requirescheck_for_nan_in_loss = False- Do not combine
moescope andmoe_routerscope - Tensor shapes must be static (fixed seq_length, fixed micro_batch_size)
- MoE token-dropless routing limits graphable scope to dense modules
- With
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, setNCCL_GRAPH_REGISTER=0(MCore enforces for local impl on arch < sm_100; TE impl asserts unconditionally) - CPU offloading is incompatible with CUDA graphs
moe_preprocessscope requiresmoe_routerscope to also be set
Practical bring-up order
- Stabilize the eager run first.
- Fix sequence length and micro-batch size.
- Enable the narrowest useful graph scope.
- Confirm replay is active and memory is still acceptable.
- Compare eager against graph replay iterations after warmup and capture; do not include the capture step in steady-state timing.
- Only then widen scope or combine with overlap features.
Code Anchors
Bridge config and validation
# CUDA graph scope validation: check_for_nan_in_loss must be disabled with full_iteration graph
if self.model.cuda_graph_impl == "local" and CudaGraphScope.full_iteration in self.model.cuda_graph_scope:
assert not self.rerun_state_machine.check_for_nan_in_loss, (
"check_for_nan_in_loss must be disabled when using full_iteration CUDA graph. "
"Set rerun_state_machine.check_for_nan_in_loss=False."
)
if self.model.cuda_graph_impl == "none":
self.model.cuda_graph_scope = []
TE RNG tracker requirement
if self.cuda_graph_impl != "none":
assert getattr(self, "use_te_rng_tracker", False), (
"Transformer engine's RNG tracker is required for cudagraphs, it can be "
"enabled with use_te_rng_tracker=True'."
Graph creation and capture in training loop
# Capture CUDA Graphs.
cuda_graph_helper = None
if model_config.cuda_graph_impl == "transformer_engine":
cuda_graph_helper = TECudaGraphHelper(...)
# ...
if config.model.cuda_graph_impl == "local" and CudaGraphScope.full_iteration in config.model.cuda_graph_scope:
forward_backward_func = FullCudaGraphWrapper(
forward_backward_func, cuda_graph_warmup_steps=config.model.cuda_graph_warmup_steps
)
TE graph capture after warmup
# Capture CUDA Graphs after warmup.
if (
model_config.cuda_graph_impl == "transformer_engine"
and cuda_graph_helper is not None
and not cuda_graph_helper.graphs_created()
and global_state.train_state.step - start_iteration == model_config.cuda_graph_warmup_steps
):
if model_config.cuda_graph_warmup_steps > 0 and should_toggle_forward_pre_hook:
disable_forward_pre_hook(model, param_sync=False)
cuda_graph_helper.create_cudagraphs()
if model_config.cuda_graph_warmup_steps > 0 and should_toggle_forward_pre_hook:
enable_forward_pre_hook(model)
cuda_graph_helper.cuda_graph_set_manual_hooks()
RNG initialization
_set_random_seed(
rng_config.seed,
rng_config.data_parallel_random_init,
rng_config.te_rng_tracker,
rng_config.inference_rng_tracker,
use_cudagraphable_rng=(model_config.cuda_graph_impl != "none"),
pg_collection=pg_collection,
)
Delayed wgrad + CUDA graph interaction
cuda_graph_scope = getattr(model_cfg, "cuda_graph_scope", []) or []
# ... scope parsing ...
if wgrad_in_graph_scope:
assert is_te_min_version("2.12.0"), ...
assert model_cfg.gradient_accumulation_fusion, ...
if attn_scope_enabled:
assert not model_cfg.add_bias_linear and not model_cfg.add_qkv_bias, ...
Perf harness override helper
def _set_cuda_graph_overrides(
recipe, cuda_graph_impl=None, cuda_graph_scope=None
):
# Sets impl, scope, and auto-enables te_rng_tracker
Graph cleanup
def _delete_cuda_graphs(cuda_graph_helper):
# Deletes FullCudaGraphWrapper and TE graph objects to free NCCL buffers
MCore classes (in 3rdparty/Megatron-LM)
CudaGraphManager:megatron/core/transformer/cuda_graphs.pyTECudaGraphHelper:megatron/core/transformer/cuda_graphs.pyFullCudaGraphWrapper:megatron/core/full_cuda_graph.pyCudaGraphScopeenum:megatron/core/transformer/enums.py
Positive recipe anchors
src/megatron/bridge/perf_recipes/deepseek/gb300/deepseek_v3.pysrc/megatron/bridge/perf_recipes/qwen/gb300/qwen3_moe.pysrc/megatron/bridge/perf_recipes/gpt_oss/gb300/gpt_oss.py
Tests
| File | Coverage |
|---|---|
tests/unit_tests/training/test_config.py | full_iteration NaN-check constraint |
tests/unit_tests/training/test_comm_overlap.py | delay_wgrad + CUDA graph interaction |
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py | TE autocast with CUDA graphs |
tests/functional_tests/test_groups/recipes/test_llama_recipes_pretrain_cuda_graphs.py | End-to-end local and TE graph smoke tests |
tests/unit_tests/recipes/kimi/test_kimi_k2.py | TE + CUDA graph recipe config |
tests/unit_tests/recipes/gpt/test_gpt3_175b.py | TE + CUDA graph recipe config |
tests/unit_tests/recipes/qwen_vl/test_qwen25_vl_recipes.py | VLM CUDA graph settings |
Pitfalls
-
TE RNG tracker is mandatory: Setting
cuda_graph_implwithoutuse_te_rng_tracker=Trueandrng.te_rng_tracker=Truewill assert in the provider. -
full_iterationrequires NaN checks disabled: The entire fwd+bwd is captured, so loss-NaN checking cannot inspect intermediate values. -
MoE scope restrictions:
moescope andmoe_routerscope are mutually exclusive. Token-dropless MoE can only graphmoe_routerandmoe_preprocess, not the full expert dispatch. -
Memory overhead: CUDA graphs pin all intermediate buffers for the graph's lifetime (no memory reuse). TE scoped graphs add a few GB; full-iteration graphs can increase peak memory by 1.5–2×.
PP > 1compounds overhead since each stage holds its own graph. -
Delayed wgrad interaction: When
delay_wgrad_compute=Trueand attention or MoE router is incuda_graph_scope, additional constraints apply: TE >= 2.12.0,gradient_accumulation_fusion=True, and no attention bias. -
Variable-length sequences break graphs: Sequence lengths must be constant across steps. Use padded packed sequences if packing is needed.
-
Graph cleanup is required: CUDA graph objects hold NCCL buffer references. Bridge handles this in
_delete_cuda_graphs()at the end of training, but early exits must call it explicitly. -
Older GPU architectures: On GPUs with compute capability < 10.0 (pre-Blackwell), set
NCCL_GRAPH_REGISTER=0when usingPYTORCH_CUDA_ALLOC_CONF=expandable_segments:True. Enforced in MCoreCudaGraphManager(cuda_graphs.py:1428) andTECudaGraphHelper(cuda_graphs.py:1697). The TE impl asserts unconditionally regardless of arch. -
CPU offloading incompatible: CUDA graphs cannot be used with CPU offloading. Enforced in MCore
transformer_config.py:1907. -
MoE recompute + moe_router scope: MoE recompute is not supported with
moe_routerCUDA graph scope when usingcuda_graph_impl = "transformer_engine". Enforced in MCoretransformer_config.py:1977. -
Layer-level recompute requires
full_iterationscope: Usingrecompute_granularity="full"withrecompute_num_layers(recompute N whole transformer layers) is incompatible with TE-scoped graphs. MCore calls this "full" granularity even though you're selecting how many layers — the name refers to recomputing the full layer, not full model. Any TE-scoped scope (attn,mlp,moe_router, etc.) will assert:AssertionError: full recompute is only supported with full iteration CUDA graph.This commonly hits FP8 configs that default to TE-scoped graphs (e.g.LLAMA3_70B_SFT_CONFIG_H100_FP8_CS_V1usescuda_graph_impl= "transformer_engine",cuda_graph_scope="mlp"). Fix: use submodule recompute (recompute_granularity="selective"+recompute_modules), disable CUDA graphs, or switch tolocal+full_iteration. Enforced in MCoretransformer_config.py:2001-2005. See also @skills/nemo-mbridge-perf-activation-recompute/SKILL.md. -
Benchmark numbers are workload-specific: graph wins are usually real when host overhead is visible, but the exact gain depends on batch shape, PP depth, recompute, dispatcher backend, and whether the eager baseline was already optimized.
-
A successful capture is not a speedup guarantee: On 2026-05-18, Qwen3 30B A3B H100 BF16 pretrain with the all-to-all dispatcher captured TE-scoped
attn,moe_router,moe_preprocessgraphs successfully (48graphable layers, about6.9 scapture time on rank 0), but replay iterations 5-8 averaged42.00 sversus41.36 sfor eager. Treat scoped graphs as a bring-up candidate and validate on the target stack.
Verification
Unit tests
uv run python -m pytest \
tests/unit_tests/training/test_config.py -k "cuda_graph" \
tests/unit_tests/training/test_comm_overlap.py -k "cuda_graph" \
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cuda_graph" -q
Functional smoke test (requires GPU)
uv run python -m pytest \
tests/functional_tests/test_groups/recipes/test_llama_recipes_pretrain_cuda_graphs.py -q
Success criteria
- Unit tests pass, covering config validation for both
localandtransformer_engineimplementations. - Functional test completes training steps with both CUDA graph implementations.
- No NCCL errors or illegal memory access in logs.
Files
5- BENCHMARK.md
9709f2e8623.9 KB - SKILL.md
24c276f48113.8 KB - card.yaml
729e6cfe6e13.5 KB - evals/evals.json
5db93d51191.4 KB - skill-card.md
5121ebcada4.1 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related ai-ml skillsscan passed
Engineering operating model for teams where AI agents generate a large share of implementation output. Use when setting team process, review gates, or ownership rules for a codebase largely written by agents.
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v
Manages project directory setup and artifact organization. Use when starting a new project, resuming an existing one, or when a PLAN.md needs to be associated with a project directory. Creates the project folder structure (specs/, scripts/, notebooks/, manifests/, agent_memory/) and resolves project