nemo-mbridge-perf-tp-dp-comm-overlap
Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 5
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 afa3a6854c507a05… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
TP / DP / PP Communication Overlap Skill
For stable background and recommendation level, see:
- @docs/training/communication-overlap.md
Enablement
Minimal Bridge override:
from megatron.bridge.training.comm_overlap import CommOverlapConfig
cfg.model.tensor_model_parallel_size = 4
cfg.model.sequence_parallel = True
cfg.model.pipeline_model_parallel_size = 4
cfg.model.virtual_pipeline_model_parallel_size = 2
cfg.comm_overlap = CommOverlapConfig(
tp_comm_overlap=True,
)
cfg.ddp.use_distributed_optimizer = True
cfg.ddp.overlap_grad_reduce = True
cfg.ddp.overlap_param_gather = True
Optional TP preset:
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
Precision knobs belong to mixed precision:
cfg.mixed_precision.grad_reduce_in_fp32 = False
cfg.mixed_precision.fp8_param_gather = False
Code Anchors
Bridge overlap gating:
if self.user_comm_overlap_cfg.tp_comm_overlap is True:
if model_cfg.tensor_model_parallel_size < 2:
...
elif not model_cfg.sequence_parallel:
...
elif not HAVE_TE:
...
PP overlap selection:
if model_cfg.pipeline_model_parallel_size > 1:
if vp_size > 1:
comm_overlap_cfg.overlap_p2p_comm = True
comm_overlap_cfg.batch_p2p_comm = False
else:
comm_overlap_cfg.overlap_p2p_comm = False
comm_overlap_cfg.batch_p2p_comm = True
DP overlap defaults:
if self.data_parallel_size > 1:
comm_overlap_cfg.bucket_size = 128 * 1024 * 1024
comm_overlap_cfg.overlap_grad_reduce = True
comm_overlap_cfg.overlap_param_gather = True
Launch-time env tuning:
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)
...
executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
Pitfalls
- TP overlap silently disables itself if
sequence_parallel=Falseor Transformer Engine is unavailable. - PP overlap is not enabled for all PP cases. Bridge only auto-selects
overlap_p2p_comm=TruewhenPP > 1andVPP > 1. bucket_sizeis a parameter-count knob, not a byte-size knob.grad_reduce_in_fp32andfp8_param_gathershould be set through mixed precision, not as standalone DDP tuning first.CUDA_DEVICE_MAX_CONNECTIONSand LayerNorm SM margin are launch-time plugin settings, notCommOverlapConfigfields.
Verification
Use the checked-in overlap unit coverage first:
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q
Optional second check if nemo_run is available:
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q
Success criteria:
- first command reports
26 passed - second command validates plugin-owned env wiring when not skipped
Files
5- BENCHMARK.md
68da81dc314.3 KB - SKILL.md
0cc72662743.6 KB - card.yaml
b46cb2ee442.1 KB - evals/evals.json
7d1e338c831.4 KB - skill-card.md
a6fd6e99913.7 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.