tao-train-mask-auto-label
MAL (Mask Auto-Label) for weakly-supervised segmentation. Produces segmentation masks from minimal annotations
- 0
- Installs
- —
- Rating
- —
- Success rate
- 13
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 5162c72b6ac0efc3… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
MAL
Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
MAL (Mask Auto-Label) for weakly-supervised segmentation. Produces segmentation masks from minimal annotations (e.g., point or box annotations). Uses ViT-MAE backbone.
Set train.pretrained_model_path for ViT-MAE pretrained weights.
Quick Start (docker run)
Docker-native launch — no TAO SDK and no Python on the host. Use the local Docker/platform skill instead when it gives a stricter environment-specific command (non-root UID mapping, cache redirects, remote daemons).
TAO_PYT_IMAGE_DEFAULT=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt # versions-key: images.tao_toolkit.pyt
TAO_PYT_IMAGE="${TAO_PYT_IMAGE:-$TAO_PYT_IMAGE_DEFAULT}"
RUN_ROOT="${RUN_ROOT:-$PWD}"
DOCKER_COMMON=(
--rm --gpus all --shm-size=8g
--shm-size=8g
--ulimit memlock=-1
--ulimit stack=67108864
-v "$RUN_ROOT/data:/data:ro"
-v "$RUN_ROOT/specs:/specs:ro"
-v "$RUN_ROOT/results:/results"
)
Train:
docker run "${DOCKER_COMMON[@]}" "$TAO_PYT_IMAGE" \
mal train -e /specs/train.yaml
Evaluate:
docker run "${DOCKER_COMMON[@]}" "$TAO_PYT_IMAGE" \
mal evaluate -e /specs/evaluate.yaml
Inference:
docker run "${DOCKER_COMMON[@]}" "$TAO_PYT_IMAGE" \
mal inference -e /specs/inference.yaml
Every action takes its spec with -e; results_dir is set in the spec or
overridden on the command line. Mount any pretrained-weights directory the spec
references, and keep every in-container path consistent across actions.
Dataclass Schemas
Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML for an action requires schemas/<action>.schema.json and references/spec_template_<action>.yaml to exist and parse. Use the packaged selected-action schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.
Train Action Policy
This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.
Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.
Training Requirements
- Dataset type: segmentation
- Formats: default
- Monitoring metric: mIoU
- AutoML metric contract: Use
mIoUemitted by the evaluate action and maximize it. Compare the recorded AutoML objective with the evaluator's emittedmIoU; do not substitute training loss.
Per-Action Dataset Requirements
| Action | Spec Key | Source | Files | List? |
|---|---|---|---|---|
| evaluate | dataset.val_img_dir | eval_dataset | images.tar.gz | No |
| evaluate | dataset.val_ann_path | eval_dataset | annotations.json | No |
| inference | inference.img_dir | inference_dataset | images.tar.gz | No |
| inference | inference.ann_path | inference_dataset | annotations.json | No |
| train | dataset.train_img_dir | train_datasets | images.tar.gz | No |
| train | dataset.train_ann_path | train_datasets | annotations.json | No |
| train | dataset.val_img_dir | eval_dataset | images.tar.gz | No |
| train | dataset.val_ann_path | eval_dataset | annotations.json | No |
Typical Spec Overrides
Data source overrides are mandatory for every action — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in spec_overrides.
MAL expects COCO-style annotation JSON plus image paths that match the JSON
file_name entries after the data source is prepared. Archive-only CSV/image
datasets are not compatible unless they are converted to this format first.
S3_TRAIN = "s3://bucket/data/train"
S3_EVAL = "s3://bucket/data/eval"
train (mandatory data sources):
{
"train.num_gpus": 1,
"train.gpu_ids": [
0
],
"train.num_epochs": 5,
"train.checkpoint_interval": 5,
"train.validation_interval": 5,
"dataset.train_img_dir": f"{S3_TRAIN}/images.tar.gz",
"dataset.train_ann_path": f"{S3_TRAIN}/annotations.json",
"dataset.val_img_dir": f"{S3_EVAL}/images.tar.gz",
"dataset.val_ann_path": f"{S3_EVAL}/annotations.json",
}
evaluate (mandatory data sources):
{
"evaluate.checkpoint": "<selected train/AutoML checkpoint>",
"dataset.val_img_dir": f"{S3_EVAL}/images.tar.gz",
"dataset.val_ann_path": f"{S3_EVAL}/annotations.json",
}
inference (mandatory data sources):
{
"inference.checkpoint": "<selected train/AutoML checkpoint>",
"inference.img_dir": f"{S3_EVAL}/images.tar.gz",
"inference.ann_path": f"{S3_EVAL}/annotations.json",
}
For checkpoint-dependent actions, use the model resolver declared in
references/skill_info.yaml. Select the exact epoch/step checkpoint requested
by the user or the best checkpoint when a best-checkpoint action is requested.
The mal_model_latest.pth symlink is only appropriate when the user explicitly
asks for the latest checkpoint.
Eval Dataset
Optional. Val images and annotations configured alongside train paths.
Important Parameters
- model.arch: ViT-MAE backbone variant. Default vit-mae-base/16.
Avoid
vit-deit-tiny/16; the current runtime rejects tiny ViT variants. - train.lr: Learning rate. Default 1e-6 (very low — fine-tuning ViT).
- dataset.crop_size: Training crop size. Default 512. Use this key, not
model.crop_size. - train.warmup_epochs: Warmup epochs before full learning rate.
- dataset.load_mask: Whether annotations contain pre-computed segmentation
masks. Set this to
falsefor COCO annotations that contain boxes but nosegmentationfields when performing training or inference without mask ground truth. Keep ittruewhen every annotation contains segmentation data. AutoML validation that selects bymIoUrequires segmentation ground truth anddataset.load_mask: true; bbox-only evaluation emits non-finitemIoUand must not be accepted as a valid AutoML objective.
AutoML / HPO Notes
For MAL AutoML launches, keep the default smoke search space narrow and pass
automl_hyperparameters=["train.lr", "train.wd"]. Use conservative Bayesian
ranges around the ViT-MAE fine-tuning defaults, for example
train.lr from 1e-7 to 1e-5 and train.wd from 1e-5 to 1e-2.
The packaged train schema marks these two parameters as the default AutoML
parameters; pass them explicitly when using a runtime that still derives MAL
search metadata from its bundled config module.
Multi-GPU / Multi-Node
Launch method: Lightning-managed (single python process, Lightning spawns workers).
| Spec Key | Description | Default |
|---|---|---|
train.num_gpus | Number of GPUs | 1 |
train.gpu_ids | GPU device indices | [0] |
train.num_nodes | Number of nodes | 1 |
- Multi-GPU strategy:
ddp_find_unused_parameters_true - No fsdp support
- LR auto-scaling:
lr = lr * num_devices * batch_size(learning rate is scaled automatically by device count and batch size)
Multi-node env vars (set by orchestrator): WORLD_SIZE, NODE_RANK, MASTER_ADDR, MASTER_PORT, NUM_GPU_PER_NODE.
Hardware
Minimum 1 GPU(s), recommended 2 GPU(s). 24GB+ (A100 recommended) VRAM per GPU. ViT-MAE backbone at crop_size=512 needs 24GB+ GPU memory.
Error Patterns
CUDA out of memory: Reduce dataset.crop_size (512 -> 384 -> 256) or use a smaller ViT-MAE variant (base vs large).
Key crop_size not in MALModelConfig: The crop-size override was placed
under model.crop_size. Move it to dataset.crop_size.
Spec Param / Parent Model Inference
Model-specific inference mappings belong in this MD file, not in config.json. Generated runners should read this section and apply the mappings with SDK helpers before create_job(). This mirrors the old microservices infer_params.py flow.
Inference mappings from TAO Core mal.config.json:
| Action | Spec Field | Inference Function | Meaning |
|---|---|---|---|
| evaluate | evaluate.checkpoint | parent_model | model file inferred from the parent job results folder |
| evaluate | results_dir | output_dir | current job results directory |
| inference | inference.checkpoint | parent_model | model file inferred from the parent job results folder |
| inference | inference.label_dump_path | create_inference_result_file_mal | MAL inference JSON path |
| inference | results_dir | output_dir | current job results directory |
| train | train.pretrained_model_path | ptm_if_no_resume_model | optional pretrained model when not resuming |
| train | train.resume_training_checkpoint_path | resume_model | exact checkpoint for resume runs |
| train | results_dir | output_dir | current job results directory |
For parent_model or parent_model_folder, pass the upstream train/export/AutoML child job id as parent_job_id. The SDK lists the parent result folder, filters checkpoint artifacts, and returns the selected model file or folder. Do not add these mappings back to config.json and do not patch generated runner scripts to guess checkpoint paths.
Files
13- BENCHMARK.md
0de55874746.9 KB - SKILL.md
5784de209911.1 KB - config/skillspector-baseline.yaml
11f48bf7871.0 KB - evals/evals.json
442b070811821 B - references/skill_info.yaml
618d07e6993.4 KB - references/spec_template_evaluate.yaml
dfcd1022ba1.8 KB - references/spec_template_inference.yaml
0f84115c3b1.8 KB - references/spec_template_train.yaml
c5e0bb5d561.5 KB - schemas/evaluate.schema.json
46185b04e520.2 KB - schemas/inference.schema.json
6c09077c6a20.1 KB - schemas/manifest.json
a9e80216c33.7 KB - schemas/train.schema.json
5a4e928c1d17.3 KB - skill-card.md
3921242f754.3 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related knowledge skillsscan passed
Restrict file edits to a specific directory for the session. (gstack)
PostHog logs for Java
Operate execution flow across GitHub and Linear by triaging issues and pull requests, linking active work, and keeping GitHub public-facing while Linear remains the internal execution layer. Use when the user wants backlog control, PR triage, or GitHub-to-Linear coordination.