tao-train-metric-learning-recognition
Metric-learning recognition (ml-recog) for fine-grained visual recognition. Learns embeddings for
- 0
- Installs
- —
- Rating
- —
- Success rate
- 22
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 04782d6b24aef389… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
ML Recog
Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Metric learning recognition for fine-grained visual recognition. Learns embeddings for retrieval-based matching (e.g., retail product recognition). Uses triplet/contrastive losses.
Set model.pretrained_model_path for pretrained backbone.
For TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and TensorRT inference), read references/tao-deploy-metric-learning-recognition.md first. Deploy spec templates live in this skill's references/ folder with the spec_template_deploy_*.yaml prefix.
Dataclass Schemas
Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML for an action requires schemas/<action>.schema.json and references/spec_template_<action>.yaml to exist and parse. Use the packaged selected-action schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.
Train Action Policy
This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.
Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.
Training Requirements
- Dataset type: ml_recog
- Formats: default
- Monitoring metric: val Precision at Rank 1
- AutoML metric contract: Use
val Precision at Rank 1emitted during training and maximize it. Usetest Precision at Rank 1only for standalone selected-checkpoint validation. - Standalone evaluation metric:
test Precision at Rank 1. Use the trainingval Precision at Rank 1KPI for AutoML recommendation ranking and thetestKPI only to verify the selected checkpoint on the evaluation reference/query split.
Per-Action Dataset Requirements
| Action | Spec Key | Source | Files | List? |
|---|---|---|---|---|
| evaluate | dataset.val_dataset | train_datasets | reference: metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/reference.tar.gz, query: metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/test.tar.gz | No |
| inference | dataset.val_dataset | train_datasets | reference: metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/reference.tar.gz, query: | No |
| inference | inference.input_path | train_datasets | metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/test.tar.gz | No |
| train | dataset.train_dataset | train_datasets | metric_learning_recognition/retail-product-checkout-dataset_classification_demo/known_classes/train.tar.gz | No |
| train | dataset.val_dataset | train_datasets | reference: metric_learning_recognition/retail-product-checkout-dataset_classification_demo/known_classes/reference.tar.gz, query: metric_learning_recognition/retail-product-checkout-dataset_classification_demo/known_classes/val.tar.gz | No |
Typical Spec Overrides
Data source overrides are mandatory for every action — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in spec_overrides.
S3_TRAIN = "s3://bucket/data/train"
train (mandatory data sources):
{
"train.num_epochs": 30,
"train.checkpoint_interval": 10,
"train.validation_interval": 10,
"train.num_gpus": 1,
"dataset.train_dataset": f"{S3_TRAIN}/metric_learning_recognition/retail-product-checkout-dataset_classification_demo/known_classes/train.tar.gz",
"dataset.val_dataset": {"reference": f"{S3_TRAIN}/metric_learning_recognition/retail-product-checkout-dataset_classification_demo/known_classes/reference.tar.gz", "query": f"{S3_TRAIN}/metric_learning_recognition/retail-product-checkout-dataset_classification_demo/known_classes/val.tar.gz"},
}
evaluate (mandatory data sources):
{
"evaluate.checkpoint": "<selected train/AutoML checkpoint>",
"dataset.val_dataset": {"reference": f"{S3_TRAIN}/metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/reference.tar.gz", "query": f"{S3_TRAIN}/metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/test.tar.gz"},
}
inference (mandatory data sources):
{
"inference.checkpoint": "<selected train/AutoML checkpoint>",
"dataset.val_dataset": {"reference": f"{S3_TRAIN}/metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/reference.tar.gz"},
"inference.input_path": f"{S3_TRAIN}/metric_learning_recognition/retail-product-checkout-dataset_classification_demo/unknown_classes/test.tar.gz",
}
Eval Dataset
Required. Evaluation requires reference and query datasets for retrieval metrics.
Important Parameters
- model.backbone: Default resnet_50. Options: resnet_50, resnet_101, fan_small, fan_base, fan_large, fan_tiny, nvdinov2_vit_large_legacy.
- model.feat_dim: Embedding dimension. Default 256. Output feature vector size for similarity matching.
- train.batch_size: Per-GPU batch size. Default 4.
val_batch_sizealso 4. For training and AutoML search,train.batch_sizemust be divisible bydataset.num_instance. - dataset.num_instance: Instances per identity in a batch (P/K sampling). Default 4. Controls how many images of the same class appear together. If using a custom AutoML range for
train.batch_size, use explicit options that are multiples of this value. - train.optim.trunk.base_lr: Learning rate for the trunk (backbone). Default 3.5e-4 (Adam).
- train.optim.embedder.base_lr: Learning rate for the embedding head. Default 3.5e-4.
- train.optim.triplet_loss_margin: Margin for triplet loss. Default 0.3. smooth_loss=True by default.
- train.optim.miner_function_margin: Hard mining margin. Default 0.1. Controls pair mining difficulty.
- train.optim.steps: LR decay steps. Default [40, 70] with gamma=0.1.
- dataset.train_dataset: Path to training images organized in class folders.
- dataset.val_dataset: Dict with 'reference' and 'query' keys pointing to ImageNet-format directories for retrieval evaluation.
Multi-GPU / Multi-Node
Launch method: Lightning-managed (single python process, Lightning spawns workers).
| Spec Key | Description | Default |
|---|---|---|
train.num_gpus | Number of GPUs | 1 |
train.gpu_ids | GPU device indices | [0] |
- Strategy:
auto(Lightning picks best strategy automatically) - No explicit
num_nodesordistributed_strategyconfig — single-node oriented
Hardware
Minimum 1 GPU(s), recommended 2 GPU(s). 16GB+ VRAM per GPU. Metric learning benefits from larger batch sizes for better triplet sampling but is otherwise moderate on memory.
Error Patterns
Reference/query mismatch: Ensure reference and query datasets share compatible class namespaces for evaluation.
PyTorch 2.6 checkpoint load failure on checkpoint actions: Current TAO
ML-Recog checkpoints may contain OmegaConf objects. For checkpoints produced by
the same trusted TAO train/AutoML workflow, set
TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1 in downstream evaluate, inference, export,
or resume/retrain job env vars so Lightning can load the full checkpoint. Do not
use this env var for untrusted checkpoints.
Spec Param / Parent Model Inference
Model-specific inference mappings belong in this MD file, not in config.json. Generated runners should read this section and apply the mappings with SDK helpers before create_job(). This mirrors the old microservices infer_params.py flow.
Inference mappings from TAO Core ml_recog.config.json:
| Action | Spec Field | Inference Function | Meaning |
|---|---|---|---|
| evaluate | evaluate.checkpoint | parent_model | model file inferred from the parent job results folder |
| evaluate | results_dir | output_dir | current job results directory |
| export | export.checkpoint | parent_model | model file inferred from the parent job results folder |
| export | export.onnx_file | create_onnx_file | output ONNX path |
| export | results_dir | output_dir | current job results directory |
| inference | inference.checkpoint | parent_model | model file inferred from the parent job results folder |
| inference | results_dir | output_dir | current job results directory |
| train | model.pretrained_model_path | ptm_if_no_resume_model | PTM when no resume checkpoint exists |
| train | results_dir | output_dir | current job results directory |
| train | train.resume_training_checkpoint_path | resume_model | model file inferred from the current job results folder |
For parent_model or parent_model_folder, pass the upstream train/export/AutoML child job id as parent_job_id. The SDK lists the parent result folder, filters checkpoint artifacts, and returns the selected model file or folder. Do not add these mappings back to config.json and do not patch generated runner scripts to guess checkpoint paths.
Deployment
Files
22- BENCHMARK.md
ef827a042f7.1 KB - SKILL.md
8d516bbf9311.3 KB - config/skillspector-baseline.yaml
8b379b90c3834 B - evals/evals.json
f1a9fbf49f887 B - references/skill_info.yaml
a69317bb5c4.8 KB - references/spec_template_deploy_evaluate.yaml
d48e7d8fb7169 B - references/spec_template_deploy_gen_trt_engine.yaml
5a5156489c582 B - references/spec_template_deploy_inference.yaml
a91f49dc2e308 B - references/spec_template_evaluate.yaml
6069c1ee461.8 KB - references/spec_template_export.yaml
f1b74b61261.7 KB - references/spec_template_gen_trt_engine.yaml
6acaded85f2.0 KB - references/spec_template_inference.yaml
cf2ed3ad1c1.8 KB - references/spec_template_train.yaml
41fde60c901.6 KB - references/tao-deploy-metric-learning-recognition.md
79646eaf115.9 KB - references/tao-deploy-metric-learning-recognition.skill_info.yaml
9f687245574.0 KB - schemas/evaluate.schema.json
41ff49ab1635.6 KB - schemas/export.schema.json
521f5f1abd34.6 KB - schemas/gen_trt_engine.schema.json
481b28190d40.9 KB - schemas/inference.schema.json
9aa68ba46935.9 KB - schemas/manifest.json
ad232dba4c10.3 KB - schemas/train.schema.json
20fd568d3f32.5 KB - skill-card.md
c6b287134b4.1 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related ai-ml skillsscan passed
Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest. Use after ito-compute has booked GPU nodes and the user wants pre-training, fine-tuning, or RL on that metal. ECC implemen
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). Use when the user says "is my dataset okay", "evaluate my data", "check my training data", "I have my own data", or before starting any fine-tuning job. Detects file format, checks schema compliance against