tao-artifacts
The contract home for TAO's SDK-free execution pipeline — authoritative JSON Schemas for the four typed artifacts (spec-bundle, job-record, results_dir layout, best_rec) plus the fixed job-status vocabulary and the nested-not-dotted spec rule. Use when authoring or validating a spec-bundle before su
- 0
- Installs
- —
- Rating
- —
- Success rate
- 10
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 d34de29856224d90… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
tao-artifacts
Four typed artifacts flow through every TAO job. Their schemas live here and
nowhere else — producers (model/data skills) and consumers (platform skills)
both validate against this skill's references/.
| Artifact | Schema | Produced by → consumed by |
|---|---|---|
| spec-bundle | references/spec_bundle.schema.json | model/data skill → platform skill (at the submit seam) |
| job-record | references/job_record.schema.json | scripts/tao_job_record.py (the ONLY writer) → any re-attaching agent/poller |
| results_dir layout | references/results_dir.contract.md | platform skill at submit → whoever collects outputs |
| best_rec | references/best_rec.schema.json | tao-run-automl adapter → DEFT warm-start |
Quick Start — validate an artifact
python - <<'PY'
import json, yaml, jsonschema, pathlib
ref = pathlib.Path("${TAO_SKILL_BANK_PATH:?}/skills/core/tao-artifacts/references")
schema = json.loads((ref / "spec_bundle.schema.json").read_text())
bundle = yaml.safe_load(open("/path/to/bundle.yaml")) # or a dict built in-context
jsonschema.validate(bundle, schema) # raises on violation
print("bundle OK")
PY
Validate the bundle before the verify-before-launch gate; validate a job-record only when debugging (the writer script already enforces the schema).
The two rules the schemas enforce structurally
- Nested, not dotted. A
specis a nested dict mirroring the container's config shape —{"train": {"num_epochs": 12}}. Any key containing.at any depth is rejected ({"train.num_epochs": 12}is the #1 authoring mistake). Note the distinction:declared_inputs[].spec_keyandgpu_spec_keyare dotted/indexed pointers into the spec (dataset.train_data_sources[0].image_dir) — dots are correct there. - Mode discrimination.
mode: configrequiresspec+config_formatand acommandcontaining{config_path}, and forbidsargs.mode: argsrequiresargsand forbidsspec. There is no other mode.
Optional action lifecycle
Use execution when an action needs more than its primary command. This is the
shared model-to-platform seam; do not add a model-specific Docker, Kubernetes,
or SLURM renderer merely to carry runtime environment, attestations,
post-processing, or helper dependencies.
- The producing model/data skill owns
environment, orderedpre_commands, orderedpost_commands, distributed-launch intent, and completion evidence. - The platform owns container mounts, scheduler/container syntax, task/rank binding, timeouts, log paths, and preservation of the real child exit code.
environmentis non-secret. Credential values continue to use the selected platform's secret/sidecar contract and never enter a spec-bundle.- Commands, environment values, and string values in
specmay use{config_path},{job_id}, and{results_dir}. The platform binds them only after the job record has been opened; the job record'sresults_diris authoritative over any pre-review display path. Persist hashes of both the producer bundle and the bound runtime config. supporting_filesnames checked-in orchestration helpers relative to the producing skill root. The platform stages the closed set, verifies every declared SHA256, and rejects traversal, undeclared siblings, or overwrite of a different bundle. Supporting files orchestrate an action; they must never shadow or patch code inside the selected image.- A
torchrundeclaration expresses process topology, not SLURM/Kubernetes syntax. Each platform maps it to its native distributed launcher.
Fixed status vocabulary
Every job state anywhere in the pipeline is exactly one of:
PENDING · RUNNING · COMPLETE · ERROR · CANCELED · UNKNOWN
Platform-native sub-states (ImagePullBackOff, PENDING-because-resources,
Insufficient-GPU, slurm COMPLETING…) are never new states — they ride in
the transition's message field. Terminal = COMPLETE | ERROR | CANCELED.
This is what lets the in-turn poll loop and the detached poller share one code
path across docker/slurm/kubernetes/brev.
Ordering invariants (enforced at the seam, stated here)
- The verify-before-launch gate runs on the spec-bundle, before any job id exists.
tao_job_record.py openwritesPENDING+ the resolvedresults_dirfirst and returns the id — the only handle a launch can use. A submit that skipped the gate has no id, so it cannot launch.transitionsis append-only;.tao/lives outside every synced results tree.
Files
10- BENCHMARK.md
03bf4a7bed6.9 KB - SKILL.md
188bed2dc45.3 KB - config/skillspector-baseline.yaml
8b379b90c3834 B - evals/evals.json
44482288d9903 B - references/best_rec.schema.json
b508e7b0282.3 KB - references/job_record.schema.json
539958a7aa3.9 KB - references/results_dir.contract.md
a9d1c3f2c11.2 KB - references/spec_bundle.schema.json
f833d7d79c7.7 KB - references/tests/test_artifact_schemas.py
fe34b50d6812.3 KB - skill-card.md
8273488c234.3 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.