tao-run-on-virtualenv
Run a Python training/eval script directly in an existing local virtualenv — no docker, no container. Implements the four-verb consumer contract (submit/status/logs/cancel) over a vendored process-lifecycle runner with durable on-disk state, PID-reuse-safe identity, and process-group cleanup. Use fo
- 0
- Installs
- —
- Rating
- —
- Success rate
- 8
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 bd44ea217aacbffc… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Virtualenv — docker-free local Python execution
Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
The virtualenv platform runs a Python script natively in an existing venv —
as an argv vector whose first element is <venv>/bin/python, never through a
shell, never activating anything. The vendored runner
(references/virtualenv_runner.py) is this platform's "native CLI" — the role
docker/kubectl/sbatch play elsewhere — and owns only the process
lifecycle. Job records stay with tao_job_record.py; specs are authored by the
agent, exactly like every other platform.
When to use
- The workload is a plain Python script (its dependencies pip-installed in a venv), not a TAO container action.
- No docker on the host, or container startup cost isn't worth it (fast smokes, AutoML trial loops over lightweight models).
- Single node only. For TAO container actions use
tao-run-on-docker; for clusters use-slurm/-kubernetes.
Preflight
# 1. The venv is real and has an executable interpreter.
[ -f "$VENV/pyvenv.cfg" ] && [ -x "$VENV/bin/python" ] || echo "MISSING: $VENV is not a venv"
# 2. The script's top-level imports resolve inside it (catches wrong-venv early);
# substitute the real modules your script imports.
"$VENV/bin/python" -c "import torch" || echo "MISSING: script dependency not in $VENV"
# 3. GPU visibility only if the script needs CUDA.
nvidia-smi >/dev/null 2>&1 || echo "note: no GPU visible (fine for CPU scripts)"
No credentials are required by the platform itself; model-specific env vars
(e.g. HF_TOKEN) pass through by NAME with -e (values never land on argv).
Storage
Tier A by definition — everything is local paths. Datasets must already be
on local disk (stage with tao-data-io first if they live in S3). Outputs land
in the job record's results_dir, which IS the runner's --job-dir.
Execution — the four verbs
$BANK = ${TAO_SKILL_BANK_PATH}; $RUNNER =
$BANK/skills/platform/tao-run-on-virtualenv/references/virtualenv_runner.py.
submit
- Author the spec (if the script takes one) at a local path — nested
dicts, never flat dotted keys — and lint the assembled command with
redact_secrets.py lint. - Open the record — mints the id, binds
results_dirBEFORE launch:JOB_ID=$("$BANK/scripts/tao_job_record.py" open --platform virtualenv \ --image "$VENV/bin/python" --network-arch "$ARCH" --action "$ACTION" \ --storage-tier A --results-root "$RESULTS_ROOT") RESULTS_DIR="$RESULTS_ROOT/$JOB_ID" - Launch detached (the runner writes a durable wrapper that gates start,
records identity, and cleans up the process group on exit):
Placeholdersset -a; source /path/to/.env; set +a # omit if already exported python3 "$RUNNER" submit --job-dir "$RESULTS_DIR" --venv "$VENV" \ --script train.py --job-id "$JOB_ID" --config-path "$SPEC" \ --arg train --arg=--config={config_path} --arg=--out={results_dir} \ --gpu-ids 0 -e HF_TOKEN{config_path}{results_dir}{job_id}render inside--argtokens. A token starting with-must use the--arg=TOKENform (argparse).--gpu-idssetsCUDA_VISIBLE_DEVICES;--gpus 0hides GPUs; neither reserves anything. - Record RUNNING with the pid the runner printed:
"$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state RUNNING --backend-ref "pid:<pid>"
One submit per job dir — a retry gets a NEW record (--retry-of), never a
re-submit into the same dir.
status
python3 "$RUNNER" status --job-dir "$RESULTS_DIR" # {"status": "...", ...}
Prints the fixed vocabulary directly: PENDING RUNNING COMPLETE ERROR CANCELED UNKNOWN — no mapping table needed. Status is derived from durable files
(exit_status.json, launcher identity) and is safe to poll from any process,
any time, including after reboots of the polling agent. On a terminal status,
mark the record.
logs
python3 "$RUNNER" logs --job-dir "$RESULTS_DIR" --tail 200
cancel
python3 "$RUNNER" cancel --job-dir "$RESULTS_DIR"
"$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state CANCELED --source agent
Cancel marks first (a not-yet-started wrapper self-cancels at its start gate),
verifies process identity (never kills a reused PID), then SIGTERM→SIGKILLs the
whole process group. already_terminal in the reply means the job finished
before the cancel — mark the record with the status it reports instead.
Platform caveats
- Linux first-class. Identity and group cleanup use
/proc; on macOS the runner falls back tops/pgrep— fine for local smokes, but GPU training targets are Linux hosts. - No multi-node, no image resolution — there is no container. The "image" recorded is the venv's interpreter path.
- The runner never downloads anything. Remote inputs are the agent's job to
stage first (
tao-data-io).
Files
8- BENCHMARK.md
1d4d5ece2e7.0 KB - SKILL.md
88c47293886.0 KB - config/skillspector-baseline.yaml
4d2667c0841.5 KB - evals/evals.json
eefd3a62171022 B - references/skill_info.yaml
dc89777d91605 B - references/tests/test_virtualenv_runner.py
731cb637a020.5 KB - references/virtualenv_runner.py
79633f192136.6 KB - skill-card.md
fdaff107654.1 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related devops skillsscan passed
Build monitoring dashboards that answer real operator questions for Grafana, SigNoz, and similar platforms. Use when turning metrics into a working dashboard instead of a vanity board.
Configure deployment settings for /land-and-deploy.
Design, configure, troubleshoot, or review Cloudflare One Zero Trust and SASE deployments. Use cloudflare-one-migrations for migration planning from other vendors.
Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea
Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil