skills/ NVIDIA/skills

i4h-workflow-dataset-convert

Convert workflow HDF5 recordings to LeRobot datasets for training or browser inspection. Use for conversion; do not use for replay, augmentation, or raw-data repair.

0
Installs
—
Rating
—
Success rate
4
Files scanned
Scan passedai-ml
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

4 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 e9bc41027dda1ab3… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Convert Workflow HDF5 to LeRobot

Purpose

Preserve recorded actions, state, cameras, task text, and embodiment labels in a local LeRobot dataset.

Instructions

  1. Run the checkout resolver and select the source HDF5.
  2. Read the workflow, Scene, embodiment, and instruction.
  3. Run conversion for the selected successful episodes.
  4. Inspect metadata, parquet, videos, and feature widths.

Resolve and inspect

export I4H_WORKFLOWS_REPO_URL="${I4H_WORKFLOWS_REPO_URL:-https://github.com/isaac-for-healthcare/i4h-workflows}"
I4H_REPO_DIR_NAME="${I4H_WORKFLOWS_REPO_URL%/}"
I4H_REPO_DIR_NAME="${I4H_REPO_DIR_NAME##*/}"
I4H_REPO_DIR_NAME="${I4H_REPO_DIR_NAME##*:}"
I4H_REPO_DIR_NAME="${I4H_REPO_DIR_NAME%.git}"
[ -n "$I4H_REPO_DIR_NAME" ] || { echo "Cannot derive a checkout name from I4H_WORKFLOWS_REPO_URL" >&2; exit 2; }
ROOT="${I4H_WORKFLOWS:-$(git rev-parse --show-toplevel 2>/dev/null)}"
if [ ! -d "$ROOT/workflows/i4h_workflows" ]; then
  ROOT="${I4H_WORKFLOWS:-$HOME/$I4H_REPO_DIR_NAME}"
  [ -d "$ROOT/workflows/i4h_workflows" ] || git clone "$I4H_WORKFLOWS_REPO_URL" "$ROOT"
fi
export I4H_WORKFLOWS="$ROOT"
cd "$ROOT"
HDF5_PATH=/absolute/path/to/recording.hdf5
uv run --project tools/dataset i4h-dataset inspect "$HDF5_PATH" --segments

Treat the resolver above as part of the skill contract: a hosted copy may run outside the base repository, so never assume the current checkout contains workflows/i4h_workflows. I4H_WORKFLOWS_REPO_URL selects the clone source. When I4H_WORKFLOWS is unset, derive the fallback directory from that URL; set I4H_WORKFLOWS only to reuse or choose a specific destination. Never replace an existing checkout.

Use the explicit/current-chain recording. Resolve its workflow and Scene from recording metadata/context, then read the Scene manifest for the embodiment and instruction. Use the embodiment manifest for labels. Do not assume state width equals action width; the converter derives both from the recording.

Convert

RUN_DIR="$(pwd)/runs/<workflow>/$(date +%Y%m%d_%H%M%S)"
DATASET_DIR="$RUN_DIR/lerobot/local/<name>"
mkdir -p "$(dirname "$DATASET_DIR")"
[ ! -e "$DATASET_DIR" ] || { echo "Destination already exists: $DATASET_DIR" >&2; exit 2; }
uv run --project tools/dataset i4h-dataset convert \
  "$HDF5_PATH" "$DATASET_DIR" \
  --robot <embodiment> \
  --repo-id "local/<name>" \
  --successful-only \
  --task "<instruction>"

Use --fps or --skip-frames only when the user requests it or source metadata justifies it. Keep the default H.264 video codec for compatibility with GR00T's fast decord loader; select another --video-codec only when the target consumer requires it.

Conversion writes aggregate meta/stats.json for downstream policy loaders. Native G1 rule-based WBC recordings already contain 43-D state and 50-D action; the converter recognizes that contract and writes GR00T's required semantic meta/modality.json automatically. For a G1 recording made through the legacy 23-D Pink/keyboard contract and destined for a 50-D G1 WBC policy Task, add --g1-wbc-policy-actions. That explicit mapping combines the measured 43-joint state with the recorded navigation, base-height, and torso commands; require source action width 23 and state width 43.

Verify

Require:

  • meta/info.json
  • meta/stats.json
  • meta/modality.json when the target trainer requires semantic modality slices
  • episode parquet data
  • video files for every recorded camera
  • converted episode count matching selected successful sources
  • action/state feature widths and names matching the recording plus embodiment descriptor

For G1, require modality metadata for both supported paths: native state=43/action=50, or explicitly mapped state=43/source-action=23/output-action=50. Treat a native 50-D dataset without meta/modality.json as incomplete.

Treat missing inputs or zero converted episodes as failure. If conversion leaves a partial destination, quarantine or remove that exact incomplete directory before retrying; never report it as usable.

Troubleshooting

On dimension errors, resolve the source workflow and embodiment again. On missing videos, confirm frames existed before conversion.

Prerequisites

Require a readable HDF5 recording and its matching Scene plus embodiment manifests.

Limitations

Conversion cannot reconstruct missing cameras, actions, state, task text, or successful episodes.

Examples

  • Convert my scissor pick-and-place recording into a LeRobot dataset. → resolve so101, convert successful episodes, and verify metadata, parquet, and both camera videos.

Completion gate

Report source HDF5/workflow, embodiment, task text, source/converted/skipped counts, action/state widths, output directory/repo id, aggregate-stats/modality/parquet/video checks, and any missing modality.

Files

4
18.5 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from NVIDIA/skills8

accelerated-computing-cudf

Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.

Needs review 0
aiq-deploy

Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.

Needs review 0
aiq-research

Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.

Scan passed 0
ambient-healthcare-agent-with-nemotron-voice-agent

Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.

Needs review 0
amc-run-rtsp-calibration

Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.

Scan passed 0
amc-run-sample-calibration

Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.

Scan passed 0
amc-run-video-calibration

Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.

Scan passed 0
amc-setup-calibration-stack

Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.

Needs review 0

Related ai-ml skillsscan passed

ito-training

Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest. Use after ito-compute has booked GPU nodes and the user wants pre-training, fine-tuning, or RL on that metal. ECC implemen

Scan passed 0
pair-agent

Pair a remote AI agent with your browser. (gstack)

Scan passed 0
ce-noslop

Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market

Scan passed 0
superjson

Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi

Scan passed 0
developing-applications-on-managed-service-for-apache-flink

MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v

Scan passed 0
dataset-evaluation

Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). Use when the user says "is my dataset okay", "evaluate my data", "check my training data", "I have my own data", or before starting any fine-tuning job. Detects file format, checks schema compliance against

Scan passed 0