i4h-workflow-dataset-convert
Convert workflow HDF5 recordings to LeRobot datasets for training or browser inspection. Use for conversion; do not use for replay, augmentation, or raw-data repair.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 4
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 e9bc41027dda1ab3… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Convert Workflow HDF5 to LeRobot
Purpose
Preserve recorded actions, state, cameras, task text, and embodiment labels in a local LeRobot dataset.
Instructions
- Run the checkout resolver and select the source HDF5.
- Read the workflow, Scene, embodiment, and instruction.
- Run conversion for the selected successful episodes.
- Inspect metadata, parquet, videos, and feature widths.
Resolve and inspect
export I4H_WORKFLOWS_REPO_URL="${I4H_WORKFLOWS_REPO_URL:-https://github.com/isaac-for-healthcare/i4h-workflows}"
I4H_REPO_DIR_NAME="${I4H_WORKFLOWS_REPO_URL%/}"
I4H_REPO_DIR_NAME="${I4H_REPO_DIR_NAME##*/}"
I4H_REPO_DIR_NAME="${I4H_REPO_DIR_NAME##*:}"
I4H_REPO_DIR_NAME="${I4H_REPO_DIR_NAME%.git}"
[ -n "$I4H_REPO_DIR_NAME" ] || { echo "Cannot derive a checkout name from I4H_WORKFLOWS_REPO_URL" >&2; exit 2; }
ROOT="${I4H_WORKFLOWS:-$(git rev-parse --show-toplevel 2>/dev/null)}"
if [ ! -d "$ROOT/workflows/i4h_workflows" ]; then
ROOT="${I4H_WORKFLOWS:-$HOME/$I4H_REPO_DIR_NAME}"
[ -d "$ROOT/workflows/i4h_workflows" ] || git clone "$I4H_WORKFLOWS_REPO_URL" "$ROOT"
fi
export I4H_WORKFLOWS="$ROOT"
cd "$ROOT"
HDF5_PATH=/absolute/path/to/recording.hdf5
uv run --project tools/dataset i4h-dataset inspect "$HDF5_PATH" --segments
Treat the resolver above as part of the skill contract: a hosted copy may run outside the base repository, so never assume the current checkout contains workflows/i4h_workflows. I4H_WORKFLOWS_REPO_URL selects the clone source. When I4H_WORKFLOWS is unset, derive the fallback directory from that URL; set I4H_WORKFLOWS only to reuse or choose a specific destination. Never replace an existing checkout.
Use the explicit/current-chain recording. Resolve its workflow and Scene from recording metadata/context, then read the Scene manifest for the embodiment and instruction. Use the embodiment manifest for labels. Do not assume state width equals action width; the converter derives both from the recording.
Convert
RUN_DIR="$(pwd)/runs/<workflow>/$(date +%Y%m%d_%H%M%S)"
DATASET_DIR="$RUN_DIR/lerobot/local/<name>"
mkdir -p "$(dirname "$DATASET_DIR")"
[ ! -e "$DATASET_DIR" ] || { echo "Destination already exists: $DATASET_DIR" >&2; exit 2; }
uv run --project tools/dataset i4h-dataset convert \
"$HDF5_PATH" "$DATASET_DIR" \
--robot <embodiment> \
--repo-id "local/<name>" \
--successful-only \
--task "<instruction>"
Use --fps or --skip-frames only when the user requests it or source metadata justifies it. Keep the default H.264 video codec for compatibility with GR00T's fast decord loader; select another --video-codec only when the target consumer requires it.
Conversion writes aggregate meta/stats.json for downstream policy loaders. Native G1 rule-based WBC recordings already contain 43-D state and 50-D action; the converter recognizes that contract and writes GR00T's required semantic meta/modality.json automatically. For a G1 recording made through the legacy 23-D Pink/keyboard contract and destined for a 50-D G1 WBC policy Task, add --g1-wbc-policy-actions. That explicit mapping combines the measured 43-joint state with the recorded navigation, base-height, and torso commands; require source action width 23 and state width 43.
Verify
Require:
meta/info.jsonmeta/stats.jsonmeta/modality.jsonwhen the target trainer requires semantic modality slices- episode parquet data
- video files for every recorded camera
- converted episode count matching selected successful sources
- action/state feature widths and names matching the recording plus embodiment descriptor
For G1, require modality metadata for both supported paths: native state=43/action=50, or explicitly mapped state=43/source-action=23/output-action=50. Treat a native 50-D dataset without meta/modality.json as incomplete.
Treat missing inputs or zero converted episodes as failure. If conversion leaves a partial destination, quarantine or remove that exact incomplete directory before retrying; never report it as usable.
Troubleshooting
On dimension errors, resolve the source workflow and embodiment again. On missing videos, confirm frames existed before conversion.
Prerequisites
Require a readable HDF5 recording and its matching Scene plus embodiment manifests.
Limitations
Conversion cannot reconstruct missing cameras, actions, state, task text, or successful episodes.
Examples
Convert my scissor pick-and-place recording into a LeRobot dataset.→ resolveso101, convert successful episodes, and verify metadata, parquet, and both camera videos.
Completion gate
Report source HDF5/workflow, embodiment, task text, source/converted/skipped counts, action/state widths, output directory/repo id, aggregate-stats/modality/parquet/video checks, and any missing modality.
Files
4- BENCHMARK.md
388457496b7.0 KB - SKILL.md
51cae38f065.2 KB - evals/evals.json
e1242a00f72.3 KB - skill-card.md
156e2441fd4.0 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related ai-ml skillsscan passed
Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest. Use after ito-compute has booked GPU nodes and the user wants pre-training, fine-tuning, or RL on that metal. ECC implemen
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). Use when the user says "is my dataset okay", "evaluate my data", "check my training data", "I have my own data", or before starting any fine-tuning job. Detects file format, checks schema compliance against