kermt-finetune
Finetune a pretrained KERMT encoder on a labeled CSV. Validate the checkpoint and data, prepare features, and run containerized training. Use a local checkpoint or optionally download a pinned Hugging Face model bundle using HF_TOKEN if configured. Write model bundles, prepared data, logs, and train
- 0
- Installs
- —
- Rating
- —
- Success rate
- 14
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 25342402b51800a0… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
kermt-finetune
Finetune a pretrained KERMT encoder on a user-supplied labeled CSV. The skill is the workflow orchestrator: validate ckpt, validate data, prepare data, launch the runner detached, return a run directory + container name.
Skill and runtime paths
Set SKILL_DIR to the absolute path of this installed skill directory. Export
KERMT_REPO as the absolute path to the KERMT checkout used for model
execution. The bundled container helper mounts that checkout at
/workspace and this skill at /skill (read-only). Commands inside
the container use /skill/scripts/; defaults are bundled in config/.
See Released models for checkpoint bundle requirements.
Downloads and local outputs
The optional released-model branch reads config/released_model.json for the
Hugging Face repository, pinned revision, and filenames. The bundled
scripts/fetch_released_model.py downloads the model bundle over HTTPS into
the host directory the user selects. Public models work without credentials;
if HF_TOKEN is set, the container helper forwards it for Hugging Face
authentication. Prepared data, logs, and workflow results go into the chosen
run directory.
Hardware requirements
- GPUs: 1 by default (single-GPU); pass
--gpus 0(or whichever id) to select one. For faster training on a multi-GPU host, pass--num-gpus N(N>1) to run data-parallel DDP across N GPUs —--batch-sizeis then per-GPU (effective global batch = batch_size × N). - VRAM: ≥ 8 GB for the default
batch_size 32configuration. Lower VRAM works at smaller batch sizes — pass--batch-size Nto override. - Disk: a few GB per run (checkpoint + features + logs).
- Driver / CUDA: any host supporting CUDA 12.6 (the kermt image base).
kermt-setupvalidates this up-front.
Inputs
Required:
--csv <path>— labeled CSV. First column issmiles; every other column is a target.
Checkpoint (optional — defaults to the released model if omitted):
--ckpt <path>— input pretrain checkpoint (grover_base / cmim / hybrid). The validator refuses already-finetuned ckpts with a redirect tokermt-infer. If omitted, the skill offers to download the released pretrained hybrid model nvidia/NV-KERMT-70M-v2 and finetune from it — see "Resolve & validate the checkpoint" (workflow step 3).--pretrained-release— explicit opt-in to use the released model without the interactive prompt (for non-interactive / agent runs). Mutually exclusive with--ckpt.--model-dir <dir>— where to save the downloaded bundle (default$KERMT_REPO/models/NV-KERMT-70M-v2/). An already-complete bundle there is reused, not re-downloaded.
Optional:
-
--dataset-type {regression | classification | multiclass}— defaultregression(fromdefaults_finetune.json). Drives loss, metric defaults, and head initialization. For classification tasks pass--dataset-type classification. -
--targets COL [COL ...]— explicit target column names. If omitted, the validator auto-detects numeric non-smiles columns and the skill confirms with the user before proceeding. -
--val-csv <path>and--test-csv <path>— user-provided val + test splits. Either pass both or pass neither (the skill auto-splits using the configured--split-type). -
--split-type {random | scaffold_balanced | index_predetermined}— defaultscaffold_balancedfromdefaults_finetune.json.randomandscaffold_balanced: build the val/test split internally from the train CSV. No--val-csv/--test-csvneeded.index_predetermined: requires pre-split CSVs passed via--val-csv+--test-csv(and, separately, per-fold index files — seekermt/util/utils.split_data). Use this when the dataset ships its own canonical split (e.g.tests/data/Biogen_for_grover/scaffold/ balance/<endpoint>/{train,val,test}.csv).
-
--metric NAME—mae(regression default),auc(classification default), or any namekermt.util.metrics.get_metric_funcaccepts. -
--epochs N/--batch-size N/--init-lr F/--max-lr F/--final-lr F/--warmup-epochs F/--weight-decay F/--dropout F/--bond-drop-rate F/--dist-coff F/--early-stop-epoch N/--seed N— training-hyperparameter overrides. Anything not given is filled fromconfig/defaults_finetune.json. -
--ffn-hidden-size N/--ffn-num-layers N— shared FFN trunk dims. -
--ffn-num-task-specific-layers N/--ffn-task-specific-hidden-size H— per-target FFN heads (default 0 = off; useful for heterogeneous multi-target finetunes). Both must be set together when N > 0. -
--ensemble-size N/--num-folds N— multi-model / k-fold CV. Default 1 each. -
--gpus 0— single GPU id for single-process finetune (default 0). Ignored when--num-gpus > 1. -
--num-gpus N— number of GPUs for data-parallel DDP finetune. Default 1 (single-process, unchanged). N>1 runsmain.py finetunewithWORLD_SIZE=N(one process per GPU);--batch-sizeis per-GPU. -
--from-prepare <dir>— skip the prepare step and reuse an existingprepare_data.jsonin<dir>. Useful when iterating on hyperparameters.
Workflow
Let $KERMT_REPO be the path to your kermt repo checkout, and assume
kermt-setup has built kermt:latest. All paths below are on the host; the
helper bind-mounts them at known container paths.
-
Pre-flight: ensure container + system probe.
"$SKILL_DIR/scripts/kermt_container.sh" check_system | python -c " import json, sys; d = json.load(sys.stdin) if not d['ok']: print('System check failed:', d['gaps']); sys.exit(1) print(f'OK: {len(d[\"gpus\"])} GPU(s); CUDA via container toolkit') "Refuse to proceed if
ok: false. -
Compute run directory.
RUN_DIR=$KERMT_REPO/runs/finetune_$(date -u +%Y-%m-%dT%H-%M-%SZ) -
Resolve & validate the checkpoint.
Resolve — only if
--ckptwas omitted. Default to the released pretrained hybrid model nvidia/NV-KERMT-70M-v2:- Consent gate. Unless
--pretrained-releasewas passed, ask the user: "No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2 (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2) and finetune from it? [y/N]". Never download without an explicit yes (or--pretrained-release). If both--ckptand--pretrained-releaseare given, abort — they conflict. - Save location. Default
$KERMT_REPO/models/NV-KERMT-70M-v2/; honor--model-dir <dir>if given. An already-complete bundle is reused. - Download (foreground; ~282 MB on first fetch):
Parse the JSON; abort on"$SKILL_DIR/scripts/kermt_container.sh" run --model-dir <save-dir> -- \ "python /skill/scripts/fetch_released_model.py --out /model"ok: false(surfaceerrors). On success set<user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt.
Validate the resolved (or user-provided) ckpt:
"$SKILL_DIR/scripts/kermt_container.sh" run --ckpt <user-ckpt> -- \ "python /skill/scripts/check_checkpoint.py --mode finetune_init --ckpt /ckpt"Parse the JSON. Abort on
ok: false. The validator rejects already- finetuned ckpts (has_task_ffn: true) with a redirect tokermt-infer. - Consent gate. Unless
-
Validate the data.
"$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> -- \ "python /skill/scripts/check_data.py --mode finetune --csv /data/<basename> [--targets COL1 COL2 ...]"If
--targetswas not given by the user, surfaceauto_detected_targetsfrom the JSON and ask the user to confirm before continuing. Abort onok: false. -
Prepare the data (skip if
--from-preparegiven).Pre-flight: check for sibling val.csv / test.csv. Before invoking prepare_data, inspect the parent directory of
<user-csv>. If a canonical-looking siblingval.csv(orval_*.csv— common variants includeval_T.csv,val_clean.csv) AND a matchingtest.csv/test_*.csvexist next to the train CSV, the dataset ships its own pre-defined split. In that case set--split-type index_predeterminedAND pass--val-csv/--test-csv— otherwise the configuredsplit_type(defaultscaffold_balanced) will re-split the train CSV from scratch and silently discard the user's val/test files. When in doubt — or when the sibling files use non-canonical suffixes (_T,_v2, etc.) — surface the situation to the user and ask which they want.Quoting target names. If any of the
--targetscolumn names contain shell metacharacters (>,&,|,(,),$, etc.), single-quote each one when passing on the CLI to keep the shell from eating part of the name. Example:--targets 'Log_Caco2_Papp_A>B' 'logD'. The CSV header itself is read directly by the downstream trainer and is unaffected, but the prepare_data.json manifest'stargets[]field captures whatever the shell delivers — unquoted metacharacters get truncated there.Mount note:
kermt_container.sh --data <host-csv>mounts the parent directory of<host-csv>at/data.--val-csvand--test-csvmust therefore reference files in that same parent directory. If val/test live in a separate directory (e.g. a siblingsplits/folder), mount the parent of all three using--data <dir>on a directory rather than a file."$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> --run-dir $RUN_DIR -- \ "python /skill/scripts/prepare_data.py --mode finetune \\ --csv /data/<basename> --out /runs/data \\ --split-type <split_type> \\ [--val-csv /data/<val-basename> --test-csv /data/<test-basename>] \\ [--val-frac 0.1 --test-frac 0.1 --seed 0] \\ --targets <COL1> [COL2 ...]"Outputs land at
$RUN_DIR/data/prepare_data.json. Forscaffold_balancedandindex_predetermined, prep emits a singleclean_full_csv+.npz; the runner passes them through tomain.py finetunewhich callssplit_datainternally with the user-supplied seed. -
Estimate runtime + echo applied defaults.
- Finetune wall time is typically minutes-to-hours on 1 GPU.
- Surface a summary of every flag that was filled from the defaults
vs user-supplied, so the user knows what was assumed. The runner
records this in
args_applied. - Sample message:
"Filling from defaults_finetune.json: epochs=30, batch_size=32, split_type=scaffold_balanced. Override any of these with --<flag>."
-
Targets confirmation gate (hard requirement). Before launching the runner, regardless of how the targets list was determined (CLI
--targets, auto-detection in step 4, or a user natural-language request like "finetune on Caco2 and HLM"), echo the final targets list to the user with an explicit count:"Will finetune on N target(s): COL1, COL2, ...". If the user's request specified a subset that doesn't match this list (e.g., they asked for 2 tasks via natural language but the list still has 4), treat it as a discrepancy and re-prompt with the diff — never silently proceed on the wrong target set. Wait for explicit confirmation before launching unless--yeswas given. -
Launch the runner detached. (Consistent with the pretrain skills.)
"$SKILL_DIR/scripts/kermt_container.sh" run_detached \\ --name kermt-finetune-<ts> \\ --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\ "python /skill/scripts/run_finetune_local.py \\ --ckpt /ckpt \\ --prepare-manifest /runs/data/prepare_data.json \\ --dataset-type <type> \\ --out /runs \\ [--gpus 0] \\ [--num-gpus N] \\ [--epochs N --batch-size N --init-lr F ...] \\ [--ffn-num-task-specific-layers N --ffn-task-specific-hidden-size H]"Returns the container name + id + log file path.
-
Report to the user. Output a short summary:
- Container name + id
$RUN_DIR/run.json(manifest with cmd_replay + image digest)- Log file:
$RUN_DIR/logs/finetune.log - TensorBoard:
$RUN_DIR/logs/tb(open withtensorboard --logdir $RUN_DIR/logs/tb) - Final checkpoints land at
$RUN_DIR/ckpt/fold_0/model_0/model.pt(best-val) andlast_checkpoint.pt(sibling, auto-resume target). Held-out test predictions + metrics land at$RUN_DIR/ckpt/fold_0/test_result.csv. Paths vary with--num-folds/--ensemble-size. - To follow progress:
kermt-monitor <RUN_DIR>(one-shot) ordocker logs -f <container-name>(streaming). - To block until the run finishes (useful for short test runs):
docker wait <container-name>— prints the exit code on completion.
Hard rules
- Never download the released model without consent. When
--ckptis omitted, downloadnvidia/NV-KERMT-70M-v2only after an explicit user "yes" or an explicit--pretrained-releaseflag.--ckptand--pretrained-releaseare mutually exclusive. - Never modify the user's input ckpt. The runner passes its path via
--checkpoint_path;task/train.pyloads it read-only into the model and attaches a new FFN head. The source file stays untouched. - Arch comes from the ckpt, not from CLI/defaults. The runner extracts
hidden_size,depth,num_attn_head,activation,embedding_output_type,self_attention(+attn_hidden/attn_outwhen applicable) from the ckpt's saved_args. There is no--hidden-sizeflag on this runner. - Never block on the long-running finetune. The skill launches via
run_detachedand returns immediately after step 9. Usekermt-monitor. - Echo applied defaults back to the user. The
args_appliedfield ofrun.jsonrecords every flag's value + source (user / default-config). Surface a one-line summary of every filled-from-default flag so the user knows what was assumed.
Common errors
finetune_init requires a pretrain ckpt (grover_base / cmim / hybrid)→ the ckpt you passed is already finetuned (has task FFN heads). Pick a pretrain ckpt instead, or usekermt-inferif you want to run predictions with the existing finetuned model. To resume a finetune on the SAME dataset, bypass the skill and callpython main.py finetune --checkpoint_path <ckpt> ...directly — the agent skill doesn't support resume because saved-task identity can't be machine-verified against the new training data.prepare_data manifest reports ok=False→ checkerrorsfor the failed step (typically clean_smiles or save_features). Fix and re-run.ffn_num_task_specific_layers=N>0 but ffn_task_specific_hidden_size is unset→ MTL heads need an explicit hidden size. Pass--ffn-task-specific-hidden-size H.finetune is single-GPU(from--gpus 0,1) →--gpusselects one device for single-process finetune. For multi-GPU, use--num-gpus N(DDP) instead.
Replayability
The run.json cmd_replay field is a single-line command that re-runs the
finetune with the same inputs, hyperparameters, and arch. To replay inside
the kermt container:
$(jq -r .cmd_replay $RUN_DIR/run.json)
If ok_to_replay: false in the manifest (because the kermt repo working
tree was dirty at launch time), the replay may not be bit-exact — pin the
exact commit via the repo.commit field and git checkout it
first.
Files
14- BENCHMARK.md
fc4b07b20f7.6 KB - SKILL.md
bf157bd45316.1 KB - config/defaults_finetune.json
095c0f1fa12.6 KB - config/released_model.json
2858044932398 B - evals/evals.json
d0a1695d6e7.3 KB - references/released-models.md
8566f0a72f1.6 KB - scripts/_utils.py
026220a22614.1 KB - scripts/check_checkpoint.py
0bc6cc872920.1 KB - scripts/check_data.py
689268187311.9 KB - scripts/fetch_released_model.py
4d0e6485297.9 KB - scripts/kermt_container.sh
fcdab595c018.9 KB - scripts/prepare_data.py
1426289ab836.5 KB - scripts/run_finetune_local.py
aeb5355f0521.4 KB - skill-card.md
4346e6790d4.5 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related ai-ml skillsscan passed
Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest. Use after ito-compute has booked GPU nodes and the user wants pre-training, fine-tuning, or RL on that metal. ECC implemen
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v
Validates the user's environment for SageMaker AI operations — checks SDK version, AWS region, and execution role. Use when the user says "set up", "getting started", "check my environment", "configure SDK", or as the first step in any plan involving SageMaker/Bedrock training, evaluation, or deploy