kermt-pretrain-scratch
Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is
- 0
- Installs
- —
- Rating
- —
- Success rate
- 11
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 ec363d3d866f431c… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
kermt-pretrain-scratch
Pretrain a brand-new KERMT model from scratch on a user-provided corpus. Useful
when you want to retrain a model on a custom chemistry domain rather than
extending one of the released checkpoints. Significantly more expensive than
kermt-continue-pretrain — no warm start, so the loss curves need to descend
from scratch over many epochs.
Skill and runtime paths
Set SKILL_DIR to the absolute path of this installed skill directory. Export
KERMT_REPO as the absolute path to the KERMT checkout used for model
execution. The bundled container helper mounts that checkout at
/workspace and this skill at /skill (read-only). Commands inside
the container use /skill/scripts/; defaults are bundled in config/.
Hardware requirements
Same as kermt-continue-pretrain:
-
GPUs: 1–N CUDA-capable. The runner auto-detects via
torch.cuda.device_count();--gpus 0,2overrides. Single-GPU fallback:--batch_size 32 --save_interval 500. Multi-GPU keeps defaults (--batch_size 256etc.). Note:--gpus Nuses torch.cuda indexing, which can differ fromnvidia-smi's display order on multi-GPU hosts (PCI bus vs. CUDA enumeration). To target a specific physical GPU, setCUDA_VISIBLE_DEVICESbefore invoking, or runpython -c "import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])"to confirm which device you're picking. -
VRAM: the default
--batch-size 256is sized for A100-class hardware (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:GPU class VRAM Suggested --batch-sizeL4, T4, V100 16 GB 16–24 GB 32–64 A100 40 GB, L40, A40 40–48 GB 128 A100 80 GB, H100, H200 80 GB 256 (default) These are rough starting points — pass
--batch-size Nto override. -
Disk: tens of GB for shards + vocab + checkpoints, scaled by epochs.
-
Wall time: this is the big difference. Pretraining from scratch on an 11M-mol corpus at 100 epochs typically takes days even on a multi-GPU box. The skill prints an estimate before launching; confirm with the user.
When to invoke
- User wants to train a new model on a custom corpus (e.g. domain-specific chemistry that the released ckpts don't cover).
- User wants to reproduce a pretrain config end-to-end without depending on a released ckpt.
For continuing an existing released ckpt, use kermt-continue-pretrain. For
adding a cMIM decoder to an encoder-only grover_base ckpt, use
kermt-add-cmim-pretrain.
Inputs
Required:
--csv <path>— the pretrain corpus CSV with asmilescolumn. Single file by convention; multi-file corpora deferred. Use--val-csvfor a separate validation set.--pretrain-target-mode {vocab|cmim|hybrid}— which pretrain objective to use. No default — must be set explicitly so the user makes an informed choice:vocab— original GROVER-style atom + bond vocab prediction (encoder-only output, lightweight).cmim— contrastive + SMILES reconstruction objective. Requires building a SMILES vocab from the corpus.hybrid— both vocab and contrastive objectives jointly (the state-of-the-art config from the KERMT manuscript).
Optional:
--val-csv <path>— separate validation CSV. Without it, prepare_data auto-splits the input by--val-frac 0.1(random shuffle with--seed).- Training-hyperparameter overrides:
--epochs N/--batch-size N/--init-lr F/--max-lr F/--final-lr F/--warmup-epochs F/--weight-decay F/--dropout F/--save-interval N/--seed N. Anything not given is filled fromconfig/defaults_pretrain.json. --vocab-loss-weight F(hybrid only) /--latent-dim N/--contrastive-temperature F(cmim and hybrid only).--wandb-project NAME/--wandb-run-name NAME— optional Weights & Biases logging. When--wandb-projectis set, rank 0 logs train/val losses; the run name is honored only alongside a project. Off by default.--gpus 0,2— restrict to a GPU subset.
Workflow
Let $KERMT_REPO be the path to your kermt repo checkout.
-
Pre-flight: ensure container + system probe (same as
kermt-continue-pretrainstep 1). Refuse to proceed ifcheck_systemreports gaps. -
Compute run directory.
RUN_DIR=$KERMT_REPO/runs/pretrain-scratch_$(date -u +%Y-%m-%dT%H-%M-%SZ) -
Validate the corpus (no ckpt to validate, so this is the only input check):
"$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> -- \ "python /skill/scripts/check_data.py --mode pretrain --csv /data/<basename>"Abort on
ok: false. -
Prepare the data — no vocab pass-through (we want fresh vocab from corpus):
"$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> --run-dir $RUN_DIR -- \ "python /skill/scripts/prepare_data.py --mode pretrain \\ --csv /data/<basename> --out /runs/data \\ [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"Outputs land at
$RUN_DIR/data/prepare_data.jsonwithvocab_source: "built_fresh". -
Estimate runtime + warn loudly. This is critical for pretrain-from-scratch:
- "Pretraining from scratch is days-scale even on multi-GPU; the released
KERMT checkpoints were each trained on millions of molecules for hundreds
of GPU-hours. If you mainly want to leverage existing knowledge for a
downstream task, consider
kermt-continue-pretrainfrom a released ckpt instead, which converges in hours instead of days." - Show the corpus size × epochs × GPU count → estimated wall time.
- Ask for explicit confirmation unless
--yeswas given.
- "Pretraining from scratch is days-scale even on multi-GPU; the released
KERMT checkpoints were each trained on millions of molecules for hundreds
of GPU-hours. If you mainly want to leverage existing knowledge for a
downstream task, consider
-
Launch the runner detached.
"$SKILL_DIR/scripts/kermt_container.sh" run_detached \\ --name kermt-pretrain-scratch-<ts> \\ --run-dir $RUN_DIR -- \\ "python /skill/scripts/run_pretrain_local.py \\ --from-scratch --pretrain-target-mode <vocab|cmim|hybrid> \\ --prepare-manifest /runs/data/prepare_data.json \\ --out /runs \\ [--epochs N --batch-size N ...]"Note: NO
--ckptflag (the runner refuses if both--from-scratchand--ckptare given). The runner uses thearchgroup fromconfig/defaults_pretrain.jsonto size the model. -
Report to the user. Always include all of the following — do not omit the TensorBoard line under output-length pressure:
- Container name + id
$RUN_DIR/run.json(the manifest withworkflow: pretrain-scratch,from_scratch: true,vocab_check: null,archfrom defaults, fullcmd_replay)- Log file:
$RUN_DIR/logs/pretrain_ddp.log - TensorBoard:
$RUN_DIR/logs/tb(open withtensorboard --logdir $RUN_DIR/logs/tb) - Suggest
kermt-monitor <RUN_DIR>for progress.
Hard rules
- Never accept a
--ckptflag. From-scratch is exclusive with input ckpt — the runner enforces this; the skill should too. - Never silently default
--pretrain-target-mode. This is a significant architectural choice (vocab = lightweight, hybrid = SOTA). Prompt the user if not given on the CLI. - Strong warning before launching. From-scratch pretrain is the most expensive workflow. The user needs to know what they're committing to.
Common errors
--pretrain-target-mode is required when --from-scratch is set→ user forgot the mode flag. Prompt.--from-scratch is incompatible with --ckpt→ user provided both; ask which one they meant.defaults_pretrain.json has no arch group→ repo state issue (should never happen on a fresh clone); points the user at runningkermt-setupagain.
What's in the manifest after a from-scratch run
Same reproducibility fields as continue-pretrain (repo.commit, kermt_image,
cmd_replay, args_applied), plus:
workflow:"pretrain-scratch"from_scratch:trueinputs.ckpt:nullckpt_symlink:nullvocab_check:null(not verified — vocab built from corpus is authoritative for from-scratch)arch: the values pulled fromconfig/defaults_pretrain.json'sarchgroup (with any future CLI overrides applied).
Replayability
Same as continue-pretrain: cmd_replay is a copy-pasteable command. If
ok_to_replay: false, the kermt repo working tree was dirty at launch
time — check repo.commit and git checkout it first.
Files
11- BENCHMARK.md
394b5866af7.6 KB - SKILL.md
33f9eeb51b9.3 KB - config/defaults_pretrain.json
d4faa67a942.6 KB - evals/evals.json
9a9860397b5.6 KB - scripts/_utils.py
026220a22614.1 KB - scripts/check_checkpoint.py
0bc6cc872920.1 KB - scripts/check_data.py
689268187311.9 KB - scripts/kermt_container.sh
fcdab595c018.9 KB - scripts/prepare_data.py
1426289ab836.5 KB - scripts/run_pretrain_local.py
3c54d5371534.8 KB - skill-card.md
8ce73fe7304.9 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.