kermt-add-cmim-pretrain
Convert a grover_base checkpoint (encoder-only or encoder + vocab heads) into a hybrid checkpoint by adding a randomly-initialized cMIM decoder + latent_dist, then continue pretraining on the user's corpus as hybrid (vocab + contrast). Effectively kermt-continue-pretrain with a one-time ckpt-convers
- 0
- Installs
- —
- Rating
- —
- Success rate
- 12
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 dd1d1891338ac441… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
kermt-add-cmim-pretrain
Convert a grover_base checkpoint (legacy original-GROVER grover.encoders.*
or modern kermt.encoders.*, with or without vocab heads) into a fully-formed
hybrid (cMIM + vocab) checkpoint, then continue pretraining on the user's
corpus as hybrid.
This is a thin wrapper: upgrade_to_hybrid.py produces a new ckpt that
classifies as model_type: hybrid via check_checkpoint.py, and the rest of
the workflow is identical to kermt-continue-pretrain.
Status: experimental. This workflow is functional end-to-end but has not been benchmarked against the manuscript's from-scratch hybrid training (which produces the released checkpoint). Use as an experimental alternative to
kermt-pretrain-scratchwhen you want to extend an existing grover_base checkpoint rather than restart from random init. Validate downstream performance on your own benchmark before relying on the upgraded ckpt for production work.
Skill and runtime paths
Set SKILL_DIR to the absolute path of this installed skill directory. Export
KERMT_REPO as the absolute path to the KERMT checkout used for model
execution. The bundled container helper mounts that checkout at
/workspace and this skill at /skill (read-only). Commands inside
the container use /skill/scripts/; defaults are bundled in config/.
Hardware requirements
Same as kermt-continue-pretrain (the cMIM decoder adds parameters but not
substantially; VRAM headroom should be fine). The upgrade step itself is
fast (~5 s) and CPU-only — only the subsequent continue-pretrain consumes
GPU.
When to invoke
- User has a grover_base checkpoint (encoder-only or with vocab heads) and wants to extend it into a hybrid (vocab + cMIM contrastive) pretrain.
- Useful for adding the SMILES-reconstruction contrastive objective to a
pretrained encoder without restarting pretraining from scratch (which
kermt-pretrain-scratchwould do at days-scale).
For continuing an existing hybrid or cmim ckpt: use kermt-continue-pretrain
directly. For training a fresh model on a custom corpus: use
kermt-pretrain-scratch.
Inputs
Required:
--ckpt <path>— grover_base ckpt to upgrade. Validated viacheck_checkpoint.py --mode upgrade_to_hybrid; rejected if the ckpt already has a contrast head or task FFN.--csv <path>— pretrain corpus CSV. Same shape askermt-continue-pretrain's--csvinput.
Optional (same as kermt-continue-pretrain):
--val-csv <path>— separate validation CSV. Without it, prepare_data auto-splits by--val-frac 0.1.- Training-hyperparameter overrides (
--epochs N,--batch-size N, lr triple,--warmup-epochs F, etc.). --vocab-loss-weight F/--latent-dim N/--contrastive-temperature F.--wandb-project NAME/--wandb-run-name NAME— optional Weights & Biases logging (run name honored only alongside a project). Off by default.--gpus 0,2.
Workflow
Let $KERMT_REPO be the path to your kermt repo checkout.
-
Pre-flight: check_system (same as
kermt-continue-pretrainstep 1). -
Compute run directory:
RUN_DIR=$KERMT_REPO/runs/add-cmim-pretrain_$(date -u +%Y-%m-%dT%H-%M-%SZ) -
Validate the input ckpt with
check_checkpoint --mode upgrade_to_hybrid. Abort onok: false. The validator rejects ckpts that already have contrast head (suggestkermt-continue-pretrain) or task FFN heads (the ckpt has been finetuned; suggest using the original pretrain checkpoint). -
Validate the corpus via
check_data --mode pretrain. Abort onok: false. -
Prepare the data with
--mode pretrain— without--vocab-dir. The upgrade builds fresh vocab heads sized to the corpus's vocab, so we wantprepare_datato produce a new vocab from the corpus rather than passing through the ckpt's old vocab (which may not even exist for encoder-only legacy grover_base ckpts):"$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> --run-dir $RUN_DIR -- \ "python /skill/scripts/prepare_data.py --mode pretrain \\ --csv /data/<basename> --out /runs/data \\ [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"The output manifest has
vocab_source: "built_fresh"and includes asmiles_vocab(built from the corpus, needed for the new decoder). -
Upgrade the ckpt.
"$SKILL_DIR/scripts/kermt_container.sh" run --ckpt <user-ckpt> --run-dir $RUN_DIR -- \ "python /skill/scripts/upgrade_to_hybrid.py \\ --ckpt /ckpt \\ --prepare-manifest /runs/data/prepare_data.json \\ --out /runs/upgraded.pt"Surface the JSON summary to the user — especially
warnings[], which includes any encoder-arch drift notes (e.g. legacy GROVER had two extraact_func_*keys that modern KERMTEmbedding doesn't) and the pretrain_ddp.py--backboneargparse-restriction note if the upgraded ckpt's backbone is anything other thangtrans. -
Estimate runtime + confirm with the user. Same heuristic as
kermt-continue-pretrain(corpus size × epochs × GPU count → wall time). -
Launch the runner detached.
"$SKILL_DIR/scripts/kermt_container.sh" run_detached \\ --name kermt-add-cmim-pretrain-<ts> \\ --run-dir $RUN_DIR -- \\ "python /skill/scripts/run_pretrain_local.py \\ --ckpt /runs/upgraded.pt \\ --prepare-manifest /runs/data/prepare_data.json \\ --out /runs \\ [--epochs N --batch-size N ...]"The runner sees the upgraded ckpt as
model_type: hybrid, so it auto-dispatches--pretrain_mode hybrid --vocab_loss_weight 1.0with smiles_vocab plumbed through. -
Report to the user with the upgraded ckpt path + the same run.json pointer / log path / tensorboard URL pattern as
kermt-continue-pretrain.
Hard rules
- Never modify the user's input ckpt. The upgrade writes a new file at
<run_dir>/upgraded.pt; the source ckpt stays untouched. - Vocab heads are always fresh. Even if the input grover_base has vocab heads, they're discarded and rebuilt sized to the new corpus's vocab. Continue-pretraining the upgraded ckpt will train those new heads alongside the decoder.
- Don't auto-relax
--backbonechoices. If the upgrade warning fires because the input ckpt's backbone isn'tgtrans(e.g. legacydualtrans), surface the warning and ask the user. Do NOT silently modify parsing.py to add the legacy backbone to the choices list.
Common errors
check_checkpoint rejected the ckptwith model_type=hybrid or cmim → user's ckpt already has a contrast head. Redirect tokermt-continue-pretrain.check_checkpoint rejected the ckptwith task_ffn=true → the ckpt has been finetuned. The upgrade workflow only supports pretrain checkpoints.prepare manifest missing smiles_vocab→ prepare_data was invoked with--skip-vocabor some equivalent that omitted the smiles vocab. Re-run prepare without those flags.unexpected key(s) in encoder loadwarning → legacy GROVER architectures saved a couple ofact_func_*weights that modern KERMTEmbedding doesn't use. Benign; the rest of the encoder loaded correctly.
What's in run.json after a successful run
Same reproducibility fields as kermt-continue-pretrain, plus the upgrade step's
summary.json is captured under the inputs.upgrade_summary path so the
provenance of the upgraded ckpt is auditable.
Replayability
Same as kermt-continue-pretrain: cmd_replay rebuilds the
run_pretrain_local.py --ckpt <upgraded.pt> ... invocation. To redo the
full add-cmim flow end-to-end, the user also needs the input grover_base
ckpt and the corpus — both are captured in the prepare_data and upgrade
manifests by absolute path.
Files
12- BENCHMARK.md
78129201e27.6 KB - SKILL.md
eed773f2958.4 KB - config/defaults_pretrain.json
d4faa67a942.6 KB - evals/evals.json
dff010908d5.6 KB - scripts/_utils.py
026220a22614.1 KB - scripts/check_checkpoint.py
0bc6cc872920.1 KB - scripts/check_data.py
689268187311.9 KB - scripts/kermt_container.sh
fcdab595c018.9 KB - scripts/prepare_data.py
1426289ab836.5 KB - scripts/run_pretrain_local.py
3c54d5371534.8 KB - scripts/upgrade_to_hybrid.py
cc62c6424c18.2 KB - skill-card.md
169bb07a6a4.4 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related knowledge skillsscan passed
Restrict file edits to a specific directory for the session. (gstack)
PostHog logs for Java
Operate execution flow across GitHub and Linear by triaging issues and pull requests, linking active work, and keeping GitHub public-facing while Linear remains the internal execution layer. Use when the user wants backlog control, PR triage, or GitHub-to-Linear coordination.