codonfm-score
Validate, prepare, or run public CodonFM Encodon masked-codon variant scoring and review compatibility of its scoring workflows. Use only when the user explicitly requests CodonFM or Encodon, or that context is already established in the conversation. Do not select this skill for a generic variant-s
- 0
- Installs
- —
- Rating
- —
- Success rate
- 7
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 b48b2c84272a8e4f… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Score variants with public Encodon
Run general masked-codon mutation_prediction only. This produces a research
signal, not a clinical diagnosis or an expression-direction prediction.
Instructions
First confirm CodonFM or Encodon context in the user's request or established conversation. If that context is missing, ask for any missing variant details and the intended analysis before choosing a model or inspecting model-specific files. The presence of this skill or source files alone does not establish the user's intent.
For source reviews and command preparation, inspect the supplied source and
metadata without installing the ML runtime. Use an available Python 3
interpreter with standard-library zipfile, json, and csv; do not assume
the python alias or unzip exists. Read archive members directly with
ZipFile.namelist() and ZipFile.read() where possible. If extraction is
needed, use a fresh directory from tempfile.mkdtemp() or mktemp -d and
preserve existing checkouts and scratch directories. Check whether rg is
available; use grep or Python if it is absent. Read the source sections
needed for the requested command or compatibility question.
Check whether the request is executable in public v1 before installing or downloading anything. For synonymous-codon aggregation or Decodon, inspect the parser and model configuration, explain the missing feature, and finish. Do not implement the missing workflow, search private code, or keep retrying unsupported commands.
Resolve the variant CSV, checkpoint, and output directory from the request and available files. Validate inputs before inference. Execution requires the project's ML dependencies and a compatible NVIDIA GPU. If a required resource is unavailable, return the validated inputs where possible and a command with the missing prerequisite identified. When scoring is requested and resources are ready, execute and verify the score arrays. A request for preparation ends with the inputs and command. If variants are missing, report the required schema; do not invent variants or silently switch to a public dataset.
Default to the public 80M checkpoint for demonstrations:
nvidia/NV-CodonFM-Encodon-80M-v1, revision
399ca9fe17b57941a7bebc6788033919b417413c, file
NV-CodonFM-Encodon-80M-v1.safetensors and sibling config.json.
Reuse an existing checkpoint or download it when needed for the requested work.
Preserve an explicitly requested model size.
Preflight
- Confirm
src/runner.py,src/data/mutation_dataset.py, andsrc/inference/encodon.pyexist. - Accept only
encodon_80m,encodon_600m, orencodon_1bas--model_name. The public parser lists larger names, but its model configuration does not implement them. - For model execution, require a
.ckptfile, or a.safetensorsfile with siblingconfig.json. Input preparation can use a planned path. - Validate the CSV headers before starting a GPU job.
Inputs
Require these CSV columns:
id: unique row identifier.ref_seq: reference coding sequence, not genomic DNA with introns, UTR-only sequence, or protein sequence.ref_codonandalt_codon: three-nucleotide codons.codon_position: zero-based codon position relative to the CDS.
With --extract-seq, MutationDataset extracts an appropriate sequence window
from ref_seq; it does not derive or require alt_seq.
Before running, normalize sequences and codons to uppercase DNA (A/C/G/T),
require CDS lengths divisible by three, and check every row satisfies:
0 <= codon_position < len(ref_seq) / 3
ref_seq[3 * codon_position : 3 * codon_position + 3] == ref_codon
The public extractor asserts the second condition and otherwise stops the job.
Examples
Set CODONFM_DATA_PATH to the variant CSV, CODONFM_CHECKPOINT_PATH to the
checkpoint, and CODONFM_RUN_DIR to your chosen output directory:
Use the interpreter from the configured ML environment for inference. The
example uses python; substitute that environment's interpreter path if the
alias is unavailable.
python -m src.runner eval \
--exp_name variant_scoring \
--model_name encodon_80m \
--checkpoint_path "$CODONFM_CHECKPOINT_PATH" \
--data_path "$CODONFM_DATA_PATH" \
--process_item mutation_pred_mlm \
--dataset_name MutationDataset \
--task_type mutation_prediction \
--extract-seq \
--mask_mutation \
--num_nodes 1 \
--num_gpus 1 \
--num_workers 0 \
--val_batch_size 2 \
--out_dir "$CODONFM_RUN_DIR" \
--predictions_output_dir "$CODONFM_RUN_DIR/predictions"
Do not remove --mask_mutation: without it, the reference codon remains
visible at the scored position and invalidates masked-codon LLR scoring.
For preparation requests, inspect the CSV directly against the input schema
and reference-position checks above, then report the rows checked and provide
the scoring command. Extra columns are allowed; use --ref_seq_col if the
reference sequence has a different column name. These checks do not require
the ML runtime. The command above performs inference when resources are ready.
The existing --dryrun optionally builds runtime configuration and skips
execution. It requires the ML dependencies, can create the prediction directory,
and does not read the CSV or load weights. Do not use it as evidence that inputs,
checkpoint compatibility, or prediction quality have been validated.
Outputs
--predictions_output_dir receives:
ref_likelihoods_merged.npyalt_likelihoods_merged.npylikelihood_ratios_merged.npyids_merged.npy
Load the arrays with NumPy and align scores by ids_merged.npy. The reported
LLR is log p(ref_codon) - log p(alt_codon); a larger positive value means the
alternate codon is less probable in context. It does not say whether
expression goes up or down.
Reporting
Keep the final answer concise and self-contained, with the requested command or compatibility conclusion near the start. For command preparation, include each row's validation result, the complete command, all four output filenames, and the LLR definition and sign interpretation. Cite the inspected source locations for the command, outputs, and scoring semantics. State whether inference ran; report numerical scores only when execution produced them.
Boundaries
- General
mutation_predictionhandles both synonymous and missense changes. - Do not use
missense_prediction,missense_inference,MissenseDataset,mutation_pred_clm,--organism_token, or--causal; those are newer unavailable public-release features. - If a user asks specifically for synonymous-codon-aggregated missense scoring, explain that public v1 only provides the general ref/alt LLR. Do not silently substitute the two methods.
Files
7- BENCHMARK.md
18147b1abc7.3 KB - SKILL.md
18f7662a537.2 KB - agents/openai.yaml
ac3bc5e5c7236 B - evals/config.yml
d00aabc354103 B - evals/evals.json
d9e87affb23.4 KB - evals/files/encodon_checkpoint.json
698fa354141.0 KB - skill-card.md
0d3507b8fd4.4 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.