skills/ K-Dense-AI/scientific-agent-skills

pacsomatic

Prepares and launches nf-core/pacsomatic matched tumor-normal PacBio HiFi genomics workflows from unaligned BAM inputs. Supports samplesheet generation, pinned Nextflow launch artifacts, local checks, LSF/Slurm/PBS Pro/SGE launcher submission, and startup troubleshooting. Use for pacsomatic run prep

0
Installs
—
Rating
—
Success rate
6
Files scanned
Scan passedmethodology
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

6 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 f1ccb507b8e0cc93… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

pacsomatic

When to use

Use scripts/run_pacsomatic.py to prepare one matched PacBio HiFi tumor/normal pair, generate a samplesheet and reproducible launch artifacts, and launch locally or submit the Nextflow driver to a scheduler. The pipeline realigns input BAMs; this helper targets unaligned HiFi BAMs and optional PacBio .pbi indexes. Do not substitute short reads or treat a BAM filename as evidence of platform, matched identity, or methylation information.

The reviewed upstream dev commit is 24c84cb371b0339c1d65a4de9451671945e19772. GitHub had no releases or tags on 2026-10-01, despite the internal manifest saying 1.0.0. The helper pins that commit by default for nf-core/pacsomatic; it does not invent a release tag. This is a source-reviewed development workflow, not a clinically validated assay. See references/pacsomatic_guide.md for sources and scientific checks.

Workflow

  1. Obtain distinct tumor and normal BAM paths, patient ID, distinct sample IDs, output directory, and exactly one reference mode: --fasta or --genome. IDs and BAM/PBI/FASTA paths must have no whitespace. Local inputs must be nonempty regular files. Remote BAM/PBI/FASTA URIs are passed through without downloading or authenticating; use managed filesystem/cloud credentials, never embed secrets or signed URLs in generated files.
  2. Confirm PacBio HiFi read groups and sample identity from acquisition metadata; confirm that MM/ML modification tags needed for methylation have been retained. Verify reference sequence/contig compatibility for every annotation resource.
  3. Generate artifacts with --dry-run. This performs helper checks and writes files, but does not invoke the pipeline, validate BAM contents, check remote availability, resolve every pipeline parameter, or verify biological suitability. Missing runtime tools are warnings here. --dry-run cannot be combined with --run/--submit, cloning, or environment creation.
  4. Review samplesheet, generated params YAML, script, pipeline revision, profiles, and branch-specific resources/skips. Existing artifacts require explicit --overwrite; input files can never be artifact targets. config.yaml is an operator reference, not an automatically loaded configuration file.
  5. For execution, select the actual runtime with --use-current-path or an existing --conda-env. Load cluster modules before invoking the helper; --module-load only repeats those commands in the generated script. No Conda YAML is bundled; creating an environment needs --conda-env-file explicitly.
  6. Use --run only for requested execution. For HPC, distinguish the outer launcher scheduler (--executor) from Nextflow's per-task process.executor, configured by a site profile or --nextflow-config. Driver CPU/memory requests do not constrain task resources. Read references/config-and-output.md.
  7. Report artifact paths, revision, checks/warnings, run type, submission ID if present, and a concrete next QC or failure-triage step. Scheduler acceptance is not pipeline completion. Keep the output directory and work/cache state stable for --resume; scripts run with the output directory as their cwd.

Examples

Run these from the repository root. Paths and site settings are illustrative; local tests use synthetic placeholders only, not human genomic data.

python skills/pacsomatic/scripts/run_pacsomatic.py \
  --tumor-bam /data/P001_T.bam --normal-bam /data/P001_N.bam \
  --patient-id P001 --tumor-sample-id P001_T --normal-sample-id P001_N \
  --outdir /results/P001 --fasta /refs/GRCh38.fa \
  --profile apptainer --use-current-path --dry-run

After reviewing artifacts, a Slurm launch can use the same inputs plus the following options (replace --dry-run with --run):

--executor slurm --queue compute --project my_account
--cpus 2 --memory-gb 8 --walltime 48:00
--nextflow-config /configs/slurm.config --overwrite --run

Those resources are for the driver, assuming the reviewed infrastructure config sets process.executor = 'slurm' and suitable task queue/resources. The helper normalizes 48:00 to Slurm 48:00:00 (48 hours). Do not add a sanger profile unless actually using that institution's LSF infrastructure.

Custom pipeline parameters go in --params-file; infrastructure goes in --nextflow-config (-c). The helper's explicit input/outdir/reference options win over params-file values. --extra-args is tokenized and shell-quoted, but cannot override these managed inputs/configuration options. Keep paths inside external params/config files absolute because the launcher cwd is the outdir.

Verification and references

uv run skills-ref validate skills/pacsomatic
python tests/run_all.py --isolated pacsomatic

The standard-library suite checks local artifact behavior, path protections, CLI modes, runtime failures and mocked scheduler submissions. Native Nextflow checks use a tiny local workflow; they do not establish that pacsomatic's full containerized scientific pipeline succeeds on a given dataset or cluster.

Files

6
56.3 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from K-Dense-AI/scientific-agent-skills8

13c-metabolic-flux

Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional

Scan passed 0
adaptyv

Uses the Adaptyv Bio Foundry API and Python SDK to design protein characterization experiments, estimate costs, submit sequences, monitor laboratory progress, and retrieve results. Applies to Adaptyv Foundry, its target catalog, binding screening and affinity assays, thermostability, expression, flu

Scan passed 0
aeon

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorit

Scan passed 0
alphagenome

Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), scores varia

Scan passed 0
analytical-method-validation

Plans, executes, and documents validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and lig

Scan passed 0
anndata

Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

Scan passed 0
arbor

Applies Arbor Hypothesis Tree Refinement to research artifacts with repeatable evaluators, including model training, agent harnesses, data synthesis and benchmark optimization. Uses persistent hypotheses, isolated experiments, evidence propagation and held-out candidate comparison for multi-experime

Scan passed 0
arboreto

Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3. Use for transcription factor-target association ranking, compatible Dask execution, sparse expression inputs, and network stability checks.

Scan passed 0

Related methodology skillsscan passed