pacsomatic
Prepares and launches nf-core/pacsomatic matched tumor-normal PacBio HiFi genomics workflows from unaligned BAM inputs. Supports samplesheet generation, pinned Nextflow launch artifacts, local checks, LSF/Slurm/PBS Pro/SGE launcher submission, and startup troubleshooting. Use for pacsomatic run prep
- 0
- Installs
- —
- Rating
- —
- Success rate
- 6
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 f1ccb507b8e0cc93… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
pacsomatic
When to use
Use scripts/run_pacsomatic.py to prepare one matched
PacBio HiFi tumor/normal pair, generate a samplesheet and reproducible launch
artifacts, and launch locally or submit the Nextflow driver to a scheduler.
The pipeline realigns input BAMs; this helper targets unaligned HiFi BAMs and
optional PacBio .pbi indexes. Do not substitute short reads or treat a BAM
filename as evidence of platform, matched identity, or methylation information.
The reviewed upstream dev commit is
24c84cb371b0339c1d65a4de9451671945e19772. GitHub had no releases or tags on
2026-10-01, despite the internal manifest saying 1.0.0. The helper pins that
commit by default for nf-core/pacsomatic; it does not invent a release tag.
This is a source-reviewed development workflow, not a clinically validated assay.
See references/pacsomatic_guide.md for sources
and scientific checks.
Workflow
- Obtain distinct tumor and normal BAM paths, patient ID, distinct sample IDs,
output directory, and exactly one reference mode:
--fastaor--genome. IDs and BAM/PBI/FASTA paths must have no whitespace. Local inputs must be nonempty regular files. Remote BAM/PBI/FASTA URIs are passed through without downloading or authenticating; use managed filesystem/cloud credentials, never embed secrets or signed URLs in generated files. - Confirm PacBio HiFi read groups and sample identity from acquisition metadata; confirm that MM/ML modification tags needed for methylation have been retained. Verify reference sequence/contig compatibility for every annotation resource.
- Generate artifacts with
--dry-run. This performs helper checks and writes files, but does not invoke the pipeline, validate BAM contents, check remote availability, resolve every pipeline parameter, or verify biological suitability. Missing runtime tools are warnings here.--dry-runcannot be combined with--run/--submit, cloning, or environment creation. - Review samplesheet, generated params YAML, script, pipeline revision, profiles,
and branch-specific resources/skips. Existing artifacts require explicit
--overwrite; input files can never be artifact targets.config.yamlis an operator reference, not an automatically loaded configuration file. - For execution, select the actual runtime with
--use-current-pathor an existing--conda-env. Load cluster modules before invoking the helper;--module-loadonly repeats those commands in the generated script. No Conda YAML is bundled; creating an environment needs--conda-env-fileexplicitly. - Use
--runonly for requested execution. For HPC, distinguish the outer launcher scheduler (--executor) from Nextflow's per-taskprocess.executor, configured by a site profile or--nextflow-config. Driver CPU/memory requests do not constrain task resources. Read references/config-and-output.md. - Report artifact paths, revision, checks/warnings, run type, submission ID if
present, and a concrete next QC or failure-triage step. Scheduler acceptance
is not pipeline completion. Keep the output directory and work/cache state
stable for
--resume; scripts run with the output directory as their cwd.
Examples
Run these from the repository root. Paths and site settings are illustrative; local tests use synthetic placeholders only, not human genomic data.
python skills/pacsomatic/scripts/run_pacsomatic.py \
--tumor-bam /data/P001_T.bam --normal-bam /data/P001_N.bam \
--patient-id P001 --tumor-sample-id P001_T --normal-sample-id P001_N \
--outdir /results/P001 --fasta /refs/GRCh38.fa \
--profile apptainer --use-current-path --dry-run
After reviewing artifacts, a Slurm launch can use the same inputs plus the
following options (replace --dry-run with --run):
--executor slurm --queue compute --project my_account
--cpus 2 --memory-gb 8 --walltime 48:00
--nextflow-config /configs/slurm.config --overwrite --run
Those resources are for the driver, assuming the reviewed infrastructure config
sets process.executor = 'slurm' and suitable task queue/resources. The helper
normalizes 48:00 to Slurm 48:00:00 (48 hours). Do not add a sanger profile
unless actually using that institution's LSF infrastructure.
Custom pipeline parameters go in --params-file; infrastructure goes in
--nextflow-config (-c). The helper's explicit input/outdir/reference options
win over params-file values. --extra-args is tokenized and shell-quoted, but
cannot override these managed inputs/configuration options. Keep paths inside
external params/config files absolute because the launcher cwd is the outdir.
Verification and references
uv run skills-ref validate skills/pacsomatic
python tests/run_all.py --isolated pacsomatic
The standard-library suite checks local artifact behavior, path protections, CLI modes, runtime failures and mocked scheduler submissions. Native Nextflow checks use a tiny local workflow; they do not establish that pacsomatic's full containerized scientific pipeline succeeds on a given dataset or cluster.
Files
6- SKILL.md
54253193876.3 KB - config.yaml
c8a88b938f1.0 KB - references/agent-playbook.md
e57f8d5b5f2.0 KB - references/config-and-output.md
6a8e29b8b15.8 KB - references/pacsomatic_guide.md
1e274439b75.2 KB - scripts/run_pacsomatic.py
853ad3170335.9 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from K-Dense-AI/scientific-agent-skills8
Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional
Uses the Adaptyv Bio Foundry API and Python SDK to design protein characterization experiments, estimate costs, submit sequences, monitor laboratory progress, and retrieve results. Applies to Adaptyv Foundry, its target catalog, binding screening and affinity assays, thermostability, expression, flu
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorit
Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), scores varia
Plans, executes, and documents validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and lig
Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
Applies Arbor Hypothesis Tree Refinement to research artifacts with repeatable evaluators, including model training, agent harnesses, data synthesis and benchmark optimization. Uses persistent hypotheses, isolated experiments, evidence propagation and held-out candidate comparison for multi-experime
Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3. Use for transcription factor-target association ranking, compatible Dask execution, sparse expression inputs, and network stability checks.