skills/ K-Dense-AI/scientific-agent-skills

depmap

Retrieves and analyzes Cancer Dependency Map (DepMap) release data, including CRISPR Chronos gene effects, cancer model annotations, omics biomarkers, and PRISM drug sensitivity. Supports cancer-selective dependency, co-essentiality, and candidate synthetic-lethality analyses with release-aware iden

0
Installs
—
Rating
—
Success rate
3
Files scanned
Scan passedknowledge
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

3 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 4f2cc9925e1844ef… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

DepMap — Cancer Dependency Map

When to use

Use this skill to rank dependencies in cancer models, compare prespecified biomarker cohorts, inspect co-essentiality, or relate genetic dependencies to PRISM compound response. These analyses generate target and synthetic-lethality hypotheses; they do not establish clinical efficacy, a therapeutic window, or a causal interaction.

DepMap integrates Broad and Sanger CRISPR screens and hosts independent RNAi and drug-screen datasets. Match each question to its assay and pinned release. RNAi DEMETER2 is a different perturbation modality, not an older CRISPR scoring method.

1. Discover a release, then obtain its files

Use DepMap Downloads and the selected release's README and release notes. On review, the live catalogue's latest DepMap Public release was 26Q1, dated 2026-04-01. Rediscover before new work; a release name or expected quarter is not evidence a release is available.

Verified public endpoint

GET https://depmap.org/portal/api/no-captcha/download/files returns a complete CSV catalogue, with no request body, API key, or pagination parameters. Required metadata columns are release, release_date, filename, and md5_hash. The live response also has url: older files can have links, while recent release rows have blank links. All 85 entries for 26Q1 had blank url values at review. Treat availability per row, not per endpoint.

The staff announcement introduced this metadata route in July 2026. The older https://depmap.org/portal/api/download/files catalogue can supply download links but returned an HTML verification page during review, including with HTTP 200. Validate content before parsing. Obtain current files through the portal when verification is required. Refresh expiring signed links immediately before use; do not fabricate storage URLs or assume every new release is on Figshare.

No supported contract was verified for the former /api/gene?gene_id=... or /api/data/gene_dependency examples. Use local release matrices for gene slices. Do not assume a Python package named depmap is an official portal client.

From the skill root, use the bundled helpers:

from scripts.depmap_data import fetch_catalogue, select_file, verify_md5

catalogue = fetch_catalogue()
public = catalogue[catalogue["release"].str.startswith("DepMap Public ")]
print(public[["release", "release_date"]].drop_duplicates()
      .sort_values("release_date", ascending=False).head(12))

# Explicit example pin; inspect discovery output before selecting your release.
release = "DepMap Public 26Q1"
record = select_file(catalogue, release, "CRISPRGeneEffect.csv")
print(record[["release", "filename", "md5_hash"]])
# After obtaining this file from the portal:
# verify_md5("CRISPRGeneEffect.csv", record["md5_hash"])

Save the selected metadata, retrieval date, release citation, local checksum, and README alongside analysis outputs. Record source dataset terms separately from this skill's license; newer release terms differ from older Figshare releases. For a missing published checksum, retain a local SHA-256 and explicitly state that it was not compared with a publisher checksum.

2. Inspect file contracts before joining

The following names were verified in the 26Q1 catalogue. Check that release's README for units and identifier level before treating a file as a numeric matrix.

FileUse and important contract
CRISPRGeneEffect.csvIntegrated, model-level Chronos effects; rows are ModelID, columns preserve Symbol (EntrezID)
CRISPRGeneEffectUncorrected.csvUncorrected effects; not interchangeable with corrected effects
CRISPRGeneDependency.csvRelease-specific dependency statistics; confirm probability/FDR definition and direction in the README
Model.csvModel metadata keyed by ModelID
ModelCondition.csvGrowth/treatment conditions keyed by ModelConditionID, mapping to ModelID
ScreenGeneEffect.csv, ScreenSequenceMap.csv, CRISPRScreenMap.csvScreen-level effects and mappings; multiple screens can belong to one model
OmicsExpressionTPMLogp1HumanProteinCodingGenes.csvExpression output with sequencing/model metadata and default-entry flags; exclude metadata columns from expression calculations
OmicsSomaticMutationsMatrixDamaging.csv, OmicsSomaticMutationsMatrixHotspot.csvDistinct mutation annotations; damaging and hotspot calls answer different biological questions
OmicsCNGeneWGS.csv, OmicsCNGeneMC_WES.csvSeparate WGS/WES copy-number products; do not concatenate as one uniformly processed cohort
PortalOmicsCNGeneLog2.csvTransformed portal copy-number product; establish the exact transform before inversion
OmicsProfiles.csv, Gene.csvOmics provenance/mappings and gene annotations
AchillesScreenQCReport.csv, AchillesSequenceQCReport.csvScreen/sequence QC; apply documented eligibility rules rather than invented universal cutoffs

Current model annotations include CellLineName, OncotreeLineage, OncotreePrimaryDisease, and OncotreeSubtype. Legacy sample_info.csv fields (DepMap_ID, lineage, primary_disease) require an explicit conversion. Model.csv includes models without CRISPR measurements: absence from the effect matrix is not a nondependency score.

Omics may be indexed by SequencingID, ModelConditionID, or ModelID. For a basal model analysis, use the release's IsDefaultEntryForModel == Yes flag, then require one row per ModelID. The condition-level flag IsDefaultEntryForMC answers a different question. Subset the assay/datatype before selecting defaults from a mapping table. Never average repeated conditions or count them as independent models without an explicit scientific design.

See the mapping guide source and metadata definitions.

3. Interpret Chronos and inspect a target

Chronos gene effect is continuous and unbounded. More negative values indicate stronger loss of fitness. In the normalized release matrix, approximately 0 is the nonessential-control anchor and −1 is the common-essential-control anchor; −1 is not a dependency boundary. A filter such as ≤ −0.5 is exploratory, not a p-value, FDR, or probability cutoff. Positive effects may reflect growth advantage or technical noise and need validation.

For binary calls, inspect the selected file's statistic and direction. A high dependency probability and a low FDR are different rules. Do not silently convert between them. Use release-matched positive/negative control lists and QC outputs. See score and QC caveats.

The following local-data examples are illustrative until run against your chosen files. Helper behavior is tested using synthetic fixtures; current bulk matrices were not downloaded during this review.

from scripts.depmap_data import load_gene_effect, load_models, gene_profile

effects = load_gene_effect("CRISPRGeneEffect.csv")
models = load_models("Model.csv")
profile = gene_profile(effects, models, "KRAS (3845)")
print(profile[["CellLineName", "OncotreeLineage", "gene_effect"]].head(20))

# Inspect actual annotation values, then choose an exact cohort.
print(profile["OncotreeLineage"].value_counts())
known = profile.dropna(subset=["OncotreeLineage"])
lung = known.loc[known["OncotreeLineage"].eq("Lung"), "gene_effect"]
other = known.loc[~known["OncotreeLineage"].eq("Lung"), "gene_effect"]
if lung.empty or other.empty:
    raise ValueError("Both cohorts require observed effects and known lineage")
print({"n_lung": len(lung), "n_other": len(other),
       "other_minus_lung_mean": other.mean() - lung.mean(),
       "lung_fraction_below_exploratory_cutoff": (lung <= -0.5).mean()})

This is a descriptive comparison against other cancer models, not against normal tissue. Keep the denominator and missingness counts. Preserve complete gene labels; a symbol can be ambiguous, and the helper refuses ambiguous symbol resolution.

4. Test biomarker associations

  1. Define the alteration before inspecting target scores. A damaging-call matrix does not establish biallelic loss; activating KRAS hotspots require a different definition from generic damaging mutations.
  2. Map profiles to the correct model/condition. Use 0 only for an assayed negative, 1 for the prespecified biomarker, and missing for unknown status.
  3. Restrict to scientifically comparable models (lineage, culture conditions, screen source and related-patient structure). Inspect confounding before testing.
  4. Test every eligible gene in the planned family, then adjust all its p-values before selecting hits. Report group sizes, effect sizes, p-values, and q-values.
  5. Validate candidate synthetic lethality with matched/isogenic perturbations and rescue or orthogonal evidence; association alone is insufficient.
import pandas as pd
from scripts.depmap_data import biomarker_scan

# User-prepared, release-matched annotation after the mapping and curation above.
status = pd.read_csv("curated_biomarker_status.csv", index_col="ModelID")["status"]
results = biomarker_scan(effects, status, min_n=5)
# Positive effect_size means stronger dependency in biomarker-positive models.
candidates = results.loc[(results["qval"] < 0.1) & (results["effect_size"] > 0)]

This one-sided Mann–Whitney screen is exploratory and does not adjust for covariates. min_n=5 is an explicit example floor, not a power guarantee. For inference across lineages, fit an appropriate adjusted model and inspect effect stability within lineages; do not report the helper's q-value as confounder-adjusted.

5. Co-essentiality and drug sensitivity

from scripts.depmap_data import coessentiality

correlates = coessentiality(effects, "KRAS (3845)", min_n=500)
print(correlates[["gene", "n", "r", "pval", "qval"]].head(20))

min_n is configurable and should be prespecified for the analysis. Sparse genes can dominate rankings with spurious correlations; always retain the pairwise sample count. The helper uses pairwise complete observations and skips constant genes. Pearson/BH output still needs lineage, library, and screen-quality checks; co-essentiality is not proof of a physical interaction or shared pathway. The DepMap team's discussion documents this specific sparse-coverage failure mode.

For PRISM, first select the exact screen/release and endpoint. Treatment-info CSVs are annotations, not sensitivity values. The original primary screen's primary-screen-replicate-collapsed-logfold-change.csv has response rows keyed by cell-line row_name and columns keyed by treatment column_name. Join the corresponding cell-line and treatment-info files; retain compound, dose, and screen identity. Lower log2 fold change indicates greater loss of viability. Secondary-screen AUC comes from dose-response curve parameters and is not the same quantity as single-dose log fold change. See the PRISM workflow before loading these files.

Validation and sources

Before reporting: verify release/file identity, unique joins, identifier level, complete gene labels, cohort membership, missingness, screen eligibility, score units/direction, multiple-testing family, and whether evidence is observational. Expression, copy number, and CRISPR can share technical confounders; low expression does not guarantee a measured score of zero. Broad essentiality motivates normal cell/selectivity studies; it does not by itself prove a target is undruggable.

The live no-captcha catalogue and original PRISM READMEs were fetched successfully. The helper uses Python urllib, which succeeded at review; the same public endpoint returned HTTP 403 with a default requests client. Surface access failures instead of treating an error page as data. Protected portal endpoints returned verification pages; authenticated downloads and current-matrix end-to-end analysis remain unverified. This review makes no claim of a working bearer-token API or undocumented gene-query endpoints.

Files

3
32.5 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from K-Dense-AI/scientific-agent-skills8

13c-metabolic-flux

Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional

Scan passed 0
adaptyv

Uses the Adaptyv Bio Foundry API and Python SDK to design protein characterization experiments, estimate costs, submit sequences, monitor laboratory progress, and retrieve results. Applies to Adaptyv Foundry, its target catalog, binding screening and affinity assays, thermostability, expression, flu

Scan passed 0
aeon

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorit

Scan passed 0
alphagenome

Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), scores varia

Scan passed 0
analytical-method-validation

Plans, executes, and documents validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and lig

Scan passed 0
anndata

Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

Scan passed 0
arbor

Applies Arbor Hypothesis Tree Refinement to research artifacts with repeatable evaluators, including model training, agent harnesses, data synthesis and benchmark optimization. Uses persistent hypotheses, isolated experiments, evidence propagation and held-out candidate comparison for multi-experime

Scan passed 0
arboreto

Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3. Use for transcription factor-target association ranking, compatible Dask execution, sparse expression inputs, and network stability checks.

Scan passed 0

Related knowledge skillsscan passed