skills/ K-Dense-AI/scientific-agent-skills

glycoengineering

Analyzes and engineers protein glycosylation by scanning canonical N-glycosylation sequons, describing S/T-rich regions, checking curated glycan evidence, and preparing NetNGlyc, NetOGlyc and GlycoSHIELD workflows. Use for glycoprotein engineering, antibody Fc glycosylation, glycan shielding, and si

0
Installs
—
Rating
—
Success rate
5
Files scanned
Scan passedmethodology
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

5 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 7ef85ad096c70510… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Glycoengineering

When to use

Use for canonical N-glycosylation candidate analysis, O-GalNAc candidate triage, antibody glycoform comparison, or structural glycan shielding. Keep four kinds of result separate: sequence motif, prediction, experimentally supported occupancy, and glycan structure/composition. None alone establishes all the others.

This workflow focuses on eukaryotic secretory-pathway N-glycosylation and mucin-type O-GalNAc. O-GlcNAc, O-mannose, O-fucose and other O-linked modifications require appropriate evidence and predictors; S/T enrichment does not identify the type.

Workflow

  1. Record the protein accession, isoform/version, exact sequence, expression host, signal peptide/transmembrane topology and construct boundaries. Preserve a mapping from submitted sequence coordinates to mature protein, PDB chain/residue identifiers and antibody numbering where relevant.
  2. Scan canonical N-X-[S/T] candidates with X not Pro. Retain overlaps: NNST has candidates starting at 1 and 2. A missing canonical motif does not rule out unusual N-glycosylation; a present motif does not establish occupancy.
  3. Prioritize secreted/luminal/extracellular regions using topology and curated evidence. Sequence-only results in a cytoplasmic region are not evidence of secretory-pathway glycosylation.
  4. For O-GalNAc, use S/T density only as a descriptive feature, then inspect a suitable predictor and cell/transferase context. Do not exclude SP/TP motifs: adjacent Pro can be favorable, depending on the GalNAc-transferase and context. See the GalNAc-T substrate study.
  5. Propose explicitly numbered mutations, list every changed residue and rescan the full product for lost/created overlapping motifs. Evaluate structural and expression effects independently of glycosylation.
  6. Validate occupancy and glycoforms experimentally, with localization evidence, glycan composition/structure confidence, and the biological assay required by the engineering objective. For biosimilar comparisons, match sample handling, analytical coverage and quantification before comparing glycan percentages.

Local sequence analysis

The standard-library helper scripts/glycoengineering_tools.py replaces copied snippets. Import it with this skill's scripts/ directory on PYTHONPATH, or run Python from that directory. It accepts raw canonical 20-amino-acid sequences, normalizes whitespace/case, and rejects FASTA headers, gaps, unknown residues and empty input. Parse FASTA records first; do not remove unknown residues because that changes coordinates. Coordinates are 1-based in the normalized submitted sequence.

from glycoengineering_tools import (
    normalize_sequence, find_n_glycosylation_sequons,
    eliminate_glycosite, add_glycosite, find_st_rich_sites,
)

sequence = normalize_sequence("NNST")
sites = find_n_glycosylation_sequons(sequence)
assert [site["position"] for site in sites] == [1, 2]

mutant = eliminate_glycosite(sequence, 1, "Q")
assert mutant == "QNST"  # the overlapping sequon at position 2 remains
assert [site["position"] for site in find_n_glycosylation_sequons(mutant)] == [2]

# Three intended changes: A1N, P2A, A3T. P->A must be explicit.
parent = "APA"
product = add_glycosite(parent, 1, "T", allow_proline_substitution=True)
changes = [(i, before, after) for i, (before, after) in
           enumerate(zip(parent, product), 1) if before != after]
assert product == "NAT" and len(changes) == 3

# Density is a fraction of residues, not an O-glycosylation probability.
o_candidates = find_st_rich_sites("STPST", window=7, min_st_fraction=0.4)
assert [site["position"] for site in o_candidates] == [1, 2, 4, 5]

eliminate_glycosite requires a complete canonical sequon and a one-residue replacement other than N. add_glycosite can alter up to three residues; it retains an existing S/T at +2. Neither function predicts a mutant's fold, glycan occupancy, function or tolerability. An N-to-Q substitution can change protein behavior even without a glycan effect.

The S/T density window must be a positive odd integer. Terminal windows shorten, and their denominator is the actual window length. Dense regions may help triage mucin-like sequence; isolated real O-GalNAc sites can be missed.

For batch analysis, keep identifiers and normalized sequence lengths:

sequences = {"overlap": "NNST", "control": "APAA", "mucin_like": "STPST"}
rows = []
for name, raw_sequence in sequences.items():
    seq = normalize_sequence(raw_sequence)
    positions = [site["position"] for site in find_n_glycosylation_sequons(seq)]
    rows.append({"protein": name, "length": len(seq),
                 "n_sequon_positions": positions,
                 "n_sequons_per_100_residues": 100 * len(positions) / len(seq),
                 "st_rich_positions": [site["position"] for site in find_st_rich_sites(seq)]})

Prediction services

Use the official forms, which accept FASTA and return web results; this skill does not provide a stable submission API or job-polling endpoint. Saving a CGI URL is not a submitted job. Preserve the output, input and service version together.

ServiceCurrent documented scope and interpretation
NetNGlyc 1.0Human N-glycosylation; reports network potential and jury agreement. Default threshold 0.5; not a calibrated occupancy probability. Up to 2,000 sequences / 200,000 residues total / 4,000 per sequence. SignalP runs, but extracellular topology is not checked. The server may score N-P-S/T; exclude these from the canonical candidate set unless independent evidence warrants review.
NetOGlyc 4.0Mammalian mucin-type O-GalNAc; outputs GFF2 confidence scores, with scores greater than 0.5 marked positive. Up to 50 sequences / 200,000 residues total / 4,000 per sequence. Prefer full sequence including signal peptide; isolated sites need 15 residues of flanking context on both sides. A positive supports regional likelihood, not guaranteed site occupancy or glycan type beyond this model's O-GalNAc scope.

The service documentation and sample output were reviewed; new prediction jobs were not submitted during this refresh.

Structural shielding with GlycoSHIELD

GlycoSHIELD grafts pre-simulated glycan conformers onto protein structures and filters steric clashes. It does not predict which sequons are occupied or run fresh molecular dynamics for each query. See Tsai et al., 2024.

Read references/glycoshield.md for the pinned source, installation, direct Python API, input mapping and SASA analysis. The reviewed upstream CLI silently ignores several parsed options, including --mode, --threshold and --dryrun; use the documented direct API workaround. Outputs include per-site PDB/XTC ensembles and shielding encoded in a PDB B-factor column; those values are not crystallographic temperature factors.

Database evidence and glycan notation

Read references/glycan_databases.md for the live GlyTouCan SPARQL lookup, GlyConnect access limitations, curated resources and notation. Query exact accessions and preserve dataset/source dates, species, tissue/cell context and supporting publications. Missing records or service errors cannot establish that a protein is unglycosylated.

A monosaccharide composition such as Hex:5 HexNAc:4 dHex:1 does not resolve linkages, branch positions, or distinguish GlcNAc from GalNAc. A cartoon without linkages is schematic, not IUPAC condensed notation. Use an actual sequence (WURCS/GlycoCT/IUPAC with its uncertainty retained) and accession when available. Core fucose attaches to the innermost GlcNAc of an N-glycan, not to core mannose.

Antibody engineering decisions

Number Fc mutations in an explicit scheme (commonly EU), then map them to the actual construct. EU N297 is not residue 297 of an isolated Fc FASTA. Fc glycans and any Fab glycans must be distinguished analytically.

ObjectiveCandidate strategyRequired interpretation
Increase FcγRIIIa engagement / ADCCReduce Fc core fucosylation while preserving the glycanEffect size depends on antibody, receptor and assay; do not assume a universal fold gain. Structural evidence.
Remove canonical Fc N-glycosylationN297Q/A/D or disrupt the +2 residue with T299ASequon disruption; not a guarantee of otherwise unchanged structure or effector function. Verify the actual product.
Alter serum persistenceCharacterize glycoform-specific clearance with the relevant proteinIgG high-mannose clearance can increase; sialylation is not a universal IgG half-life recipe. Human PK study.
Investigate anti-inflammatory Fc activityCompare defined sialylated glycoformsLinkage, preparation and biological model matter. Positive results in particular models do not establish a universal IVIG mechanism. Defined Fc study.
Reduce non-human glycan epitopesMeasure α-Gal and Neu5Gc and select compatible production conditionsSequence editing alone does not control the host's glycan processing.
Study epitope accessibilityIntroduce or remove a mapped surface sequonConfirm occupancy, folding, binding and antigenicity; shielding calculations are geometric hypotheses.

Fc sequence variants such as S298A/E333A/K334A or F243L-containing combinations can alter both receptor interaction and host-dependent glycan processing. Do not label F243L alone as a deterministic defucosylation switch. See the 2026 Fc-variant glycan study.

Experimental interpretation and verification boundary

Glycopeptide searches can support peptide identity, composition and sometimes site localization, without resolving a full glycan structure. Review localization fragments, competing assignments, search-space choices and error control. In Byonic, composition does not determine topology; O-glycosite ambiguity may remain even with an identified glycopeptide. Relative signal among detected glycoforms is not automatically absolute site occupancy; occupancy needs the appropriate modified and unmodified denominator and analytical response considerations.

Local sequence behavior is covered by synthetic tests. Public accession lookup was executed; API error handling was mocked. GlycoSHIELD source/argument handling was checked, but full conformer grafting, GROMACS SASA, DTU jobs, commercial MS software and biological performance were not executed. Supporting references record the service/source review date and unresolved access gaps.

Files

5
33.6 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from K-Dense-AI/scientific-agent-skills8

13c-metabolic-flux

Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional

Scan passed 0
adaptyv

Uses the Adaptyv Bio Foundry API and Python SDK to design protein characterization experiments, estimate costs, submit sequences, monitor laboratory progress, and retrieve results. Applies to Adaptyv Foundry, its target catalog, binding screening and affinity assays, thermostability, expression, flu

Scan passed 0
aeon

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorit

Scan passed 0
alphagenome

Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), scores varia

Scan passed 0
analytical-method-validation

Plans, executes, and documents validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and lig

Scan passed 0
anndata

Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

Scan passed 0
arbor

Applies Arbor Hypothesis Tree Refinement to research artifacts with repeatable evaluators, including model training, agent harnesses, data synthesis and benchmark optimization. Uses persistent hypotheses, isolated experiments, evidence propagation and held-out candidate comparison for multi-experime

Scan passed 0
arboreto

Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3. Use for transcription factor-target association ranking, compatible Dask execution, sparse expression inputs, and network stability checks.

Scan passed 0

Related methodology skillsscan passed