ai.genomicintelligence/genomic-intelligence

Genomic Intelligence

Hosted DNA language models: promoter, splice, enhancer, chromatin, expression, annotation

1.0.0
Version
remote
Transport
15
Tools

Security review

Review passed

Reviewed Jan 1, 2000.

  • tools: 15 tools scanned
  • metadata: scanned

No findings.

Tools (15)

  • list_models

    List available models for a task. Use to discover model ids before passing one as the `model` argument to a predict tool. The same catalog is also available as the resource `gi://models`. Returns a FLAT object — {task, default_model, models: [...]} — not the {data, meta} envelope the predict tools return. Each model carries a `bio_spec`, whose useful fields are `request_max_bp` (the enforced ceiling, 500,000 everywhere) and `context_window_bp` (what the model reads in one step — compare your sequence length against it: a shorter one is scored against a padded window). `trained_window_bp` is the fixed receptive field where there is no sliding window (null for a model without one). For expression, `recommended_flank_bp` is how many bp the model reads on each side of the TSS; longer TSS-centred input is better, and fetch_gene_for_expression already fetches the largest listed v

  • fetch_ensembl_sequence

    Fetch a gene's reference sequence from Ensembl and store it. Returns a handle ({ref, name, length, preview, ...}). Pass the `ref` to predict_* tools — the bases stay server-side. For expression, use fetch_gene_for_expression instead (it prepares the TSS-centred window that model needs).

  • fetch_region

    Fetch a genomic region by coordinates from Ensembl and store it. For "find the genes in chr8:127,680,000-127,800,000"-style requests: resolves a coordinate range to reference sequence and returns a handle ({ref, name, length, ...}) to pass to find_genes / predict_* — the bases stay server-side. Plus strand by default, which is what the gene-finder expects. For a gene by name use fetch_ensembl_sequence; for expression use fetch_gene_for_expression.

  • fetch_gene_for_expression

    Fetch a gene's sequence prepared for expression prediction. Resolves the gene's canonical-transcript TSS via Ensembl and stores a gene-sense sequence centred on it, as a handle to pass to predict_expression(sequence_ref=...). Longer TSS-centred sequence gives better predictions, so this already fetches as much flank on each side of the TSS as the listed expression models recommend (the largest bio_spec.recommended_flank_bp from list_models). One handle serves every expression model: the API reads what the chosen model needs around the TSS and ignores the rest. The handle records `tss_index` (the TSS offset into it), and predict_expression uses it when you do not pass one, so no offset arithmetic is needed. Every expression model accepts at least 9,198 bp with the TSS at least 4,599 bp from each end. Near a chromosome end the flank is shortened to what fits (reported as `fla

  • load_demo_sequence

    Load a bundled demo reference sequence and return a handle. The server ships one curated, task-correct positive control per task (list them via the gi://sequences resource) — e.g. `expression_hbb_k562` is a ready-to-use K562 expression window for predict_expression. Stores the demo and returns a handle to pass to a predict_* tool: no Ensembl fetch, no quota. Handy for smoke-testing a prediction end-to-end.

  • store_inline_sequence

    Store a human-pasted sequence and return a handle to re-use it. For a sequence you've already pasted into the conversation, this gives back a short handle so you can run several tasks on it without re-pasting the bases in each predict_* call. Note that the full sequence still passes through the LLM on THIS call — it does not save context on its own. For large sequences, prefer fetch_ensembl_sequence / fetch_gene_for_expression / load_local_fasta, which acquire the bases server-side and never round-trip them. A line-wrapped FASTA *body* may be pasted verbatim: whitespace is stripped before storing, so the handle's `length` counts bases and a later `tss_index` counts into the same string the API measures. (A FASTA `>` header line is not a sequence and is rejected by the API's alphabet check.)

  • predict_promoter

    Predict promoter regions (G0). 300–500,000 bp. Returns the {data, meta} envelope: data.regions lists predicted promoters with start/end/score. 300 bp is the task floor for every promoter model. The default g0-promoter-2000bp scans a 2,000 bp context window, so a shorter (but ≥300 bp) sequence is still scored — against a window padded out to that size. Check the chosen model's bio_spec.context_window_bp via list_models to know whether it saw real sequence or padding.

  • predict_splice

    Predict splice donor/acceptor sites (G0 BigBird). 100–500,000 bp. The model reads a 15,000 bp context window, so anything shorter is scored against a padded window — feed a whole transcript locus when you can. It is also strand-specific, and the wrong strand fails silently and plausibly — it returns sites at different positions, often still scoring above 0.9, not the near-zero scores once documented here. Nothing in the response flags it, so submit the transcript's own orientation (fetch_region takes `strand`).

  • predict_enhancer

    Predict enhancer activity (G0 DeepSTARR). 50–500,000 bp. 50 bp is the task's admission floor (the API 422s below it), not a statement about what the model reads: enhancer models score a 249 bp context window, so 50–248 bp is accepted and scored against a padded window. For a meaningful call, submit at least the 249 bp context.

  • predict_chromatin

    Chromatin annotation across 919 features (G0 DeepSEA). 200–500,000 bp. The model reads a 1,000 bp context window; 200–999 bp is accepted and scored against a padded window.

  • predict_expression

    Predict a gene's expression from the sequence around its TSS. Expression is cell-type-specific, so `description` (cell type / assay context, e.g. 'K562 cell line') is REQUIRED — the API rejects requests without it. The input rule is the same for every expression model: at least 9,198 bp, with the TSS at least 4,599 bp from each end, up to 500,000 bp. Longer TSS-centred sequence gives better predictions: each model's bio_spec.recommended_flank_bp (see list_models) says how many bp it reads on each side of the TSS, and flank beyond that is ignored. Less flank, down to 4,599 bp, is accepted and scored, but the model then reads less than it was trained on. Omit `model` for the server's default. - A 9,198 bp sequence with the TSS at offset 4,599 (the midpoint) needs no `tss_index`. - Any longer sequence needs `tss_index`: the 0-based offset of the TSS into it (a fet

  • find_genes

    Find genes (transcript intervals) in a genomic region (async, ~8-25s). Takes 1,000–500,000 bp. The floor is the strictest of the scanning tasks: gene finding needs a region, not a site. (Only expression's 9,198 bp is higher, and that is the sequence around one TSS rather than a region to search.) Gene-finding: detects transcript boundaries (TSS + PolyA) and returns one interval per predicted transcript — start/end, strand, a confidence score, and predicted TSS/PolyA positions (BED-style feature intervals, not free-text notes). Use this for "what genes are here", "find / locate genes", or "annotate this region". Each transcript also carries its type (mRNA/lnc_RNA) and internal exon/intron/CDS structure in `exons`/`introns`/`cds` arrays, plus a browser-ready GFF3 track in `data.formats.gff3`. To get each gene's *expression* from a raw region, use find_genes_and_predict_expression

  • find_genes_and_predict_expression

    Find genes in a sequence, then predict each gene's expression (composite). Server-side chaining in ONE call: finds genes (transcript intervals, with their TSS) in the sequence, then predicts expression off each discovered TSS in the given experimental context. This is the right tool whenever you want expression for a raw region or sequence — e.g. "find the genes in chr8:… and predict their expression in K562". predict_expression scores the sequence around ONE TSS and needs you to know where that TSS is (a 9,198 bp window with the TSS at its midpoint, or a longer sequence plus `tss_index`); this tool discovers every gene's TSS itself. It has no 9,198 bp floor and no tss_index; it starts with gene finding, so it takes 1,000–500,000 bp. Runs async internally at every size (the annotate stage is slow even for small inputs), so progress always streams. With wait=True (default), blocks a

  • get_job

    Poll an async job once. Returns the {data, meta} result if complete, a progress envelope if still running, or an error envelope if it failed.

  • list_jobs

    List the caller's recent async jobs (also available as gi://jobs/recent).