Mmcp.market

genomic-intelligence skill

by K-Dense-AI·K-Dense-AI/scientific-agent-skills·47k stars·MIT

Predict regulatory features, gene structure, and expression directly from DNA sequence using Genomic Intelligence's hosted transformer DNA language models — no local GPU or model weights. Six tasks over a REST API and a hosted MCP server (keyless public demo): promoter regions, splice donor/acceptor sites, enhancer activity, chromatin state, sequence-to-expression (log TPM), and de-novo gene annotation, plus a composite find-genes-then-predict-expression workflow. Use when the user has a gene symbol, a genomic region, or a DNA/FASTA sequence and wants any of these predictions, mentions Genomic Intelligence, genomicintelligence.ai, api.genomicintelligence.ai, or mcp.genomicintelligence.ai.

A100/100content scan

Is the genomic-intelligence skill safe?

Clean: nothing in its files matched our rules. We read 5 files in the folder on 2026-09-28.

No findings.

Install the genomic-intelligence skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git /tmp/scientific-agent-skills
mkdir -p ~/.claude/skills
cp -r /tmp/scientific-agent-skills/skills/genomic-intelligence ~/.claude/skills/genomic-intelligence
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Genomic Intelligence — DNA Sequence Models

Genomic Intelligence (GI) serves transformer DNA language models over six sequence-analysis tasks on managed GPUs. Give it a gene symbol, a genomic region, or a DNA/FASTA sequence; it returns structured predictions — promoter regions, splice sites, enhancer activity, chromatin state, expression (log TPM), and de-novo gene annotation. Nothing runs locally: no model weights, no GPU, no heavy Python stack. It is a thin client over a hosted, versioned inference API.

Official docs: docs.genomicintelligence.ai · REST contract at api.genomicintelligence.ai/v1/openapi.json · hosted MCP server at https://mcp.genomicintelligence.ai/mcp

When to use this skill

Use GI when the user has DNA and wants a model prediction:

  • Find promoters in a genomic region (promoter)
  • Predict splice donor/acceptor sites (splice)
  • Score enhancer activity — developmental & housekeeping (enhancer)
  • Annotate chromatin state across hundreds of tracks (chromatin)
  • Predict expression as log(TPM+1) from a sequence + cell-type context (expression)
  • Annotate genes/transcripts de novo, no reference needed (annotation)
  • Find the genes in a region and predict each one's expression (composite)

Not for local alignment, variant calling, or file I/O — use a local tool (BioPython, bcftools) for those. GI is for model inference from sequence.

Research and development use. Not for clinical or diagnostic decisions.

Two ways to call GI

Hosted MCP server (keyless; preferred on MCP hosts)

GI hosts an MCP server at https://mcp.genomicintelligence.ai/mcp (Streamable HTTP). When your agent host supports MCP, prefer it: it works keyless against a rate- and concurrency-limited public demo tier, and an optional gi bearer key raises those limits. It exposes acquisition tools that return a sequence handle (sequenceref) and predict_* tools that take that handle, so large sequences stay out of the context. See MCP workflow below and references/mcp.md.

REST API (universal)

Plain HTTP with requests against https://api.genomicintelligence.ai/v1. The REST path requires a GIAPIKEY (a gi_ bearer). Use it on any host, in scripts, or when you need the raw envelope. See Core REST workflow.

Access and authentication

Request one at contact@genomicintelligence.ai.

  1. The hosted MCP demo is keyless — try it with nothing set.
  2. The REST /v1 API needs a key, sent as Authorization: Bearer .

(or a .env via python-dotenv). Never commit keys.

  1. Never hardcode the key. Read it from the GIAPIKEY environment variable
export GI_API_KEY="gi_yourkeyhere"     # optional for MCP; required for REST
export GI_BASE_URL="https://api.genomicintelligence.ai"   # override for staging

Keys are scoped to a partner tier with concurrency and per-minute caps. A 429 means you hit a cap — back off and retry, or ask GI to raise your tier.

The six tasks

Each task is its own published operation with its own request schema, its own minimum length, and its own closed options object — POST /v1/tasks/promoter/predict, /v1/tasks/splice/predict, /v1/tasks/enhancer/predict, /v1/tasks/chromatin/predict, /v1/tasks/annotation/predict, /v1/tasks/expression/predict. Each path is a literal string, so nothing needs to be constructed, and there is no shared PredictRequest schema. Body is {sequence, sequence_name?, model?, options?}, returning a {data, meta} envelope. What differs per task:

Recommended mode is guidance, not a constraint — every task accepts both. Omit Prefer for a synchronous 200; send Prefer: respond-async for a 202 plus GET /v1/tasks/jobs/{jobid}. The one enforced limit is per operation: where /v1/openapi.json publishes x-sync-limit-bp on a POST, a synchronous request above that length is 413 synctoo_large — 200,000 bp on annotation and 50,000 bp on the composite workflow as of info.version 2026.09.10.1. Read the field rather than memorising the numbers; the other predict tasks carry no limit today.

The minimum is admission control, not regime. A request above the floor but shorter than the selected model's biospec.contextwindowbp is accepted and scored — against a window padded out to the context window. Enhancer is the sharp case: the floor is 50 bp but the context window is 249 bp, so 50–248 bp is scored mostly on padding. Compare your length against contextwindow_bp from GET /v1/tasks/{task}/models to know whether the model saw real sequence. Longer-than-context input is fine — the scanner steps a prediction window at a time and pads only the final partial window.

Under the floor and over the 500,000 bp cap are both 422 validationfailed at loc ["body","sequence"]; over-length is not* a 413. All lengths are measured after whitespace is stripped, so a line-wrapped FASTA body can be pasted verbatim (a > header line still fails the alphabet check).

options is typed and closed (additionalProperties: false) per task — an unknown key is a hard 422 validationfailed with type: "extraforbidden", never ignored:

Prefer: respond-async is a declared header on all six predict operations and on the composite, not just annotation — see Async.

Omit model and the API uses the task's default — that is the recommended call. Default model IDs are intentionally not documented here: defaults change and retired IDs fail hard, so never hardcode one. To pin a model, or to pick a non-human one (Drosophila, yeast, and Arabidopsis models exist for several tasks), discover IDs at call time with GET /v1/tasks/{task}/models (REST) or listmodels (MCP) — and never invent one**. Full per-task output shapes are in references/tasks.md.

expression is the strictest of the six: alone among them its schema requires options as well as sequence. Three hard rules it enforces — every violation is a 422, nothing is padded or clamped, and there is no opt-out flag, header, or query parameter:

sequence[tssindex-4599 : tssindex+4599]. The endpoint itself accepts 9,198–500,000 bp; anything below 9,198 bp is rejected outright.

  • It always scores exactly one 9,198 bp TSS-centred window —

0-based TSS offset into the whitespace-stripped sequence, bounded by 4599 ≤ tssindex ≤ len(sequence) − 4599. At exactly 9,198 bp it defaults to 4,599, the only legal value there. So you may submit a whole locus (up to 500 kb) and let the server cut the window — but the server does not discover the TSS for you (that is the composite workflow's job), and does not** reverse-complement: submit gene-sense sequence.

  • tssindex is required unless the sequence is exactly 9,198 bp.** It is the

is required, and is the only key expression accepts inside options. Unknown top-level body fields are rejected too.

  • options.description — a cell-type / assay string (e.g. "K562 cells") —

Note: the legal tss_index range is wide, so an offset that is merely

wrong (counted over raw FASTA characters including newlines, or relative to

a locus start rather than the submitted slice) does not error — it returns a

confident 200 for the wrong window. Assert on

meta.taskspecificcounts.scoredwindow / .tssindex in the response.

The length you submitted is meta.sequence_length (also echoed as

data.input.submittedsequencelength); the scored width is always 9,198,

i.e. scoredwindow[1] - scoredwindow[0]. (data.input.sequence_length

was removed at contract revision 13.)

Both tss_index violations — "required unless exactly 9,198 bp" and the range

check — come from a whole-model validator, so they surface at the body level

rather than under tssindex. Match on error.code == "validationfailed"

and use the message for display only. Any loc tuple quoted in this skill is

illustrative of that shape, not part of the contract: it is not published in

the schema and must not be branched on.

Sequence acquisition

You rarely start from a raw 9,198 bp string. Acquire sequence first:

coordinates → fetchregion(region=...). Both fetch public Ensembl reference sequence (no key). REST users can query Ensembl REST directly. (find_genes is the annotation task, not an acquisition tool.)

  • From a gene symbol → MCP fetchensemblsequence(gene=...); **from

9,198 bp. MCP: fetchgeneforexpression (handles the centring). Otherwise fetch a wider locus and pass the TSS as tssindex so the server cuts the window — but compute that offset on the stripped nucleotide string, not on file characters.

  • For expression → use the TSS-centred fetch so the window is exactly

for REST. (loadlocalfasta exists only in local deployments, not on the hosted server.)

  • From a local FASTA → MCP storeinlinesequence, or read the file yourself

for a keyless smoke test; name is required.

  • A demo sequence → MCP loaddemosequence(name=...) returns a ready handle

More skills from K-Dense-AI/scientific-agent-skills

  • AadaptyvHow to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
  • AaeonThis skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
  • AalphagenomeLook up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), score variants or scan windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and build Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.
  • Aanalytical-method-validationPlan, execute, and document validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and ligand binding or cell-based assays whenever the question is whether a procedure is fit for its intended purpose. Triggers include "method validation", "analytical method validation", "AMV", "validation protocol", "acceptance criteria", "linearity", "reportable range", "accuracy and precision", "repeatability", "intermediate precision", "recovery", "LOD", "LOQ", "detection limit", "quantitation limit", "specificity", "robustness", "method transfer", "method comparison", "Deming", "Passing-Bablok", "Bland-Altman", "equivalence testing", "OOS investigation", "ICH Q2", "Q2(R2)", "Q14", "USP 1225", "ICH M10", "incurred sample reanalysis", "ISR", "CLSI EP", and any request to show that an assay works.
  • AanndataData structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
  • AarborAutonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
  • AarboretoInfer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
  • AastropyCore Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.
  • AautoskillObserve the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.
  • Abenchling-integrationBenchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.
  • Abgpt-paper-searchSearch scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.
  • AbidsUse this skill when working with Brain Imaging Data Structure (BIDS) datasets: organizing neuroscience and biomedical data (MRI, EEG, MEG, iEEG, PET, microscopy, NIRS, motion capture, EMG, MR spectroscopy, behavioral), querying BIDS layouts, validating compliance, converting DICOM to BIDS, writing metadata sidecars, or creating BIDS derivatives.

All agent skills → · MCP servers