Mmcp.market

imaging-data-commons skill

by K-Dense-AI·K-Dense-AI/scientific-agent-skills·47k stars·MIT

Query and download public cancer imaging data from NCI Imaging Data Commons. Invoke for any question about IDC collections, cancer imaging datasets, DICOM data access, radiology (CT, MR, PET) or pathology AI training sets, metadata queries, visualization, or license checks — even when the user doesn't explicitly mention "IDC". No authentication required.

A100/100content scan

Is the imaging-data-commons skill safe?

Clean: nothing in its files matched our rules. We read 15 files in the folder on 2026-09-28.

No findings.

Install the imaging-data-commons skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git /tmp/scientific-agent-skills
mkdir -p ~/.claude/skills
cp -r /tmp/scientific-agent-skills/skills/imaging-data-commons ~/.claude/skills/imaging-data-commons
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Imaging Data Commons

Overview

Query and download public cancer imaging data from the National Cancer Institute Imaging Data Commons (IDC). No authentication required for data access.

Expected network access: IDC metadata is reachable three ways — a local DuckDB index shipped with the idc-index Python package (no network), or the hosted IDC service over MCP or REST (api.imaging.datacommons.cancer.gov, no authentication). File downloads use public GCS (storage.googleapis.com) and AWS S3 (s3.amazonaws.com) — no authentication required. DICOMweb access uses either the public IDC proxy (proxy.imaging.datacommons.cancer.gov, no auth) or the Google Cloud Healthcare API (healthcare.googleapis.com, requires GCP authentication). Optional BigQuery queries (bigquery.googleapis.com) also require GCP authentication. No credentials or environment variables are accessed by this skill.

Current IDC Data Version: v24 (always verify — see Best Practices)

Choose the access path first. There is no single default: the cheapest correct path depends on the session and the task.

MCP Server*.

  1. Session already has the IDC MCP server? Route discovery and metadata there — see *IDC

use idc-index for everything.

  1. Otherwise, is idc-index installed? Run python scripts/check_version.py. If it passes,

lookups, SQL under 10 000 rows, licenses, citations, viewer URLs? Use the REST API over curl; do not install anything. Installing costs ~77 MB of packaged index data plus pandas, pyarrow, and duckdb, which a metadata question does not need. See Data Access Options.

  1. Not installed, and the task is read-only metadata — counts, attribute values, collection

plotting, pydicom/SimpleITK, pathology tiling, results past 10 000 rows, or a version-pinned script the user re-runs? Install idc-index: check_version.py exits non-zero and prints the exact install command for the running interpreter. Prefer a virtual environment, then restart Python.

  1. Not installed, and the task needs more than metadata — downloading files, pandas or

idc-index (GitHub) is still the most capable path and the only one that moves image bytes; the rule is just not to pay for it before the task calls for it. check_version.py never installs anything itself — it also flags a newer idc-index or skill release when one exists.

Setup for the idc-index path:

from idc_index import IDCClient
client = IDCClient()

# Verify IDC data version (should be "v24")
print(f"IDC data version: {client.get_idc_version()}")

Core workflow: query metadata with client.sqlquery() → download with client.downloadfromselection() → visualize with client.getviewerURL(). Python examples below assume this client; Data Access Options has the REST equivalents. For current data scale, run the summary query in references/sqlpatterns.md or GET /v3/stats.

IDC MCP Server

IDC operates a hosted MCP server at https://api.imaging.datacommons.cancer.gov/mcp (streamable HTTP, no authentication). Where it is available it complements — it does not replace — the idc-index workflow below.

Identify it by the MCP resource idc://guide, or by three or more of the tool names buildcohort, getcohorturls, listanalysisresults, and getidcversion. Generic names such as runsql are not evidence on their own. If identification is ambiguous, use idc-index.

If this session has the server, treat it as authoritative for discovery and metadata — IDC version, counts, attribute values, cohort building, metadata SQL — and follow the server's own instructions rather than re-deriving them from this file. Its data version is whatever the server reports: call getidcversion instead of relying on the version pinned in this file.

Return here for what the server does not do: downloading files, local pandas/notebook analysis, DICOMweb, BigQuery, digital pathology tiling, and reproducible scripts. Hand off by passing SeriesInstanceUIDs from the server to client.downloadfromselection(...), and run scripts/check_version.py at that point.

If it is not available, the identical service is reachable with no configuration as a REST API at https://api.imaging.datacommons.cancer.gov/v3 — use it for read-only metadata rather than installing idc-index, per the routing gate in Overview. Suggest connecting the MCP server at most once, only for repeated interactive discovery, and never change the user's configuration yourself.

See references/mcp_guide.md for the tool inventory, handoff patterns, and per-host notes.

When to Use This Skill

  • Finding publicly available radiology (CT, MR, PET) or pathology (slide microscopy) images
  • Selecting image subsets by cancer type, modality, anatomical site, or other metadata
  • Downloading DICOM data from IDC
  • Checking data licenses before use in research or commercial applications
  • Visualizing medical images in a browser without local DICOM viewer software

Quick Navigation

Inline below: the MCP/REST routing rules, the IDC data model, the index tables and how they join, the core API patterns (query, download, visualize, license, cite), best practices, and troubleshooting.

Reference Guides (load on demand):

IDC Data Model

IDC adds two grouping levels above the standard DICOM hierarchy (Patient → Study → Series → Instance):

  • collectionid: Groups patients by disease, modality, or research focus (e.g., tcgaluad, nlst). A patient belongs to exactly one collection.
  • analysisresultid: Identifies derived objects (segmentations, annotations, radiomics features) across one or more original collections. Use it to find AI-generated or expert annotations, while collection_id finds original imaging data (which may itself include deposited annotations).

Key identifiers for queries:

Index Tables

The idc-index package provides multiple metadata index tables, accessible via SQL or as pandas DataFrames. The REST API exposes the same tables through GET /tables and POST /sql.

Important: client.indicesoverview is the authoritative source for current table descriptions, available columns, and their types — query it when writing SQL or exploring data structure. It also answers "which table contains column X"; see references/indextables_guide.md for that search pattern and full schema discovery.

Available Tables

Always call client.fetchindex("tablename") before querying any index table — it is safe and idempotent for all tables, including those loaded automatically at startup.

references/indextablesguide.md has the full inventory with each table's columns and contents — load it when you need to know what a specialized table actually holds.

priorversionsindex is for reproducibility only. It contains series permanently removed from IDC, with zero overlap with index. Use it only to reproduce work against a prior IDC version. Do NOT use it for version history or "what's new" questions — those use seriesinitidcversion / seriesrevisedidcversion in the main index table, which are not equivalent to this table's minidcversion / maxidcversion.

Joining Tables

SeriesInstanceUID is the universal join key for all series-level specialized tables: smindex, sminstanceindex, segindex, annindex, anngroupindex, contrastindex, volumegeometryindex, rtstructindex, ctindex, mrindex, ptindex. Always join these to index on SeriesInstanceUID. The exceptions below use different column names.

Note: subjects, updated, and description appear in multiple tables but have different meanings (counts vs identifiers, different update contexts). Joining priorversionsindex to index on SeriesInstanceUID always returns zero rows — see the warning above.

For detailed join examples, schema discovery patterns, key columns reference, and DataFrame access, see references/indextablesguide.md.

Clinical Data Access

Clinical (non-imaging) attributes — staging, demographics, therapy — live in per-collection tables. client.fetchindex("clinicalindex") loads the dictionary mapping columns to collections; client.getclinicaltable(name) returns one table as a DataFrame.

See references/clinicaldataguide.md for the discovery workflow, coded-value mapping, and joining clinical data with imaging.

Data Access Options

The IDC Portal (https://portal.imaging.datacommons.cancer.gov/) is interactive only — browser-based exploration, manual cohort selection, and download. Unlike every option above it has no programmatic interface, so point a user there to browse or click through data themselves; never use it as a step in a script or workflow.

REST API — the no-install metadata path

https://api.imaging.datacommons.cancer.gov/v3, no authentication: discovery, cohort counts and manifests, read-only SQL, clinical tables, viewer URLs, licenses, citations. It is the same service as the MCP server over plain HTTP, so it needs no configuration. It never moves image bytes — switch to idc-index to download, to get a DataFrame, or for results past 10 000 rows.

B=https://api.imaging.datacommons.cancer.gov/v3
curl -s $B/version   # idc_version, idc_index_data_version, api_version
curl -s $B/stats     # collections, patients, studies, series, instances, size_TB
curl -s "$B/attributes/Modality/values?limit=5"   # real filter values, with counts
curl -s $B/sql -H 'content-type: application/json' \
  -d '{"sql":"SELECT collection_id, COUNT(*) n FROM index GROUP BY 1 ORDER BY n DESC LIMIT 3"}'
curl -s $B/cohort/counts -H 'content-type: application/json' \
  -d '{"filters":{"terms":{"collection_id":["rider_pilot"]}}}'

The filter object always goes under filters — on cohort/counts, cohort/manifest, cohort/manifest.txt, licenses, and citations alike. A bare filter or an unrecognized key is a 422 naming the fix; an unfiltered series-enumerating request is a 400, not the whole archive. Every filtered response echoes filters_applied and warnings — read them, because they name any predicate the server dropped. A zero count with empty warnings therefore means the filter matched nothing, not that a value was miscased; miscasing produces a warning that says so.

POST /sql takes one read-only SELECT/WITH over the tables idc-index exposes plus clinical.; maxrows defaults to 5 000, caps at 10 000, and truncated flags clipping. GET /attributes lists the 19 filterable attributes — clinical values, segmented anatomy, and acquisition parameters are not among them and need SQL. There is no rate limit or quota. Use v3 only: V1 and V2 are superseded and scheduled for shutdown, so port any /v1/- or Modalitybtw-style example a user brings rather than extending it.

Both sides build on idc-index-data, so compare the API's idcindexdataversion against local idcindexdata.version before mixing them: the major is the IDC data release (24.x.y serves v24), so differing minor/patch means the series are identical. If the API is a whole release ahead, idc-index cannot download the extra series — it silently skips what its own index does not list — so either upgrade it (run scripts/checkversion.py for the right command) or transfer directly from the bucket with s5cmd --no-sign-request.

See references/restapiguide.md for the endpoint reference, filter grounding, limits, and the manifest-based download flow.

Cloud storage organization

All DICOM files live in public buckets mirrored between AWS S3 and GCS, organized by CRDC UUIDs (not DICOM UIDs) to support versioning, as /.dcm. Access is free (no egress fees) via AWS CLI, gsutil, or s5cmd with anonymous access; use the seriesawsurl column for S3 URLs. Note that idc-open-data-cr / idc-open-cr (~4% of data) is commercial-use restricted (CC BY-NC). See references/cloudstorageguide.md for the full bucket list and UUID mapping.

DICOMweb access

More skills from K-Dense-AI/scientific-agent-skills

  • AadaptyvHow to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
  • AaeonThis skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
  • AalphagenomeLook up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), score variants or scan windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and build Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.
  • Aanalytical-method-validationPlan, execute, and document validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and ligand binding or cell-based assays whenever the question is whether a procedure is fit for its intended purpose. Triggers include "method validation", "analytical method validation", "AMV", "validation protocol", "acceptance criteria", "linearity", "reportable range", "accuracy and precision", "repeatability", "intermediate precision", "recovery", "LOD", "LOQ", "detection limit", "quantitation limit", "specificity", "robustness", "method transfer", "method comparison", "Deming", "Passing-Bablok", "Bland-Altman", "equivalence testing", "OOS investigation", "ICH Q2", "Q2(R2)", "Q14", "USP 1225", "ICH M10", "incurred sample reanalysis", "ISR", "CLSI EP", and any request to show that an assay works.
  • AanndataData structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
  • AarborAutonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
  • AarboretoInfer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
  • AastropyCore Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.
  • AautoskillObserve the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.
  • Abenchling-integrationBenchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.
  • Abgpt-paper-searchSearch scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.
  • AbidsUse this skill when working with Brain Imaging Data Structure (BIDS) datasets: organizing neuroscience and biomedical data (MRI, EEG, MEG, iEEG, PET, microscopy, NIRS, motion capture, EMG, MR spectroscopy, behavioral), querying BIDS layouts, validating compliance, converting DICOM to BIDS, writing metadata sidecars, or creating BIDS derivatives.

All agent skills → · MCP servers