pydicom skill
Use pydicom to read, inspect, write, transform, and safely preflight local DICOM datasets and pixel data. Applies to DICOM metadata, transfer syntaxes, compression plugins, frames, private elements, JSON, and bounded de-identification review.
Is the pydicom skill safe?
Clean: nothing in its files matched our rules. We read 13 files in the folder on 2026-09-28.
No findings.
Install the pydicom skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git /tmp/scientific-agent-skills mkdir -p ~/.claude/skills cp -r /tmp/scientific-agent-skills/skills/pydicom ~/.claude/skills/pydicom
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
pydicom
Use pydicom for DICOM dataset I/O and pixel processing. Version 3.0.2 is the current stable release reviewed here. It fixes CVE-2026-32711, a crafted DICOMDIR path-traversal issue. pydicom 3.0.2 declares Python >=3.10; its bundled DICOM dictionary is 2024c, while the live DICOM Standard may be newer.
Mandatory safety boundary
and pixels may contain protected health information (PHI).
- Work only with local data that the user is authorized to access.
- DICOM metadata, file names, private elements, overlays, structured content,
default. Use a documented allowlist and aggregate output.
- Never print Dataset, export full metadata/JSON, or log element values by
validation, conversion, and plugin availability are not diagnostic claims.
- pydicom is a general DICOM framework, not a diagnostic viewer. Pixel output,
threat-context-specific. It requires privacy/DICOM expert verification.
- De-identification is profile-, purpose-, recipient-, jurisdiction-, and
compliance. Preserve originals and audit derived outputs.
- Never claim that a tag-removal script is DICOM PS3.15, HIPAA, GDPR, or other
secrets: use least privilege and encrypted/managed secret storage, never commit, sync, log, or share them with derivatives, and define backup, rotation, revocation, and destruction procedures. A leaked key invalidates the intended separation; rotation also changes deterministic mappings.
- Treat deterministic pseudonymization keys and UID maps as re-identification
limits before parsing untrusted or unusually large datasets.
- Set explicit input-file, file-count, frame-count, decoded-byte, and output
Installation
Create or activate an isolated environment, then install the exact reviewed release:
uv pip install "pydicom==3.0.2"Uncompressed pixel arrays and image rendering:
uv pip install "pydicom==3.0.2" "numpy==2.5.1" "Pillow==12.3.0"Install only the transfer-syntax plugins required by the deployment:
# JPEG/JPEG-LS, JPEG 2000/HTJ2K, and faster RLE through pylibjpeg
uv pip install "numpy==2.5.1" "pylibjpeg==2.1.0" \
"pylibjpeg-libjpeg==2.4.0" "pylibjpeg-openjpeg==2.5.0" \
"pylibjpeg-rle==2.2.0"
# JPEG-LS encoder/decoder
uv pip install "numpy==2.5.1" "pyjpegls==1.5.1"
# Alternative decoder with platform-specific wheels
uv pip install "python-gdcm==3.2.6"Plugin licenses and wheels differ by package/platform; review them before deployment. Pillow has documented decoding limitations and pydicom cautions that plugin output must be independently checked.
Native codec wheels widen the supply-chain and memory-safety boundary. For a controlled deployment, resolve these exact pins on a trusted build host, lock and verify wheel hashes/provenance, mirror approved artifacts internally, scan them, and install with hash enforcement rather than resolving from the public index at runtime.
Choose the workflow
scripts/transfersyntaxinspector.py.
- Need an aggregate overview: run scripts/extract_metadata.py.
- Need bounded technical checks: run scripts/dicom_inventory.py.
- Need codec deployment preflight: run
scripts/dicomtoimage.py.
- Need frame/memory planning: run scripts/pixelframeplanner.py.
- Need one non-diagnostic rendered frame: run
a site-reviewed action profile, then run scripts/anonymizedicom.py and scripts/deidentificationaudit.py.
- Need a pseudonymized derivative: read the de-identification section, create
scripts/uidmappingvalidator.py.
- Need to check a sensitive UID map: run
Read datasets safely
dcmread() returns a FileDataset, a Dataset subclass with File Format state such as file_meta, preamble, and original encoding.
from pathlib import Path
import pydicom
path = Path("authorized/input.dcm")
ds = pydicom.dcmread(
path,
stop_before_pixels=True,
specific_tags=[
"SOPClassUID",
"Modality",
"Rows",
"Columns",
"NumberOfFrames",
],
)
technical = {
"sop_class": ds.get("SOPClassUID"),
"modality": ds.get("Modality"),
"rows": ds.get("Rows"),
"columns": ds.get("Columns"),
}Use:
check; it does not prove the bytes are valid DICOM.
- stopbeforepixels=True for metadata-only work.
- specific_tags=[...] for a minimum allowlist.
- defer_size="1 MiB" when a later write must preserve large values.
- force=False (default). force=True only bypasses the File Format header
Do not call print(ds), repr(ds), or iterate values into logs on clinical data.
Dataset, DataElement, and sequences
Access standard elements by keyword and check for absence:
modality = ds.get("Modality", "UNSPECIFIED")
if "ReferencedImageSequence" in ds:
for item in ds.ReferencedImageSequence:
referenced_class = item.get("ReferencedSOPClassUID")Tag access, such as ds[0x0010, 0x0010], returns a DataElement; its .value is separate. Sequence behaves like a list of nested Dataset items. Privacy actions must recurse through every sequence item, not only the top level.
When creating a file, use FileMetaDataset for group 0002, keep dataset and file-meta SOP UIDs consistent, set a Transfer Syntax UID, and write in enforced File Format:
from pydicom import dcmwrite
from pydicom.dataset import FileDataset, FileMetaDataset
from pydicom.uid import CTImageStorage, ExplicitVRLittleEndian, generate_uid
meta = FileMetaDataset()
meta.MediaStorageSOPClassUID = CTImageStorage
meta.MediaStorageSOPInstanceUID = generate_uid()
meta.TransferSyntaxUID = ExplicitVRLittleEndian
ds = FileDataset(None, {}, file_meta=meta, preamble=b"\0" * 128)
ds.SOPClassUID = meta.MediaStorageSOPClassUID
ds.SOPInstanceUID = meta.MediaStorageSOPInstanceUID
# Add all attributes required by the selected IOD before writing.
dcmwrite("new.dcm", ds, enforce_file_format=True, overwrite=False)writelikeoriginal is deprecated in pydicom 3.0; use enforcefileformat. A successful write is not full PS3.3 IOD conformance.
UIDs and transfer syntax
The File Meta Information Transfer Syntax UID controls dataset encoding and pixel compression:
ts = ds.file_meta.TransferSyntaxUID
summary = {
"uid": str(ts),
"name": ts.name,
"compressed": ts.is_compressed,
"implicit_vr": ts.is_implicit_VR,
"little_endian": ts.is_little_endian,
}pydicom 3.0 chooses write encoding from the Transfer Syntax UID before legacy dataset flags. Do not replace structural UIDs (Transfer Syntax, SOP Class, or coding-scheme UIDs) during pseudonymization. Instance/reference UID replacement must be one-to-one and consistent across the complete declared scope.
Read references/transfer_syntaxes.md before compression, decompression, or encapsulation.
Pixel data and frames
The stable pydicom.pixels API supports path-based, frame-specific decoding:
from pydicom.pixels import pixel_array
# Reads only the selected frame where the source permits it.
frame = pixel_array("authorized/image.dcm", index=0, raw=False)Shape semantics:
- grayscale single frame: (rows, columns)
- grayscale multi-frame: (frames, rows, columns)
- color single frame: (rows, columns, samples)
- color multi-frame: (frames, rows, columns, samples)
raw=False converts YCbCr pixel data to RGB when possible; raw=True retains the decoded color space after mandatory minimal processing. Use iter_pixels(path, indices=[...]) for bounded multi-frame iteration.
More skills from K-Dense-AI/scientific-agent-skills
- AadaptyvHow to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
- AaeonThis skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
- AalphagenomeLook up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), score variants or scan windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and build Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.
- Aanalytical-method-validationPlan, execute, and document validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and ligand binding or cell-based assays whenever the question is whether a procedure is fit for its intended purpose. Triggers include "method validation", "analytical method validation", "AMV", "validation protocol", "acceptance criteria", "linearity", "reportable range", "accuracy and precision", "repeatability", "intermediate precision", "recovery", "LOD", "LOQ", "detection limit", "quantitation limit", "specificity", "robustness", "method transfer", "method comparison", "Deming", "Passing-Bablok", "Bland-Altman", "equivalence testing", "OOS investigation", "ICH Q2", "Q2(R2)", "Q14", "USP 1225", "ICH M10", "incurred sample reanalysis", "ISR", "CLSI EP", and any request to show that an assay works.
- AanndataData structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
- AarborAutonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
- AarboretoInfer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
- AastropyCore Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.
- AautoskillObserve the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.
- Abenchling-integrationBenchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.
- Abgpt-paper-searchSearch scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.
- AbidsUse this skill when working with Brain Imaging Data Structure (BIDS) datasets: organizing neuroscience and biomedical data (MRI, EEG, MEG, iEEG, PET, microscopy, NIRS, motion capture, EMG, MR spectroscopy, behavioral), querying BIDS layouts, validating compliance, converting DICOM to BIDS, writing metadata sidecars, or creating BIDS derivatives.