Mmcp.market

analytical-method-validation skill

by K-Dense-AI·K-Dense-AI/scientific-agent-skills·47k stars·MIT

Plan, execute, and document validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and ligand binding or cell-based assays whenever the question is whether a procedure is fit for its intended purpose. Triggers include "method validation", "analytical method validation", "AMV", "validation protocol", "acceptance criteria", "linearity", "reportable range", "accuracy and precision", "repeatability", "intermediate precision", "recovery", "LOD", "LOQ", "detection limit", "quantitation limit", "specificity", "robustness", "method transfer", "method comparison", "Deming", "Passing-Bablok", "Bland-Altman", "equivalence testing", "OOS investigation", "ICH Q2", "Q2(R2)", "Q14", "USP 1225", "ICH M10", "incurred sample reanalysis", "ISR", "CLSI EP", and any request to show that an assay works.

A100/100content scan

Is the analytical-method-validation skill safe?

Clean: nothing in its files matched our rules. We read 17 files in the folder on 2026-09-28.

No findings.

Install the analytical-method-validation skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git /tmp/scientific-agent-skills
mkdir -p ~/.claude/skills
cp -r /tmp/scientific-agent-skills/skills/analytical-method-validation ~/.claude/skills/analytical-method-validation
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Analytical Method Validation

When to use

Any time the question is whether an analytical procedure is fit for its intended purpose: designing a validation study, evaluating validation data, verifying a compendial procedure, transferring a procedure to another laboratory or instrument, or defending any of these in a report.

The two rules

1. Establish which framework governs before designing anything. The same assay validates differently under ICH Q2(R2), USP <1225>, ICH M10, CLSI EP, and ISO/IEC 17025. They differ in which characteristics are required, how the studies are laid out, and whether numeric acceptance criteria are supplied at all. Blending them produces a protocol that satisfies none of them.

2. State acceptance criteria before collecting data. Criteria chosen after seeing results are not acceptance criteria, and deciding them post hoc is a standing audit finding. ICH Q2(R2) deliberately supplies almost no numeric criteria — they have to come from the specification, the analytical target profile (ICH Q14 section 3), or development data. ICH M10 is the exception: it supplies explicit numbers, and they differ between chromatographic assays and ligand binding assays.

Scope

This skill plans studies, computes the statistics correctly, and structures the documentation. It does not decide that a procedure is validated, release a batch, accept or reject a run, close an investigation, or substitute for the analyst, the technical reviewer, the quality unit, or the regulator. Every script reports; none of them concludes.

Copyright boundary

ICH guidelines are published openly and licensed for reuse with acknowledgement, so their requirements are encoded directly in this skill. USP general chapters, CLSI EP documents, and ISO standards are copyrighted and paywalled. For those, this skill supplies the designation, scope, and where to obtain an authorised copy — never the text, never invented thresholds. Do not ask an agent to retrieve, transcribe, or reconstruct their content. If a number matters and it lives in a paywalled document, read it from the authorised copy.

Frameworks

cd skills/analytical-method-validation/scripts
python3 plan_validation.py --list-frameworks

Q2(R2) replaced Q2(R1) in November 2023 and restructured the characteristics. Range is now the parent characteristic (section 3.2), containing response (linearity) and validation of lower range limits (DL/QL). Accuracy and precision are section 3.3 and may be evaluated in combination against a single criterion. Robustness is treated as a development activity and cross-refers to ICH Q14. Multivariate procedures are addressed explicitly (2.5 and 3.2.2.3), and Annex 2 adds worked examples for techniques Q2(R1) never covered — quantitative ¹H-NMR, NIR, quantitative LC/MS, qPCR, biological assays, and particle size. A Q2(R1)-shaped protocol — a flat list of linearity, range, accuracy, precision, specificity, LOD, LOQ, robustness — is out of date. Note also the error correction dated 30 November 2023 to Table 5 and Tables 6–11.

Scripts

cd skills/analytical-method-validation/scripts

All take --format table|tsv|json. Provenance, guideline citations, and caveats go to stderr; data goes to stdout, so > out.tsv keeps them separate. Exit code is 0 for no findings, 1 when findings were raised, 2 for bad input — so any of them can gate a workflow.

Workflow

1. Fix the framework and the required characteristics

python3 plan_validation.py --framework ich-q2r2 --attribute assay --technique hplc --range-use assay

Q2(R2) Table 1 decides what is required from the measured attribute, not from the technique. For an assay: specificity, response, accuracy, repeatability, intermediate precision. For a limit test: specificity and DL only. For an identity test: specificity alone. Attributes accepted include assay, impurity (quantitative), impurity-limit, and identity.

Reportable range comes from the specification. Q2(R2) Table 2 gives worked examples — 80–120% of declared content for an assay, 70–130% for content uniformity, reporting threshold to 120% of the specification for an impurity.

2. Generate the protocol and fill in the criteria

python3 plan_validation.py --framework ich-q2r2 --attribute impurity --protocol > protocol.md

Every bracketed field is a decision to make and record before data collection. The protocol skeleton deliberately refuses to pre-fill acceptance criteria for Q2(R2) work, because there is no defensible default.

3. Evaluate the response

python3 check_response.py -i calibration.csv --max-back-calc-error 2

Input is level,response, one row per injection; repeated rows at the same level are replicates, and supplying them is what makes the linearity test possible.

Real output from a curve that a coefficient of determination would wave through:

statistic                           value
distinct levels                     5
slope                               166.6000
intercept                           2495.0000
intercept CI includes 0             no
coefficient of determination (r2)   0.9830
lack-of-fit F                       469.5294
lack-of-fit p                       1.5139e-06
runs test p                         0.0492

level     n  mean_response  mean_back_calculated  relative_error_pct
50.0000   2  10075.0000     45.4982               -9.0036
75.0000   2  15150.0000     75.9604               1.2805
100.0000  2  20050.0000     105.3721              5.3721
125.0000  2  24050.0000     129.3818              3.5054
150.0000  2  26450.0000     143.7875              -4.1417

r² = 0.983 and the model is unusable: −9.0% back-calculated error at the bottom of the range, lack-of-fit p = 1.5 × 10⁻⁶, non-random residual signs. r² is not evidence of linearity — it rises with range and is nearly insensitive to curvature. The lack-of-fit F test against pure error and the residual pattern are the evidence, which is why Q2(R2) 3.2.2.1 asks for an analysis of the deviation of points from the line rather than a correlation coefficient alone.

Add --weight 1/x2 for a wide-range curve. The script flags heteroscedasticity when the residual variance in the top third of the range exceeds the bottom third by more than 10×, because an unweighted fit then biases exactly the low end where a reporting threshold lives.

4. Evaluate accuracy and precision

python3 check_accuracy_precision.py -i ap.csv --accuracy-limit 2 --rsd-limit 1.0 --design-check assay

Input is level,measured,group, where group is the intermediate-precision factor — day, analyst, or instrument.

level  component                       sd      rsd_pct  df      ci90_low_sd  ci90_high_sd
100    repeatability (within group)    0.0707  0.0707   3       0.0438       0.2065
100    between-group                   1.6515  1.6515   2       n/a          n/a
100    intermediate precision (total)  1.6530  1.6530   2.0037  0.9554       7.2821

Repeatability of 0.07% RSD looks superb; intermediate precision is 1.65%, twenty-three times larger, because the variability lives entirely between days. Reporting the within-day figure as the procedure's precision would understate routine performance by more than an order of magnitude. This is why the script fits a one-way random-effects model rather than pooling.

Two traps the script handles for you:

results into one standard deviation turns the range itself into apparent imprecision. The script reports per level, plus a level-independent view as percent of nominal.

  • Precision is estimated within each level, never pooled across levels. Pooling 80/100/120%

limit, not just the mean. Q2(R2) 3.3.1.4 asks for the interval to be compatible with the criterion; a mean that scrapes inside on six replicates has not demonstrated much.

  • --require-ci-within-limit enforces that the whole confidence interval sits inside the

5. Establish DL and QL, and confirm them

python3 check_detection_limits.py --calibration lowcal.csv --blanks blanks.csv \
    --confirm-ql 0.05 --confirm-data ql_check.csv --reporting-threshold 0.05
approach                                          sigma   slope      DL      QL
sd-and-slope (sigma = residual SD of regression)  7.2816  5033.3490  0.0048  0.0145
sd-and-slope (sigma = SD of y-intercept)          4.3303  5033.3490  0.0028  0.0086
sd-and-slope (sigma = SD of 8 blanks)             3.7702  5033.3490  0.0025  0.0075

The same data give QL estimates spanning 1.9×, purely from the choice of σ. Q2(R2) 3.2.3.5 therefore requires the limit and the approach used to determine it to be reported, and an estimated limit to be confirmed with samples at or near it. For an impurity procedure the QL must be at or below the reporting threshold. Reaching for 3.3σ/slope reflexively, reporting one number with no named approach, and never confirming it are three separate findings.

6. Bioanalytical runs under ICH M10

python3 check_bioanalytical_run.py --modality chromatographic --run run1.csv
python3 check_bioanalytical_run.py --modality lba --isr isr.csv
python3 check_bioanalytical_run.py --modality lba --criteria

--modality is mandatory and has no default, because the criteria genuinely differ:

Applying the ±15% chromatographic numbers to a ligand binding assay, or importing the LBA total-error criterion into a chromatographic method, are both common and both wrong.

The run check enforces the per-level rule that gets missed: at least 2/3 of all QCs and at least 50% at each level. A run can pass the overall fraction while a single level fails completely.

finding: QC level high: 0/2 within tolerance (0%); M10 requires at least 50% at each level

7. Transfer and method comparison

python3 compare_methods.py -i paired.csv --margin 2 --relative --slope-tolerance 0.05
mean difference (%)                       1.4646
TOST margin                               2.0000
TOST p-value                              1.0528e-13
90% CI (TOST)                             1.44127 to 1.48797
equivalent at stated margin               yes
--- for contrast only ---
paired t-test p (NOT equivalence)         0.0000
OLS slope (biased here)                   1.0396
Deming slope                              1.0398
Passing-Bablok slope                      1.0351

Two errors this replaces:

detect a difference is not evidence of equivalence, and on a small transfer dataset that outcome is close to guaranteed. TOST tests the hypothesis that matters — that the true difference lies inside a pre-stated margin. Here the t test says the difference is highly significant and TOST says the methods are equivalent at ±2%; both are true, and only one answers the question.

  • "p > 0.05, no significant difference, therefore the methods are equivalent." Failing to

error, which is false when comparing two procedures, and biases the slope toward zero. Deming (with a stated error-variance ratio) and Passing–Bablok (non-parametric, outlier-resistant) are the appropriate regressions and are reported side by side with OLS for contrast.

  • Ordinary least squares for method comparison. OLS assumes the reference values carry no

The script also flags proportional bias — when the difference trends with concentration, a single mean bias and its limits of agreement are misleading regardless of how tight they look.

More skills from K-Dense-AI/scientific-agent-skills

  • AadaptyvHow to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
  • AaeonThis skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
  • AalphagenomeLook up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), score variants or scan windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and build Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.
  • AanndataData structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
  • AarborAutonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
  • AarboretoInfer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
  • AastropyCore Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.
  • AautoskillObserve the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.
  • Abenchling-integrationBenchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.
  • Abgpt-paper-searchSearch scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.
  • AbidsUse this skill when working with Brain Imaging Data Structure (BIDS) datasets: organizing neuroscience and biomedical data (MRI, EEG, MEG, iEEG, PET, microscopy, NIRS, motion capture, EMG, MR spectroscopy, behavioral), querying BIDS layouts, validating compliance, converting DICOM to BIDS, writing metadata sidecars, or creating BIDS derivatives.
  • AbiopythonComprehensive molecular biology toolkit. Use for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Best for batch processing, custom bioinformatics pipelines, BLAST automation. For quick lookups use gget; for multi-service integration use bioservices.

All agent skills → · MCP servers