Mmcp.market

hypothesis-generation skill

by K-Dense-AI·K-Dense-AI/scientific-agent-skills·47k stars·MIT

Formulate evidence-bounded scientific questions, candidate hypotheses, rival explanations, causal or associational claims, discriminating predictions, measurements, and preregistration-ready analysis plans. Use when turning observations or preliminary findings into transparent, testable research plans without treating hypotheses as facts.

A100/100content scan

Is the hypothesis-generation skill safe?

Clean: nothing in its files matched our rules. We read 27 files in the folder on 2026-09-28.

No findings.

Install the hypothesis-generation skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git /tmp/scientific-agent-skills
mkdir -p ~/.claude/skills
cp -r /tmp/scientific-agent-skills/skills/hypothesis-generation ~/.claude/skills/hypothesis-generation
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Scientific Hypothesis Generation

Turn an observation into a transparent set of candidate explanations and tests. A hypothesis is a proposal to be challenged, not a finding, fact, diagnosis, or recommendation.

Non-negotiable boundaries

Before using unpublished, sensitive, controlled, personal, proprietary, export-controlled, or security-relevant material:

  1. Confirm authorization and the applicable institutional, funder, publisher, data-use, privacy, and AI policies.
  2. Keep the material local unless an authorized human explicitly approves a named external destination and data scope.
  3. Minimize inputs. Do not place sensitive or unpublished data in web searches or external AI systems without authorization.
  4. Stop at the appropriate human, animal, biosafety, dual-use, data-governance, or regulatory gate.

Never:

  • present a hypothesis, mechanism, causal effect, citation, or apparent pattern as established evidence;
  • claim novelty because a quick search found nothing;
  • infer causation from association, temporal order alone, predictive accuracy, or model output;
  • supply patient-specific diagnosis, treatment, dose, prognosis, or other clinical advice;
  • provide harmful experimental optimization or operational detail for pathogens, toxins, weapons, evasion, or other misuse;
  • bypass IRB/REC, IACUC, IBC, biosafety, dual-use, privacy, legal, or regulatory review;
  • fabricate sources, identifiers, search coverage, data, results, approvals, or preregistration;
  • automatically score, rank, select, accept, or reject scientific hypotheses.

If a request crosses a safety gate, produce only a high-level risk/oversight note and route it to the qualified local authority. Do not continue with operational detail.

Keep the objects distinct

Do not collapse these labels. A mechanistic story is not a prediction; a prediction is not evidence; rejection of one null does not prove a mechanism; support for one candidate does not eliminate unconsidered rivals.

Workflow

1. Run the scope and safety gate

Record:

  • accountable human owner and intended use;
  • data sensitivity, authorization, retention, and permitted processing;
  • affected people, animals, ecosystems, communities, or security interests;
  • required ethics, feasibility, biosafety, dual-use, and regulatory reviews;
  • unresolved blocks and domain expertise needed.

No script approval is an ethics, safety, regulatory, or scientific approval.

2. Freeze the observation

Write the observation before interpretation:

  • measurement or source;
  • population, system, place, and time;
  • unit of observation and unit of analysis;
  • uncertainty, missingness, exclusions, and preprocessing;
  • whether the pattern was expected, exploratory, or selected after viewing results.

Use “reported,” “observed,” or “associated,” not causal language, unless a causal design and estimand justify it.

3. Frame the research question

Choose a framework only when it fits:

  • PICO/PICOT for intervention/effectiveness questions: population, intervention, comparator, outcome, and optionally time.
  • PECO for exposure questions.
  • Population–index test–reference standard–target condition for diagnostic accuracy.
  • Population–prognostic factor–outcome–time for prognosis.
  • A domain-specific construct–context–outcome frame for qualitative, descriptive, mechanistic, or theoretical work.

PICO is not a universal template. Define stakeholders, context, boundaries, feasibility, and what answer would change knowledge or practice. FINER is a question-refinement mnemonic—Feasible, Interesting, Novel, Ethical, Relevant—not a scoring system. Treat “Novel” as unresolved until a documented, fit-for-purpose search and expert review support it.

4. Establish a dated evidence boundary

Search before making literature-dependent statements. Prefer primary research, official policies, primary methods papers, current reporting guidelines, and systematic reviews used for orientation.

Record:

  • search date and cutoff;
  • databases/indexes, queries, filters, and screening boundary;
  • included and excluded source types;
  • sources supporting, challenging, or contextualizing each claim;
  • known access, language, database, and time limitations.

A search can establish what was searched, not universal absence. Say “not located within the documented search boundary,” never “no prior work exists.” Use assets/searchboundarytemplate.json, assets/evidenceledgertemplate.csv, and references/literaturesearchstrategies.md.

5. Generate rivals before choosing tests

Create multiple candidates from genuinely different explanatory classes when plausible:

  • proposed mechanism;
  • measurement or processing artifact;
  • confounding or common cause;
  • selection or attrition;
  • conditioning on a collider;
  • reverse causation;
  • temporal, contextual, or boundary-condition differences;
  • stochastic variation;
  • competing mechanisms at another scale.

Generate an initial rival set independently before AI-assisted expansion to reduce anchoring and homogenization. Do not force a fixed number or false symmetry. Keep every candidate labeled candidate.

Platt’s strong-inference pattern motivates alternative hypotheses and crucial tests, but failed alternatives do not make the survivor true. Unknown alternatives, auxiliary assumptions, measurement error, and mixed mechanisms remain possible.

6. Declare the claim type and estimand

Classify each target as:

  • descriptive;
  • associational;
  • predictive;
  • causal;
  • mechanistic.

For a causal target, define before analysis:

  • target population or system;
  • intervention/exposure and comparator;
  • outcome and time horizon;
  • population-level summary;
  • treatment versions and intercurrent-event handling where relevant;
  • identification assumptions and target-trial/design analogue.

Document confounding, selection, collider, measurement, and reverse-causation risks separately. An observational causal estimate remains assumption-dependent. Use references/causalinferenceand_claims.md.

7. Derive discriminating predictions

For every candidate:

  1. State conditions and boundary conditions.
  2. Name the observable and measurement.
  3. State the expected pattern and uncertainty.
  4. State a result incompatible with the candidate under declared assumptions.
  5. Contrast the expected result with at least one rival.
  6. Define indeterminate outcomes and what would be learned from them.

Prefer tests where rivals predict meaningfully different outcomes. Add positive, procedural, and negative controls when scientifically appropriate. A negative control must be incapable of operating through the target mechanism while sharing relevant bias pathways; it is not a decorative untreated group.

Use assets/predictionrivalmatrixtemplate.csv and assets/falsificationcontrols_template.json.

8. Operationalize and validate measurement

For every construct record:

  • variable role and operational definition;
  • population/system, unit, timing, and conditions;
  • instrument/method, calibration, quality control, and masking;
  • reliability/repeatability;
  • validity evidence and applicability;
  • missingness, detection limits, transformations, cut points, and their rationales;
  • measurement invariance or cross-group comparability when relevant;
  • foreseeable measurement bias and limitations.

Do not treat a convenient proxy as the construct itself. Validate with:

python3 scripts/check_operationalization.py local-operationalization.json

9. Match design and analysis to the claim

Specify:

  • sampling, experimental unit, allocation, randomization, masking, and controls;
  • inclusion/exclusion and stopping rules;
  • sample-size, precision, or information rationale based on declared assumptions;
  • outcomes, contrasts, estimands, models, effect measures, and uncertainty;
  • missing-data and intercurrent-event handling;
  • multiplicity across outcomes, models, subgroups, looks, and hypotheses;
  • assumptions, diagnostics, robustness, and sensitivity analyses;
  • replication or independent validation plan;
  • what is confirmatory versus exploratory.

Do not use universal sample-size minima. Do not interpret a thresholded p-value as the probability a hypothesis is true or as effect importance. See references/experimentaldesignpatterns.md.

For intervention trials, use the current SPIRIT 2025 protocol guidance and CONSORT 2025 reporting guidance where applicable. These improve completeness; they do not certify design quality, ethics, or regulatory compliance.

10. Prevent HARKing and expose deviations

Before accessing the target outcomes, timestamp the question, candidates, predictions, outcomes, exclusions, transformations, analysis, multiplicity, missing-data plan, and stopping rule when feasible.

Afterward:

  • label data-dependent ideas and analyses exploratory;
  • preserve and report planned analyses;
  • list deviations with date, rationale, who decided, and expected impact;
  • never rewrite an observed pattern as an a priori prediction.

Preregistration is a transparent plan, not a ban on adaptation. Registered Reports add results-blind peer review and in-principle acceptance under journal policy. See references/preregistrationandopen_science.md.

11. Plan replication and updating

More skills from K-Dense-AI/scientific-agent-skills

  • AadaptyvHow to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
  • AaeonThis skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
  • AalphagenomeLook up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), score variants or scan windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and build Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.
  • Aanalytical-method-validationPlan, execute, and document validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and ligand binding or cell-based assays whenever the question is whether a procedure is fit for its intended purpose. Triggers include "method validation", "analytical method validation", "AMV", "validation protocol", "acceptance criteria", "linearity", "reportable range", "accuracy and precision", "repeatability", "intermediate precision", "recovery", "LOD", "LOQ", "detection limit", "quantitation limit", "specificity", "robustness", "method transfer", "method comparison", "Deming", "Passing-Bablok", "Bland-Altman", "equivalence testing", "OOS investigation", "ICH Q2", "Q2(R2)", "Q14", "USP 1225", "ICH M10", "incurred sample reanalysis", "ISR", "CLSI EP", and any request to show that an assay works.
  • AanndataData structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
  • AarborAutonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
  • AarboretoInfer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
  • AastropyCore Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.
  • AautoskillObserve the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.
  • Abenchling-integrationBenchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.
  • Abgpt-paper-searchSearch scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.
  • AbidsUse this skill when working with Brain Imaging Data Structure (BIDS) datasets: organizing neuroscience and biomedical data (MRI, EEG, MEG, iEEG, PET, microscopy, NIRS, motion capture, EMG, MR spectroscopy, behavioral), querying BIDS layouts, validating compliance, converting DICOM to BIDS, writing metadata sidecars, or creating BIDS derivatives.

All agent skills → · MCP servers