Mmcp.market

power-analysis skill

by pedrohcgs·pedrohcgs/claude-code-my-workflow·1.6k stars·MIT

Compute statistical power, required sample size, and minimum detectable effect (MDE) for a study design, then write a registry-ready power section. Handles two-arm RCTs (with clustering / ICC and unequal allocation), multiple-arm corrections, and a simulation-based power option for non-standard designs (DiD/event-study, IV, panel). Use when user says "power analysis", "power calculation", "MDE", "minimum detectable effect", "how big a sample do I need", "is my study powered", "power for an RCT", or when /preregister needs a power section for an experiment. Produces a power/MDE table, power curves, and a methods paragraph to paste into a preregistration.

A100/100content scan

Is the power-analysis skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the power-analysis skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow
mkdir -p ~/.claude/skills
cp -r /tmp/claude-code-my-workflow/.claude/skills/power-analysis ~/.claude/skills/power-analysis
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

/power-analysis — Power / MDE for study design

Compute the three interlocking quantities of an ex-ante design calculation — power, required N, and minimum detectable effect (MDE) — and emit a power section the user can paste straight into a preregistration. Analytical for standard designs; simulation-based (reusing the /simulation-study harness pattern) for non-standard ones.

Core principle: a power calculation is a design-time commitment made before the data exist. Fix any two of {effect size, N, power} and solve for the third; never back out a "power" number from a realised estimate (that is post-hoc power, and it is uninformative — see "What this skill does NOT do").

When to use

  • Before launching an RCT / field / survey experiment — to choose N (or clusters) for a target MDE at 80–90% power.
  • Invoked by /preregister for RCTs — the AEA RCT Registry and most IRBs require a power/MDE justification; /preregister's aea-rct style follows this skill's phases to fill that section (this skill is user-invoked only, so it is read and followed, not called).
  • During R&R — when a referee asks "was this study adequately powered to detect the effect you claim?"
  • Designing a Monte Carlo — to set R and sample sizes before handing off to /simulation-study.

Inputs

$ARGUMENTS may carry flags; missing pieces are elicited in Phase 0.

  • --mode mde|n|power — solve for MDE given N+power, N given MDE+power, or power given N+MDE. Default mde.
  • --design rct|cluster|multiarm|sim — two-arm RCT, clustered RCT (ICC), multiple arms, or simulation-based. Default inferred from the elicited design.
  • --input — a spec from /interview-me (under quality_reports/specs/) to pull the RQ, outcome, and design from.

Workflow

Phase 0 — Elicit the design

Gather the design parameters; ask once for anything missing rather than fabricating. Required:

  • Estimand & test: primary outcome, one- vs two-sided test, alpha (default 0.05), and whether the target is a difference in means, a proportion, or a regression coefficient.
  • Two of {effect size, N, power}: the effect as a raw difference and in standardized units (Cohen's d = effect / SD) — record both; power default 0.80.
  • Baseline mean and SD (or baseline proportion for a binary outcome) — needed to translate raw ↔ standardized effects.
  • Allocation: treated:control ratio (default 1:1; unequal allocation costs power — note it).
  • Clustering: if randomization is at a group level (village, school, clinic), the ICC (ρ), the average cluster size (m), and number of clusters. Compute the design effect DEFF = 1 + (m − 1)·ρ and the effective N.
  • Multiplicity: number of arms / primary outcomes; the correction (Bonferroni, Holm, or none) and whether power is per-comparison or familywise.

Echo a Pre-Flight Report (design, the two fixed quantities, the one being solved for, alpha, power, allocation, ICC/clusters, multiplicity) before computing. If the estimand or the SD source is ambiguous, stop and ask.

Phase 1 — Analytical power (standard designs)

For two-arm RCTs, clustered RCTs, and multi-arm comparisons, compute analytically. Prefer R pwr / WebPower (or a closed-form power.t.test / power.prop.test); for clustered designs inflate variance by DEFF, or use pwr on the effective N. Stata users: power twomeans / power twoproportions / power, cluster; Python: statsmodels.stats.power. Emit a short script to scripts/R/power_.R (or .do / .py) so the calc is reproducible, not a one-off console number.

  • MDE mode: MDE = (z{1−α/2} + z{1−β}) · SE(effect), where SE is built from the SD, N, allocation, and DEFF. Report MDE in raw and standardized units.
  • N mode: invert the above for total N (and #clusters when clustered) given the target MDE.
  • Power mode: given N and a hypothesized effect, return achieved power.
  • Multi-arm: divide alpha by the number of comparisons in the family m (Bonferroni alpha/m): m = K−1 for all-vs-control, m = K(K−1)/2 for all-pairwise. Report per-comparison and familywise power.

Sweep a grid (N or #clusters × effect size) so Phase 3 can draw a power curve and an MDE-vs-N curve.

Phase 2 — Simulation-based power (non-standard designs)

When the design is not a clean two-arm comparison — DiD / staggered event-study, IV / 2SLS, panel with serial correlation, a non-normal or censored outcome, or any estimator with no closed-form SE — switch to simulation. Reuse the /simulation-study harness exactly (see simulation-study and .claude/rules/simulation-conventions.md):

  1. Seeded, parameterized DGP that embeds the hypothesized effect (and the null DGP for size). set.seed(YYYYMMDD) once; L'Ecuyer streams if parallel.
  2. Estimator = the one you will actually use on the real data (e.g. fixest::feols two-way FE, did::att_gt, AER::ivreg), returning est, se, ci, p, reject.
  3. Power = share of reps rejecting H0 at alpha; size = rejection rate under the null DGP (verify it is near nominal before trusting power). Report each with its Monte Carlo SE = sqrt(p(1−p)/R).
  4. Sweep N (or #clusters / #periods) to trace the power curve; save the raw per-rep tibble via saveRDS() to output/.

A simulated power number without an MCSE, or without a verified size check, is not yet an answer.

Phase 3 — Write the power section

Produce the deliverables under quality_reports/power/:

  • power.md** — a table and a methods paragraph (below).
  • powercurve.png — power vs N (and/or MDE vs N), with reference lines at the target power and the design's planned N.
  • The reproducible script under scripts/R/ (or .do / .py).
# Power Analysis: <study title>
**Date:** YYYY-MM-DD · **Design:** <rct|cluster|multiarm|sim> · **Method:** <analytical|simulation, R/Stata/Python>

| Quantity | Value |
|---|---|
| alpha (sided) | 0.05 (two-sided) |
| Target power | 0.80 |
| Baseline mean (SD) | <m0> (<sd>) |
| Allocation (T:C) | 1:1 |
| ICC / cluster size / #clusters | <ρ> / <m> / <J>  (DEFF = <…>) |
| Total N (analysis sample) | <N> |
| **MDE (raw / standardized)** | **<Δ> / <d>** |
| Achieved power at planned N | <…>  (± MCSE <…> if simulated) |

## Methods paragraph (paste into preregistration)
> Assuming a baseline outcome mean of <m0> (SD <sd>), 1:1 allocation, and a two-sided
> test at α = 0.05, a total sample of <N> [<J> clusters of <m>, ICC = <ρ>] yields 80%
> power to detect a minimum effect of <Δ> (<d> SD). [Simulation: under the hypothesized
> DGP, <P>% of <R> replications rejected H0 (MCSE <…>); size under the null was <…>.]

Phase 4 — Handoff

If invoked by /preregister, return the methods paragraph + MDE row for the preregistration's power section. If standalone, print the save paths and remind the user the MDE is a design commitment to record before data collection.

Exit behavior

  • Computation succeeds: exit 0; print the MDE / N / power result, the save paths, and (for simulation mode) the size-check value next to power.
  • Under-identified design (only one of {effect, N, power} supplied) or ambiguous SD source: halt in Phase 0 with a single specific question — never guess the SD or the ICC.
  • Simulation size check fails (empirical size far from nominal under the null DGP): report power as UNRELIABLE and surface the size value; the estimator/DGP must be fixed before the power number is trustworthy.

Flags

  • --mode — What to solve for: minimum detectable effect, required N, or achieved power.
  • --design — Design family — two-arm RCT, clustered/ICC, multi-arm with corrections, or simulation-based for non-standard designs.
  • --input — Path to an /interview-me spec or preregistration draft to read design parameters from.

Cross-references

  • .claude/skills/preregister/SKILL.md — follows this skill to fill the power/MDE section of an aea-rct (and OSF) preregistration; this skill returns the methods paragraph.
  • .claude/skills/simulation-study/SKILL.md — the Monte Carlo harness Phase 2 reuses (seeded DGP, estimator grid, % rejecting H0).
  • .claude/rules/simulation-conventions.md — the simulation contract (truth from DGP, MCSE, size-under-the-null) that Phase 2 must honor.
  • .claude/skills/data-analysis/SKILL.md · .claude/skills/stata-replication/SKILL.md — where the realised analysis (and its actual estimator/SE) lives; the power calc should use the same estimator.
  • .claude/rules/confidential-data.md — when baseline mean/SD/ICC are taken from restricted-access data, disclosure-avoidance limits apply; cite published or pilot moments rather than embedding raw confidential statistics in the (externally-uploaded) preregistration.

What this skill does NOT do

  • Post-hoc / observed power. It refuses to compute "the power we had to detect our estimate" from a realised result — that is a deterministic function of the p-value and tells you nothing. Power is ex-ante only.
  • Pick your effect size for you. The MDE is your design commitment; the skill computes consequences of an assumed effect (from theory, a pilot, or a meta-analysis), it does not invent a plausible one.
  • Submit to a registry. Like /preregister, it writes a document; the user uploads it.
  • Replace /simulation-study. Phase 2 borrows the harness for a single power question; a full bias/RMSE/coverage study is /simulation-study's job.

More skills from pedrohcgs/claude-code-my-workflow

  • Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
  • Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
  • Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
  • Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
  • AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
  • AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
  • Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
  • AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
  • Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
  • Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
  • Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
  • Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).

All agent skills → · MCP servers