simulation-study skill
Scaffold and run a reproducible Monte Carlo simulation study in R — a declared assumption regime, a parameterized DGP, an estimator grid, a seeded replication loop, and a summary of bias, RMSE, empirical SE, coverage, size/power with Monte Carlo standard errors. Use when the user says "run a Monte Carlo simulation", "simulation study", "check the bias/coverage of an estimator", "compare estimators in simulation", "size and power simulation", "Monte Carlo experiment", or wants to demonstrate an estimator's finite-sample properties. Produces a numbered R script in `scripts/R/` and saves per-replication raw results + a summary table to `output/`.
Is the simulation-study skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the simulation-study skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow mkdir -p ~/.claude/skills cp -r /tmp/claude-code-my-workflow/.claude/skills/simulation-study ~/.claude/skills/simulation-study
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
/simulation-study — Monte Carlo Simulation Study
Design and run a Monte Carlo experiment that characterizes an estimator's finite-sample behavior, then review it for the bugs that quietly invalidate simulation evidence.
Input: $ARGUMENTS — a description of the estimator(s) and DGP to study (e.g., "compare 2SLS vs LIML under weak instruments with heteroskedasticity"), or a pointer to an existing script/paper whose simulation you want to reproduce or extend.
Constraints
- Follow .claude/rules/simulation-conventions.md — the simulation contract (DGP, truth, estimand, MCSE, assumption regime) is non-negotiable.
- Declare the assumption regime in the script header and respect the firewall — an out-of-assumption run never supports a within-assumption claim (simulation-conventions.md §2).
- Follow .claude/rules/r-code-conventions.md for general R standards (header, library() at top, relative paths, numerical discipline).
- Save the script to scripts/R/ with a numbered, descriptive name (e.g., scripts/R/sim2slsvs_liml.R).
- Save outputs (per-rep raw tibble, summary table, figures) to output/.
- saveRDS() the per-replication raw results, not just the summary — re-aggregation and the review pass need them.
- Run the sim-reviewer agent on the generated script before presenting results, then address Critical/High findings.
Workflow Phases
Phase 0: Pre-Flight Report
Before writing any code, produce a Pre-Flight Report showing you have pinned down the experiment. This prevents the most common failure mode — a beautiful results table built on a mismatched estimand or a coverage-against-the-estimate bug.
## Pre-Flight Report — Simulation Design
**Research question:** [what finite-sample property is being demonstrated]
**Target estimand:** [ATT / ATE / coefficient θ — and how its TRUE value is computed from the DGP params]
**Maintained assumptions:** [the FULL list the estimator(s) under study require — A1 … An, every one]
**Regime:** [IN-ASSUMPTION — all hold | OUT-OF-ASSUMPTION — relaxes A[k] only, severity grid {…}, targeting pseudo-estimand …]
**Verification:** [per assumption, the checkable property of the DGP that establishes it — by construction or by an assertion]
**DGP:** [structure + the parameters that define it; what is held fixed vs. varied]
**Estimator grid:** [list each estimator + which estimand it targets + how it returns est/se/CI]
**Design grid:** [sample sizes, parameter values, scenarios to sweep]
**Replications R:** [value] → implied MCSE on coverage ≈ sqrt(0.95·0.05/R) = [value]
**Metrics:** bias, empirical SE, RMSE, coverage, size/power — each with MCSE
**Conventions read:** simulation-conventions.md, r-code-conventions.mdIf the estimand or its true value is ambiguous, stop and ask before writing code.
If an assumption cannot be verified — you cannot name the property of the DGP that establishes it — the run is not IN-ASSUMPTION, and per the firewall no within-assumption claim may rest on it. Say so in the Pre-Flight Report rather than letting the header assert what was never checked.
Phase 1: The DGP
Write one parameterized function that returns a dataset. Compute and return (or store) the true target value from the parameters.
generate_data <- function(n, params) {
# ... generate covariates, treatment, outcome from params ...
list(data = df, truth = compute_truth(params)) # truth from params, never from an estimate
}The header's Verified lines are earned here: every assumption the regime block claims holds by a check gets that check written into the script (a large-draw assertion, a condition number, a stopifnot() on the parameter bounds), run once at setup. An assumption whose verification exists only in the comment is asserted, not verified.
Phase 2: Estimator Grid
Each estimator is a function data -> list(est, se, cilo, cihi, converged). State the estimand each one targets; an estimator scored against a mismatched truth is a bug, not a finding.
Phase 3: Replication Engine
- set.seed(YYYYMMDD) once. For parallel reps use RNGkind("L'Ecuyer-CMRG") and furrr::furrr_options(seed = TRUE).
- One run = generate data → run every estimator → record a row per estimator with est, se, cilo, cihi, converged.
- Pre-allocate / bind results into a tibble of R × (#estimators) rows. Track non-convergence; never silently drop.
Phase 4: Metrics & Summary
Per estimator × scenario, against truth:
- Bias = mean(est) - truth (+ MCSE = sd(est)/sqrt(R))
- Empirical SE = sd(est); RMSE = sqrt(mean((est - truth)^2))
- Coverage = mean(cilo <= truth & truth <= cihi) (+ MCSE = sqrt(p(1-p)/R))
- Size / power = rejection rate under the null / alternative DGP
- Failures = count of non-converged reps
Build a tidy summary table; report MCSE next to every headline metric.
Phase 5: Figures
Use ggplot2 with the project theme: bias / coverage vs. sample size (or scenario), with reference lines (0 bias, nominal coverage). Transparent background, explicit dimensions (per r-code-conventions.md §4).
Phase 6: Save & Review
- saveRDS() the raw per-rep tibble and the summary table to output/; also write the summary as .csv/.tex.
- Run the review:
Delegate to the sim-reviewer agent:
"Review the simulation script at scripts/R/[name].R"The agent is read-only and returns its report; save it to qualityreports/[name]sim_review.md.
- Address Critical/High findings (coverage-vs-truth, estimand mismatch, missing MCSE, dropped reps, an unstated or unverified regime) before presenting.
- Apply the firewall to the presentation itself. Every claim you are about to make must cite a run whose regime can bear it — consistency, valid analytic standard errors, nominal coverage, and shipping a default require an IN-ASSUMPTION run and nothing else (simulation-conventions.md §2). Carry the regime in every caption — and per row wherever a severity grid mixes the two.
Script Structure
# ============================================================
# [Title] — Monte Carlo simulation
# Author: [project context]
# Purpose: [property being demonstrated]
# Estimand: [target + how truth is computed]
# Maintained assumptions: [A1 ... An — the FULL list the estimator requires]
# Regime: [IN-ASSUMPTION | OUT-OF-ASSUMPTION: relaxes A[k] only, severity ...,
# targeting pseudo-estimand ...]
# Verified: [per assumption, the property that was actually checked, not asserted]
# Outputs: output/[name]_raw.rds, [name]_summary.{rds,csv}
# ============================================================
# 0. Setup ----
library(tidyverse)
library(furrr) # parallel reps (optional)
plan(multisession) # enable parallel workers; omit this line to run sequentially
RNGkind("L'Ecuyer-CMRG")
set.seed(20260531) # once, YYYYMMDD (simulation-conventions.md §3)
R <- 2000L # MCSE on coverage near .95 ≈ 0.005
dir.create("output", recursive = TRUE, showWarnings = FALSE)
# 1. DGP ----
generate_data <- function(n, params) { ... } # returns list(data, truth)
# 2. Estimators ----
estimators <- list(tsls = est_tsls, liml = est_liml) # each -> est, sImportant
- State the regime, then respect the firewall. A DGP that violates a maintained assumption produces a table on which every other check here passes — and a within-assumption claim built on it is unsupported however clean the numbers look.
- The truth comes from the DGP, never from an estimate. Coverage is the CI containing the true parameter.
- No result without an MCSE. If two estimators differ by less than ~2× MCSE, say so.
- Save raw, not just summary. A number that exists only in the console cannot be audited or put on a slide.
- Count your failures. Silently dropped non-converged reps bias every metric.
Long-running simulations: use the Monitor tool
Large grids (many scenarios × large R) can run for many minutes. Background-launch via Bash with runinbackground: true, writing R stdout and stderr to a log (e.g. Rscript scripts/R/[name].R > output/[name].log 2>&1), and run the Monitor tool with a command that tails that log through grep --line-buffered, matching progress milestones (e.g. a progressr update) and failure signatures (Error, Execution halted), instead of polling with sleep. Monitor has no job-id parameter: the stdout of its own command is the event stream. See data-analysis/SKILL.md and the guide's Cost-Conscious Parallelism section.
More skills from pedrohcgs/claude-code-my-workflow
- Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
- Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
- Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
- Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
- AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
- AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
- Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
- AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
- Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
- Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
- Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
- Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).