Mmcp.market

audit-reproducibility skill

by pedrohcgs·pedrohcgs/claude-code-my-workflow·1.6k stars·MIT

Enforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.

A100/100content scan

Is the audit-reproducibility skill safe?

Clean: nothing in its files matched our rules. We read 4 files in the folder on 2026-09-28.

No findings.

Install the audit-reproducibility skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow
mkdir -p ~/.claude/skills
cp -r /tmp/claude-code-my-workflow/.claude/skills/audit-reproducibility ~/.claude/skills/audit-reproducibility
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Audit Reproducibility

Compare numeric claims in a manuscript (point estimates, standard errors, p-values, counts) against the actual outputs produced by the analysis pipeline. Report PASS / FAIL per claim against the tolerance thresholds defined in .claude/rules/replication-protocol.md.

Core principle: If the paper says ATT = -1.632 (0.584) and the code produces -1.628 (0.591), we verify — numerically — that the difference is within the documented tolerance. No more "looks close enough" eyeballing.

Two directions, not one. Vertically, each claim is checked against the output that produced it. Horizontally, it is checked against every other artifact that displays the same number — the supplement table, the slide deck, the poster. The vertical check passes contentedly while a deck quotes last month's value; only the horizontal one catches that. Declared displays live in the passport's appears_in list — see replication-protocol.md → The horizontal check.

When to use

  • Before submission. Catches the "I updated the analysis but forgot to update Table 2" bug.
  • Before presenting or teaching from the same numbers. Catches the deck, the poster, or the supplement that was never regenerated after the last rerun.
  • Before releasing a replication package. Verifies the code actually reproduces the paper.
  • After a major revision. Ensures the paper still matches the latest code.
  • Quality-gate in /commit. Pair with a pre-commit invocation on manuscript + analysis changes.

Inputs

  • $0 — path to the manuscript (.tex, .qmd, .md, .pdf). Required.
  • $1 — path to the outputs directory. Defaults to output/, where every language's pipeline writes (R, Stata via /stata-replication, Python). Recognised alternatives: targets/objects/ (R targets workflows), any directory the user-specified outputs live in. If output/ does not exist but a pre-v2.6 scripts//outputs/ does, use that and say so in the report.

Workflow

Phase 0: Pre-flight

  1. Read replication-protocol.md for the tolerance thresholds currently in effect.
  2. Verify the outputs directory exists and is non-empty. If empty or stale (older than the manuscript), prompt the user to re-run their pipeline (e.g., Rscript scripts/R/00runall.R) before auditing.
  3. Ensure a sessionInfo.txt or equivalent environment capture exists in the outputs dir.

Phase 1: Extract claims from the manuscript

Parse the manuscript for numeric claims. Patterns to match:

  • Point-estimate + SE: ATT = -1.632 (0.584), $\beta = 0.342$ (0.091), hat{\tau} = 1.28** with starred significance
  • Table cells: & -1.632$^{***}$ & 0.584 & in LaTeX table environments
  • Counts: our sample of 2,847 firms, $N = 2{,}847$
  • Summary stats: mean = 0.423, SD = 0.087
  • P-values: p < 0.01, $p = 0.003$

Record each claim as a tuple:

{
  claim_id: "Table2_col3_ATT",
  location: "Table 2, Column 3, row 'Treatment'",
  kind: "point_estimate" | "standard_error" | "p_value" | "count" | "percentage",
  reported_value: -1.632,
  uncertainty: 0.584,              # only for point estimates
  significance_stars: 3,            # 0-3 or None
  raw_context: "the ATT estimate of -1.632 (0.584) indicates..."
}

Write the extracted claims to qualityreports/reproducibilityclaims_[manuscript-name].json so the user can review the extraction before audit.

Phase 2: Extract results from outputs

Scan $1 for corresponding values. Priority order:

  1. .rds files — readRDS(path)$coef[["treatment"]] style lookups. Can use Rscript -e "saveRDS(summary(readRDS(...)), '/tmp/audit.rds')" to extract.
  2. .tex tables — parse LaTeX table cells directly; match on column headers + row labels.
  3. .csv summary files — pandas/readr parse, key-value lookup.
  4. .out / .log files (Stata, regress output) — regex extraction.
  5. .json — direct key lookup.

Record each extracted result:

{
  source: "output/results.rds",
  lookup_key: "fit_main$coefficients['treated']",
  value: -1.628,
  uncertainty: 0.591,
  p_value: 0.005
}

Phase 2b: Collect every declared display (passport mode)

For each passport claim, read its appearsin list and pull the value as displayed** at each entry:

{
  claim_id: "C3",
  displays: [
    { path: "manuscript.tex", locator: "Table 1, Col 2", display_precision: 3, shown: 0.342 },
    { path: "Slides/Lecture04_Results.tex", locator: "frame 'Main result'", display_precision: 2, shown: 0.34 }
  ]
}

Rules for this pass:

  • The primary location: is one of the displays, not a separate thing — if appears_in repeats it, that is one display, not two.
  • A declared display that cannot be resolved — missing file, or no value at the locator — is recorded as shown: NOTFOUND.** It resolves to FAIL in Phase 4c. Never skip it: a locator that has quietly stopped matching is exactly where a stale number hides.
  • A claim with no appearsin list is horizontally unchecked*, not horizontally clean. Report it that way (see Phase 5) rather than silently passing it.

Phase 3: Match claims to results

Use fuzzy heuristics when exact labels don't match:

  • Name similarity ("treatment effect" ~ "ATT" ~ "treated")
  • Magnitude similarity (if two candidates have values within 10% of the reported, prefer the one with closer SE)
  • Context hints from the claim's raw_context field (table number, row label, description)

For every claim, produce a match candidate with a confidence score. Claims below 0.7 confidence get flagged as "UNMATCHED — manual review needed" rather than silently passing.

Phase 4: Tolerance check

For each matched claim, apply the thresholds from replication-protocol.md:

Respect any tolerance overrides the user has written into their replication-protocol.md fork (they may loosen for MC noise or tighten for administrative data).

Phase 4b: Disposition — PASS / FAIL / EXPLAINED / UNMATCHED

A tolerance check resolves to one of four dispositions:

  • PASS — within tolerance.
  • FAIL — outside tolerance, with no defensible alternative recorded. Blocks (exit 1).
  • EXPLAINED — outside tolerance, but the author has recorded a concrete, named alternative specification that accounts for the gap (see the downgrade rule). Surfaced in the report and carried into a response-to-referees; does not block.
  • UNMATCHED — no computed counterpart found (Phase 3 confidence < 0.7). Never auto-downgradable.

A mismatch is not automatically a failure. In applied work the most common out-of-tolerance result is a defensible alternative spec, not a bug — reghdfe vs feols clustering df, a different bandwidth-selection rule, a different MC seed/reps, or display rounding. The skill's job is to stage the disagreement for a human auditor, not to pronounce the code right and the paper wrong. (The df-adjustment note in "Stata-specific notes" below is the canonical example of a named alternative.)

The manuscript is not the oracle. When the computed value disagrees with the manuscript, do not presume the code is correct and the paper stale — nor the reverse. A refactor may have broken a previously-correct table (the on-disk output is the buggy one), or the paper may carry an old number. The computed value is a challenger, not ground truth. Report a mismatch as "one of {paper, code} must change — isolate which," never "revert the code to match the paper." This prevents the trap of reverting a genuine bug-fix just to make the paper 'reproduce.'

Downgrade rule: FAIL → EXPLAINED

A FAIL may be downgraded to EXPLAINED only when a specific named alternative is recorded for that exact claim — in the passport entry's notes: field (passport mode) or the audit report's author-note column (default mode). Example of a valid note:

"reghdfe vs feols clustering-df adjustment; under the reghdfe small-sample correction the published value is −1.19, within rounding of the script's −1.187. CODE-CORRECTED pending."

The author is the auditor: the skill stages the two-sided comparison (reported value and computed value, both shown); the human writes the one-line named alternative; the skill records it and thereafter respects it. Tag the resolution PAPER-CORRECTED, CODE-CORRECTED, or DEFENSIBLE-ALTERNATIVE.

Hard floor — never downgradable to EXPLAINED:

  • A blank note, "unclear", "looks fine", or any note that does not name a concrete alternative spec.
  • An UNMATCHED claim (no computed counterpart to compare against).
  • A flat numerical contradiction with no alternative offered.

(Citation/existence claims are out of scope here — /verify-claims owns those, and applies the same named-alternative softening on its side.)

Repeated EXPLAINED is a signal (two-strikes)

Reuse the two-strikes rule from review-paper --adversarial and summary-parity.md: if the same claim is downgraded to EXPLAINED in two consecutive audits without ever being corrected to PASS (the author keeps invoking the alternative but never updates paper or code), stop treating it as quietly resolved. Surface it prominently in Phase 5 — "this contested number has been EXPLAINED twice but never corrected" — so a standing disagreement can't hide behind a recorded note indefinitely. In passport mode, detect this by comparing the current status/notes against the prior audit's.

Phase 4c: Horizontal check — the displays must agree with each other

Phase 4 compared the claim to the code. This phase compares the claim's displays to each other, pairwise over the displays collected in Phase 2b:

  1. Round both sides to the coarser precision. Two displays agree when they are equal after both are rounded to the smaller of their two display_precision values. 0.34 in a deck against 0.342 in the paper agrees at 2 decimals; 0.29 against 0.342 does not.
  2. No tolerance applies. The tolerance: block governs the vertical comparison, where two measurements are compared. Two displays of one number are two copies — after the rounding in step 1 they either match or they do not.
  3. Dispositions, in the vocabulary of Phase 4b:
  • PASS — every pair agrees.
  • FAIL — any pair disagrees, or any display is NOTFOUND. Never downgradable to EXPLAINED: a named alternative spec explains why code and paper differ and has nothing to say about why two copies of one number differ. If a display genuinely shows a different* specification, it is a different claim and belongs in its own passport entry.
  • UNCHECKED — the claim declares no appears_in list. Reported, non-blocking; the fix is to declare the displays.
  1. Report which side moved, not which side is right. Name the artifact whose value differs from the rest, and when it was last touched; the author decides whether the deck is stale or the paper is. Do not hand-edit one display to match its sibling — re-derive both from the output the claim points at (replication-protocol.md → Anti-patterns).

Phase 5: Report

Write qualityreports/reproducibilityaudit_[manuscript-name].md:

# Reproducibility Audit: [Manuscript Title]

**Date:** [YYYY-MM-DD]
**Manuscript:** [path]
**Outputs directory:** [path]
**Tolerance source:** .claude/rules/replication-protocol.md

## Summary

| Status | Count |
|---|---|
| PASS (vertical and horizontal) | N |
| FAIL (diff > tolerance, no named alternative) | M |
| EXPLAINED (out of tolerance, named alternative recorded) | E |
| UNMATCHED (manual review) | K |
| HORIZONTAL DRIFT (declared displays disagree, or a display not found) | H |
| HORIZONTALLY UNCHECKED (claim declares no `appears_in`) | U |
| **Overall verdict** | **PASS / FAIL** (FAIL iff M > 0 or H > 0; EXPLAINED does not fail the audit) |

## PASS (all within tolerance)
| Claim | Reported | Computed | Diff | Tolerance |
|---|---|---|---|---|
| Table2_col3_ATT | -1.632 (0.584) | -1.628 (0.591) | 0.004 / 0.007 | 0.01 / 0.05 |

## FAIL (outside tolerance — BLOCKER)
| Claim | Reported | Computed | Diff | Tolerance | Location in paper | Author note (name a concrete alternative to downgrade → EXPLAINED) |
|---|---|---|---|---|---|---|

## EXPLAINED (out of tolerance; defensible named alternative recorded — non-blocking, carry into response-to-referees)
| Claim | Reported | C

Phase 5b: Append to the replication log

The report above and the passport are rewritten on every run, so neither keeps a history. The log does. After every run — default and passport mode alike — append one block to quality_reports/replication-log.md. It is committed, so a co-author or data editor can see what was checked, against which commit, and how each number was obtained, without reading the code.

  • Append only. Write with >>; never edit or reorder a past block. A correction is a new run, not an edit. The repo-hygiene gate fails a commit that edits or removes a committed line — at the pre-commit hook against the last commit, and in CI against the branch the work merges into. It proves no entry was edited, not that every run was logged. A disclosure redaction is the one exception: commit it with ALLOWLOGREWRITE=1 and the reason in the commit message.
  • Stamp the revision. The short commit hash, plus -dirty when anything outside quality_reports/ is uncommitted (untracked files included) — so a verdict on uncommitted code says so, and the audit's own report files do not trigger it. Outside git, or before the first commit, the stamp is no-commit.
  • "How computed" is the exact accessor or command Phase 2 used — enough for someone else to recompute that one number.
  • Restricted data: the log carries the same numbers as the report, so it is committed only after disclosure clearance, like any output (confidential-data.md).
LOG=quality_reports/replication-log.md
[ -f "$LOG" ] || printf '%s\n' "# Replication log" "" \
  "Append-only record of /audit-reproducibility runs: what was checked, against which commit, and how each number was computed. Never edit a past entry; a correction is a new run." "" > "$LOG"
# -dirty = anything uncommitted outside quality_reports/, untracked files included (the audit writes its own files in quality_reports/)
if REV=$(git rev-parse --short HEAD 2>/dev/null); then
  [ -n "$(git status --porcelain -- . ':(exclude)quality_reports' 2>/dev/null)" ] && REV="$REV-dirty"
else
  REV="no-commit"   # not a git repository, or nothing committed yet
fi
printf '## %s — %s @ %s\n\n' "$(date +%F)" "<manuscript path>" "$REV" >> "$LOG"
# Quoted heredoc: nothing below is expanded, so an accessor's `$` or backtick is written as-is.
cat >> "$LOG" <<'EOF'
Outputs: `<outputs dir>` · Verdict: **<PASS|FAIL>** (<M> FAIL, <H> horizontal drift, <E> EXPLAINED, <K> unmatched)

| Claim | Location | Reported | Computed | How computed | Tolerance | Verdict |
|---|---|---|---|---|---|---|
| Table2_col3_ATT | main.tex: Table 2, col 3 | -1.632 | -1.628 | `readRDS("output/results.rds")$coef[["treatment"]]` |

More skills from pedrohcgs/claude-code-my-workflow

  • Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
  • Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
  • Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
  • AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
  • AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
  • Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
  • AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
  • Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
  • Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
  • Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
  • Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).
  • Acredible-claimsResearch-brief + claim-record discipline for delegated or AI-assisted research work. Use when starting any substantive research task or long autonomous run (write the brief first), and when reporting results that will support a claim in a paper or decision (produce the claim record). Keeps faster execution from being confused with credible evidence.

All agent skills → · MCP servers