diagnose skill
Root-cause a failing or wrong empirical result with a disciplined reproduce → minimise → hypothesise → instrument → fix loop, instead of guessing-and-poking. Use when the user says "why is my regression wrong", "this number changed", "my script errors out", "the result won't reproduce", "debug this", "this estimate looks wrong", or "it worked yesterday". Tuned for research code (R/Stata/Python): type coercion, NA/merge blow-ups, factor levels, clustering/SE choices, weighting, collinearity/convergence, seeds, package-version drift. Use `--no-fix` to localize the root cause without editing shared or load-bearing files.
Is the diagnose skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the diagnose skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow mkdir -p ~/.claude/skills cp -r /tmp/claude-code-my-workflow/.claude/skills/diagnose ~/.claude/skills/diagnose
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
/diagnose — Root-Cause a Wrong or Failing Result
Find why an analysis errors, returns the wrong number, or won't reconcile — with a structured debugging loop rather than scattershot edits. Adapted from the diagnose pattern in mattpocock/skills, reshaped for empirical research code where the bug is usually a silent wrong number, not a crash.
The discipline: never edit before you can reproduce, and never fix before you can explain. A guessed fix that makes the symptom disappear without a named root cause is how a wrong number gets laundered into a published table.
When to use
- A regression / estimate returns a value you can't explain, or one that changed when nothing should have.
- A script errors out and the stack trace doesn't point at the real cause.
- A result "won't reproduce" — different number on re-run, on another machine, or after a package update.
- A replication claim fails /audit-reproducibility and you need to localize which step drifted.
Diagnose is symptom-driven and single-target: ONE wrong number / ONE failing run. Use a sibling instead when the job is different:
- /audit-reproducibility — verify all numeric claims in a manuscript against current code (claim-driven, whole-paper). If you have one FAILing claim and want to localize which pipeline step produced it, /audit-reproducibility hands off to /diagnose; if you want to re-check every table number, start there.
- /review-r — code-quality review with no specific symptom.
- /capture-environment — snapshot the environment when version/seed drift is the suspect.
Phases
Phase 0 — Pin the symptom (expected vs. actual)
State the bug as a falsifiable gap before touching anything:
- Expected: the value/behaviour you believe is correct, and why (a prior run, a paper table, a hand calculation, a theoretical sign).
- Actual: the value/error observed now, copied verbatim (full message, not a paraphrase).
- Tolerance: the threshold that separates "same" from "different", keyed to the source of expected — prior run on the same machine → machine-epsilon + display rounding; a published table → rounding + small slack (~1e-3); a hand calculation → ~0.01; a theoretical prediction → an economic-significance band, not a decimal. Don't chase 1e-12 floating-point noise; don't wave away a 5% gap. (See replication-protocol.md.)
If expected/actual can't be stated, the task is understanding, not diagnosis — stop and clarify first.
Phase 1 — Reproduce deterministically (get a reliable red)
A bug you can't reproduce on demand can't be fixed, only hidden.
- Fix every source of nondeterminism: set the seed, pin the working directory, record sessionInfo() / pip freeze / Stata version (lean on /capture-environment).
- Re-run the smallest unit that exhibits the bug and confirm it fails every time. An intermittent failure is its own hypothesis (uninitialised RNG, order-dependent merge, race in parallel code) — note it and carry it into Phase 3.
Phase 2 — Minimise to an MWE
Shrink until the bug sits in the open:
- Data: subset to the smallest rows/columns that still reproduce (often one group, one period, a handful of rows).
- Code: strip the pipeline to the shortest path from input to wrong output; comment out everything the symptom survives without.
- Each removal that keeps the bug is information; each that kills it is a stronger signal — record which.
The MWE is the deliverable even if the fix is later trivial: it's what makes the root cause undeniable.
Phase 3 — Hypothesise (enumerate, then rank)
List candidate causes before testing any — a written list beats poking because it prevents fixating on the first idea. For research code, walk the usual suspects (all of these run cleanly with no error message — they are silent-wrong-number bugs):
- Types & coercion — a numeric read as character/factor; integer overflow; date parsed wrong; TRUE/FALSE ↔ 1/0.
- Missingness — NA dropped silently, na.rm flipping a mean, listwise deletion changing the sample mid-pipeline.
- Joins & shape — a many-to-many merge inflating rows; duplicate keys; an unbalanced panel where balance was assumed.
- Specification — wrong clustering level, fixed effects absorbed twice, a lag/lead off by one.
- Bad controls & colliders — a control that is post-treatment, a mediator on the causal path, or a descendant of treatment (adding it induces bias, invisibly). The tell: a coefficient that moves the "wrong way" or shrinks implausibly when a control enters.
- Numerical stability & convergence — an optimizer that didn't converge (check the convergence code, not just the estimates), a singular/near-singular Hessian, collinearity (high VIF, a dropped column), tolerance set too loose, under/overflow with very small/large weights or coefficients.
- Weighting & aggregation — weights silently dropped/truncated, weights renormalised wrong, frequency vs. probability vs. analytic weights confused, a weight applied after rather than before a transform.
- Sample — a filter that runs before vs. after a transform; an outlier rule applied inconsistently.
- Environment — a package/Stata version bump that changed a default; a seed that moved; locale/encoding.
For a genuinely ambiguous bug, fan out the top competing hypotheses to parallel Agent subagents (one per hypothesis, each in a fresh context), each instructed to try to confirm its own cause on the MWE and report back — the loop-first analogue of asking three colleagues at once (see orchestrator-protocol.md).
Phase 3b — Reduce the hypotheses (so you don't launder a guess)
Each hypothesis (whether tested by hand or by a fan-out Task) returns {hypothesis, evidence for, evidence against, confidence, one-line conclusion}. Then:
- One clear winner (high confidence, others refuted) → proceed to Phase 4 to confirm the mechanism.
- A near-tie (top two within ~20 percentage points) → do not pick one; go to Phase 4 instrumentation to discriminate.
- None above ~50% → report ambiguity and ask the user; do not edit on a coin-flip.
Phase 4 — Instrument & localize (bisect, don't stare)
Test the ranked hypotheses cheaply:
- Bisect the pipeline — check the intermediate value at the midpoint of the data flow; the bug is upstream or downstream of it. Repeat. Binary search finds the offending line in log2(n) steps, not n.
- Bisect history — if it "worked yesterday", compare against the last-good commit/output to pin the change that introduced it. (git bisect is fine here — it never discards work; the destructive git commands are blocked by git-guardrails.py, this is not one of them.)
- Instrument with diagnostic primitives, not guesses — at each stage inspect: str() / summary() for types & NA patterns; row & column counts before and after every transform; table(factor) to catch a silently dropped level; cor() / VIF for unexpected collinearity; weight diagnostics range(w), sum(w), table(is.na(w)); and the regression's convergence flag. The stage where a count drops unexpectedly, a factor level vanishes, correlation jumps, or weights go sparse is the culprit stage.
End Phase 4 with a one-sentence root cause naming the exact line/step and mechanism.
Phase 5 — Fix & verify (then guard against regression)
Confidence gate (the anti-laundering rule): do not apply a fix unless the root cause is named and its mechanism is explicit. If Phase 3b left a near-tie, behave as --no-fix: report the candidates and ask. Editing research code on an unproven hypothesis is exactly the laundering this skill exists to prevent.
Unless --no-fix is set:
- Apply the minimal fix at the root cause — not a downstream patch that masks it (prefer fixing the bad merge over filtering its duplicate rows afterward).
- Re-run the MWE → confirm actual == expected within the Phase-0 tolerance.
- Re-run the full unit and any dependent step → confirm the fix didn't move another number. If the result feeds a manuscript claim, re-check it (cross-ref the passport in /audit-reproducibility).
- Note a prevention — the assertion/check that would have caught this earlier. One concrete guard per bug class:
Propose the guard; don't silently install a test suite.
With --no-fix, stop after the root cause is named and report it for the user to fix by hand.
Worked example
A demand-forecasting model's held-out MAE jumped from 0.043 to 0.071 after a data refresh; nothing in the spec changed.
# Phase 1 — reproduce: set.seed(1); same script, same number every run. Red is stable.
# Phase 2 — MWE: one region, two horizons still shows the jump.
# Strip to: read panel -> merge features -> lm(). Bug survives the merge step.
# Phase 4 — instrument: row counts before/after each step
nrow(panel) # 12,400 (expected)
nrow(merge(panel, feats, by="id")) # 12,933 <-- inflated! a many-to-many merge
# Root cause: the refresh left duplicate feats rows for a subset of ids; the
# join fans those ids out, 12,400 -> 12,933 (+533 rows), re-weighting the MAE
# toward the duplicated units.
# Phase 5 — minimal fix at the root (dedup the key), NOT a downstream row filter:
feats <- feats[!duplicated(feats$id), ]
# re-run: MAE back to 0.043 within tolerance; full pipeline re-checked, no other number moved.
# Prevention (Joins & shape guard):
stopifnot(nrow(merge(panel, feats, by = "id")) == nrow(panel))Output / report format
Write a short diagnosis to qualityreports/diagnoses/YYYY-MM-DD.md (create the directory first: mkdir -p qualityreports/diagnoses). These reports may contain real data values and file paths — they are project-internal and gitignored**, like session logs. Include:
- Symptom: expected vs. actual (+ tolerance).
- MWE: the minimal input/code that reproduces it.
- Root cause: the exact line/step and mechanism.
- Fix: the diff applied (or, with --no-fix, the recommended change).
- Verification: MWE + full-run re-check results.
- Prevention: the guard that would have caught it.
Plus a chat summary leading with the one-line root cause.
Cross-language notes
The usual-suspects model is illustrated in R but the bug classes are language-neutral; the diagnostic idioms differ:
- R — anyNA() / table(is.na(x)); factors silently drop unused levels; set.seed(); sessionInfo().
- Stata — tab v, missing and explicit ./.a–.z extended missing; set seed; version; weights as [fw=] vs [pw=] vs [aw=] is a frequent silent bug.
- Python — df.isnull().sum(); numpy.nan ≠ None; pandas vs numpy NaN handling differ; np.random.seed() / a passed random_state; pip freeze.
(Forkers in other fields: the five structural classes — Types, Missingness, Joins, Sample, Environment — are discipline-neutral; the econometric suspects above are the worked instance.)
Exit behavior
Flags
- --no-fix — Diagnose only: run through naming the root cause (Phases 0–4) and write the report, but make no edit to source. Use when you want to apply the fix yourself, or when the file is shared/load-bearing and an automated edit is inappropriate.
Step 0 — State the contract before touching anything
Before proposing or writing any fix, state in 3–5 bullets and stop for confirmation:
not a bug in the callee.
- What the function's or pipeline's documented contract is.
- What the reporter claims is broken.
- Whether the reported scenario is even in contract — a caller violating the contract is
- What the correct behaviour would be.
- What evidence would settle it.
This costs thirty seconds and prevents the most expensive class of wasted work: a confident fix to a misunderstood contract. In one logged case the same semantic point had to be corrected twice before a fix was scoped right, because the agent asserted a stance on the contract rather than restating it and asking.
Severity follows the contract. A scenario outside the documented contract is at most a
documentation or validation issue, never a high-severity correctness bug. Inflating it
because it looks wrong is how a fix ends up changing behaviour users depend on.
Separate the audit pass from the repair pass
More skills from pedrohcgs/claude-code-my-workflow
- Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
- Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
- Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
- Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
- AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
- AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
- Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
- AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
- Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
- Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
- Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
- Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).