Mmcp.market

challenge skill

by pedrohcgs·pedrohcgs/claude-code-my-workflow·1.6k stars·MIT

Stress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.

A100/100content scan

Is the challenge skill safe?

Clean: nothing in its files matched our rules. We read 4 files in the folder on 2026-09-28.

No findings.

Install the challenge skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow
mkdir -p ~/.claude/skills
cp -r /tmp/claude-code-my-workflow/.claude/skills/challenge ~/.claude/skills/challenge
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Challenge — does the result survive the choices you didn't make?

A single specification is one draw from a distribution you never looked at.

Why this exists, measured rather than asserted. In a controlled study, 150 autonomous agents were given the same data and the same questions. Effect-size interquartile ranges reached ~10.7 %/yr, and the spread concentrated in discrete measure-choice forks — not in estimation noise. Within a measure family, agents agreed to ~0.25 %/yr. Two findings from that study shape this skill:

not reduce analytical-choice variance. A clean referee report is not robustness.

  • AI peer review left the spread essentially unchanged. Review catches errors; it does

by correctness.** Herding is not agreement.

  • Exposure to exemplar papers collapsed the spread by 80–99 % — **convergence by imitation, not

So the spread has to be measured, not reviewed away.

Preconditions

2013–2019, on the treated population". If you cannot state it, stop: you cannot challenge a claim you have not defined.

  • A working baseline specification that runs and produces the headline estimate.
  • The estimate's estimand stated in words — "the ATT for units treated in 2015, over

faster than intuition.

  • A fork budget (--forks, default 64). Grid size is the product of your choices; it grows

Step 1 — Enumerate the forks, before running anything

List every point where a competent, honest analyst could have chosen differently. Do this before seeing any alternative result, and write it down — the list is the pre-registration of the challenge.

Some forks change the estimand, not just the estimate — averaging over them is

meaningless. references/fork-catalog.md labels every fork;

record estimand forks separately and say so in the report.

Ship --dry-run first. Print the grid size and an estimated runtime before executing anything. A 6-fork grid with 3 options each is 729 fits.

Step 2 — Run the grid

One fit per cell, same seed, same data build. Persist every cell — including failures. A specification that does not converge is information about fragility, not a cell to drop.

Record per cell: the fork coordinates, the point estimate, the standard error, N, and the convergence status.

Step 3 — Report the distribution, not the winner

reader can see which choices move the result.

  • Specification curve: estimates sorted, with the fork coordinates shown underneath so a

levels. Report both; they answer different questions.

  • The share of specifications with the same sign, and the share significant at conventional

choice between dollar and share volume" is a far more useful sentence than a robustness paragraph.

  • Which fork drives the spread. This is the payload. "The result is robust except to the

percentile of specifications you yourself called defensible, say so.

  • Your baseline's percentile in its own distribution. If the headline sits at the 97th

The descriptive curve is not a test — read it as a description of fragility. If you need

inference over the whole curve, use specification-curve analysis's joint permutation test

(Simonsohn, Simmons & Nelson 2020), which supplies the sharp null the picture alone lacks.

Step 4 — Attack the identifying assumption

The grid varies what you can vary. The identifying assumption is what you cannot test — so bound it instead, with a named, computable statistic. See references/sensitivity-statistics.md.

Rows for the causal-identification designs are deliberately absent (unvetted-methods

veto): populate them from your field's canonical sources after vetting.

Label every statistic executable-here or describe-and-cite. Honesty about what your environment can actually run is itself a verification step; a cited-but-unrun statistic is not evidence.

Step 5 — Placebo and falsification

Where a falsification test exists, run it: a negative-control outcome that should show nothing, a negative-control exposure, a timing placebo. A passed placebo is weak positive evidence; a failed placebo is strong negative evidence. Report both with equal prominence.

Step 6 — Write the ledger entry

Append to the specification-search ledger (verification-ladder.md rung 5): the fork list, the grid size, the distribution summary, which forks moved the result, the sensitivity statistics with their values, and every attempt including the failures.

Pre-commit the interpretation. Before running the grid, write down what result would

SUPPORT and what would WEAKEN the claim. The ledger is the arbiter. A robustness exercise

interpreted after the fact is not a robustness exercise.

Anti-patterns

prevent. Fixed fork list, fixed budget, stated stopping rule.

  • Running until something interesting appears. That is the pathology this skill exists to

fishing expedition with better graphics.

  • Reporting only the specifications that agree — a curve showing only supporting cells is a
  • Treating a wide curve as failure. Wide is a finding. Publish it and say what drives it.
  • Adding forks nobody would defend to pad the denominator and dilute the fragile cells.

Reference files

Cross-references

  • verification-ladder.md — rung 4 (analytic verification) and rung 5 (the ledger)
  • /simulation-study — when the question is finite-sample performance, not robustness
  • /preregister — reserve a holdout before the search, not after

More skills from pedrohcgs/claude-code-my-workflow

  • Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
  • Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
  • Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
  • Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
  • AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
  • Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
  • AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
  • Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
  • Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
  • Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
  • Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).
  • Acredible-claimsResearch-brief + claim-record discipline for delegated or AI-assisted research work. Use when starting any substantive research task or long autonomous run (write the brief first), and when reporting results that will support a claim in a paper or decision (produce the claim record). Keeps faster execution from being confused with credible evidence.

All agent skills → · MCP servers