Mmcp.market

deep-audit skill

by pedrohcgs·pedrohcgs/claude-code-my-workflow·1.6k stars·MIT

Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects, adjudicate every finding with a separate judge, fix all confirmed defects, then re-verify. Use when correctness must be bulletproof and single-pass or round-by-round review is too slow and too shallow. Invoke for "audit this rigorously", "find ALL the bugs/gaps", "make this rock solid", "converge faster on correctness".

A100/100content scan

Is the deep-audit skill safe?

Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.

No findings.

Install the deep-audit skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow
mkdir -p ~/.claude/skills
cp -r /tmp/claude-code-my-workflow/.claude/skills/deep-audit ~/.claude/skills/deep-audit
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Deep adversarial audit

A convergent alternative to slow round-by-round review. Instead of one reviewer finding one or two issues per pass, fan out many independent skeptics over the whole artifact at once, adjudicate what they find, fix everything confirmed, and re-verify. Modeled on the multi-agent methodology behind hard formal-proof efforts (diverse independent portfolio, adversarial throughout, concrete evidence only, synthesize-challenge-repeat).

When to reach for this

  • The artifact is dense enough that a single review keeps surfacing new issues each pass (the tell that round-by-round is the wrong tool).
  • Correctness is the priority and the cost of a missed defect is high (a paper going to a top venue, a proof, a security-sensitive change, a migration).
  • The user asked to "fix ALL of it", "be deeper", "converge faster", "100% rock solid".

Requires the user to have opted into multi-agent orchestration (they asked for a workflow / deep audit / to fan out agents, or ultracode is on). If they haven't, propose it and its rough cost first.

The method

1. Decompose (diverse portfolio). Break the artifact into components by idea, not by section: each independent claim, lemma, estimator, subsystem, invariant. Add cross-cutting failure-mode lenses (see below). Aim for coverage such that every load-bearing claim is attacked by at least one agent that is looking straight at it. Don't tell the agents your favored reading — preserve independence so they don't all converge on the same attractive-but-wrong conclusion.

2. Fan out adversarial finders (one per component). Each finder is prompted to refute, defaulting to "there is a bug," and must ground every claim in the actual text/code (read it, don't paraphrase from memory). Hard rules, borrowed from what works:

  • Concrete findings only. Every finding = exact location (file:line / label + quoted text) + one-sentence defect + a failing case (specific inputs/configuration → wrong output, or the exact missing hypothesis).
  • Reject status reports, "looks fine", "this is standard/routine", vague optimism, and "the global step is straightforward."
  • A fix that re-imposes the same difficulty elsewhere, or assumes its own conclusion, is not a fix — flag it.
  • If, after genuinely attacking, nothing is found, the agent must state the specific attacks it ran and why each closed — not just "clean."

3. Adjudicate every finding (independent judge). A separate judge re-opens each cited location and decides CONFIRMED / REFUTED / DOWNGRADED, skeptical of both the artifact and the finding. This kills false positives (misreads, hypotheses that are actually present elsewhere, failing cases that don't arise under the stated conditions) — the step that keeps the fix list honest.

4. Synthesize. Dedup by location, rank fatal > major > minor, and hand back one clean defect list. Nothing is accepted as an issue until it survives this.

5. Fix all confirmed, then re-verify. Apply every confirmed fix (you, in the main loop — fixing needs care and judgment). Then re-audit the touched spots and check that no fix created a new defect. Repeat waves until two consecutive audit passes come back empty (fallback cap: 5 waves; a finding that survives waves N and N+2 goes to the user rather than a third patch). Don't stop after the first wave.

Failure-mode lenses (adapt to domain)

Beyond per-component attacks, sweep these cross-cutting modes explicitly — they are where real defects hide:

  • Overclaim: the headline/abstract claims more than the theorems/tests actually deliver.
  • Scope creep in a proof: a pointwise result used where a uniform one is needed; a both-correct property stated unqualified; a special-case argument invoked generally.
  • Silent hypotheses: a differentiability/density/continuity/boundedness/positivity condition used but never stated (in math), or an un-checked precondition/invariant (in code).
  • Edge cases: atoms/ties, endpoints/unbounded support, empty/degenerate inputs, boundary of the parameter space.
  • Circularity: an assumption that assumes its own conclusion; a result that cites itself; a citation that gives less than claimed (read the cited source).
  • Internal contradiction: a definition/notation used two ways; a table cell contradicting a proposition; main text vs appendix disagreement; a dangling/wrong cross-reference.

The fresh-eyes pass (final gate)

Every targeted wave inherits the blind spots of whoever wrote its prompts: focus hints, fix history, and expected failure modes all prime the auditors toward known territory. After all targeted waves and fixes are done, run one cold audit with little to no context: independent auditors given ONLY the artifact and a minimal instruction ("find concrete defects: location + failing case"), with no cluster assignments, no history, no special-focus lists. Diversify only the entry point (main-text-first as a journal referee would; appendix-first; tables/claims-first; a single deep dive of the auditor's own choosing). Adjudicate as usual. Clean fresh-eyes pass + clean targeted coverage + green mechanical battery is the closure standard; a fresh-eyes finding that targeted waves missed is also a diagnosis of the prompt set — add the missed failure mode to the lenses.

Full inventory — never sample

For a paper/proof artifact: enumerate every formal statement first (grep \begin{theorem|proposition|lemma|corollary} + labels) and assign each proof to a verifier — coverage must be 100% of load-bearing statements, not "a few proofs of the reviewer's choice." Sampling converges linearly and stochastically; inventories converge in one wave. Group tightly-coupled small lemmas into clusters; big proofs get their own verifier. Each verifier returns, besides findings, a steps-verified list and a hypotheses ledger (used-vs-stated; used-but-unstated is a finding).

The mechanical battery (the highest-yield check)

Written arguments can read soundly while the object they define is wrong. For every estimating equation, influence-function identity, identification claim, and population moment, write an executable check that computes the population object on adversarial toy designs — truncation (censoring endpoint below the outcome endpoint), interior atoms, misspecified nuisances, boundary/overlap failure — and asserts the claimed centering/identity numerically (analytic or fine-grid/large-N with fixed seed). Keep the scripts as a permanent test directory in the repo with a README; rerun after any change to the corresponding formula. A 5-line population computation catches classes of defects (tail-renormalized roots, sign flips, mass-deficit weighting) that neither careful reading nor model consensus reliably finds.

Fix hygiene

  • New math introduced by fixes is un-audited math: every fix wave is followed by a verification wave over exactly the fixed spots before anything is declared closed.
  • In workflow synthesis, match findings to verdicts by INDEX (require the judge to return verdicts in the findings' order), never by location string — judges paraphrase locations and silent drops follow. Verify the synthesized summary against the journal before acting on it.

Blocked routes are outcomes

If a component cannot be fixed under the stated assumptions, that is a finding, not a failure of the audit: report the exact remaining gap (the precise missing hypothesis or broken step) and the honest options (weaken the claim, add the hypothesis, restrict scope). Do not search for a favorable reading, and do not let an agent paper over a theorem-strength gap as "routine."

Orchestration

  • Fan out with parallel Agent calls in one message (one finder per component), then one judge per component over that component's findings, then synthesize; see orchestrator-protocol.md. Where the Workflow tool is available (e.g. an ultracode session), pipeline(components, finder, judge) is an optional accelerator that judges each component's findings the moment its finder returns (no barrier); it is never a requirement. Return the confirmed list; do the fixing yourself afterward.
  • Model division: finders = the strong execution/analysis model (find and attack); judges = the strong adjudication model. Match to the local convention in model-routing.md (here: both roles run on the Opus tier, and a judge gets more effort before any change of tier; the Fable tier is a per-session choice, never the fleet default). Keep judges to one-per-component (adjudicating all that component's findings at once) to conserve the judging budget.
  • Set finder effort high; give each the exact labels/locations to read and its specific attack list.
  • Persist. Don't return "best effort" or a list of why it's hard. Return the confirmed defects (and, once fixed, a clean re-audit) — or the single strongest remaining gap stated exactly.

Prompt skeletons

Finder: "You are a HOSTILE referee auditing ONE component. Read the ACTUAL text at {locations}. Attack: {failure modes}. Return CONCRETE findings only (location + defect + failing case); no 'looks fine'/'routine'/vague. A fix that re-imposes the difficulty isn't a fix. If clean, list the specific attacks you ran and why each closed."

Judge: "An adversarial referee returned these findings on component X. For each, open the cited location, verify against what the text ACTUALLY says and its proof, mark CONFIRMED/REFUTED/DOWNGRADED. Skeptical of both the artifact and the finding."

Shared doctrine lives in one place

Three things this audit depends on are not restated here, because they are the same rules every other verification surface uses and a third copy would drift:

→ /vaccinate, and verification-ladder.md rung 0.

  • Seeded-fault calibration — a check that has not caught a planted defect licenses nothing.

fail the same way. → verification-ladder.md rung 3.

  • Independence and correlated errors — agreement between models is not confirmation; they

→ external-oracle-process.md §6.

  • The five credibility questions — evidence for one never clears another.

Auditing this repository itself

For the repo-infrastructure application — surface-sync, skill/agent/rule integrity, hook and script review, doc-vs-reality drift — see references/repo-infrastructure-audit.md. Start with ./scripts/backtest.sh: the mechanical battery is already written, and an agent should never hand-check what a script decides.

More skills from pedrohcgs/claude-code-my-workflow

  • Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
  • Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
  • Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
  • Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
  • AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
  • AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
  • Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
  • AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
  • Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
  • Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
  • Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
  • Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).

All agent skills → · MCP servers