vaccinate skill
Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch. Seeds known defects into a copy of a real artifact plus a clean control, runs the checker, and reports recall and false-positive rate into a qualification ledger. Use when the user says "does this check work", "qualify the gate", "test my reviewer", "seed defects", "vaccinate", "qualify the checks", "can I trust this review", "does it pass for the right reason", or before relying on any automated check or referee simulation for a decision that matters. NOT a code fixer and NOT a reviewer itself — it grades the grader.
Is the vaccinate skill safe?
Clean: nothing in its files matched our rules. We read 7 files in the folder on 2026-09-28.
No findings.
Install the vaccinate skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow mkdir -p ~/.claude/skills cp -r /tmp/claude-code-my-workflow/.claude/skills/vaccinate ~/.claude/skills/vaccinate
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Vaccinate — grade the grader
Twenty bugs were once planted in a working codebase and the review agents were asked to check it again. They reported everything was fine. Recall: 0/20.
A vaccine is a small, controlled dose of error that strengthens the whole system. This skill administers one.
The rule it enforces: an unqualified check is not weak evidence — it is none.
When to run it
decision that matters (a submission, a release, a deposit).
- Before a referee simulation, reproducibility gate, or review agent is used to make a
- After changing a checker — a modified gate is unqualified until re-measured.
- On a schedule for gates that guard load-bearing claims. Detection decays as artifacts drift.
Protocol
1. Name the failure
State the defect class the check is supposed to catch. "Catches problems" is not a class. "Detects a coefficient in the text that no longer matches its table" is.
2. Build the seeded set + a clean control
Work on a copy, never the live artifact. Produce:
- N seeded variants, one defect each, drawn from references/defect-library.md.
- At least one clean control — an unmodified copy.
The control is not optional. Without it you measure recall and call it accuracy.
Verify each seed actually violates something. A seed that the artifact already permits creates no defect, and the checker correctly reporting "pass" will look like a broken gate. This is the most common way a qualification run produces a false alarm about itself.
3. Run the checker blind
Run the check or agent against each variant in a fresh context, one variant per run. It must not know which variant it has, how many defects exist, or that a qualification is underway. For an AI reviewer, spawn a fresh-context Agent call (not a conversation fork, which would carry the seeded answer key).
4. Score
A finding on the clean control counts as a false positive only when it is factually wrong — not merely unwelcome. A reviewer prompted to find gaps will report some in sound work; that is expected behaviour, not a failure.
The baseline is load-bearing. A five-agent panel that scores no better than grep -n has not earned its cost.
5. Write the ledger row
Append to quality_reports/qualification/LEDGER.md:
| date | target | artifact | defect classes | N | recall | FPR | baseline | verdict |Verdicts: PASS (detects its named class at an agreed threshold) · FAIL (misses it) · BLOCKED (could not be run — say why; do not record as PASS).
6. Act on the result
weaken the seed until it passes.
- FAIL → the check does not license its claim. Fix the check or stop citing it. Do not
nothing.
- PASS → record the threshold. A PASS at one difficulty is not a PASS at another.
- Either way, a checker with no ledger row is unqualified, and its green light means
Worked example
/vaccinate check-model-versions.shControl: unmodified README.md.
- Failure class: "a superseded model presented as current".
- Seed: append The newest model is Opus 4.8 and it is the default. to README.md.
Recall 1/1, FPR 0/0. Baseline: grep -c "Opus 4.8" README.md also detects — so the gate's value is its allow-marker logic, not raw detection.
- Run: bash scripts/check-model-versions.sh; echo $?
- Score: seeded → exit 1 (detected). Control → exit 0 (no false alarm).
- Ledger: PASS.
- Restore the artifact and re-run to confirm you are back to green.
Anti-patterns
broken gate.
- Seeding into an artifact that already permits the seed — measures nothing, looks like a
replicates per class where cost allows.
- Telling the reviewer it is a test — it will look harder than it does in production.
- Counting any finding as a hit — a finding at the wrong location is not detection.
- One seed, one run — a single trial does not distinguish detection from luck. Use ≥2
- Weakening the seed until it passes — that is fitting the test to the checker.
- Skipping the clean control — the most common omission, and it hides the cost.
Reference files
Doctrine: what qualification means
Do not assume more machinery is better
A second model, more agents, or a longer debate is not presumed to verify better. Before an elaborate procedure earns extra weight, show it outperforms a simpler check on the same prespecified seeded failures and valid cases, reporting both detection and false alarms. Complexity that has not beaten a baseline is cost, not assurance.
Treat AI verdicts as predictions, not facts
When a model grades, triages, or reviews at scale:
- keep a sampled set for qualified human review, and record how it was sampled (retain coverage of hard subgroups — do not sample only the easy middle);
- keep fitting/prompt-tuning cases separate from evaluation cases;
- report where AI and expert judgments diverge;
- remember a well-calibrated average score certifies no individual verdict;
- agreement between models is not independent evidence — they share failure modes and converge on the same wrong answer at a meaningful rate.
Any material change to the model, prompt, rubric, or target population requires fresh human labels and recalibration.
Requalify after material change
A check qualified against an old interface, schema, or scale may silently stop testing anything. Re-run the seeded-defect proof after material changes to the object under test or to the check itself.
Distinguish qualified checks from scientific judgments
regenerates, an estimator recovers an analytic special case, a seeded fault triggers a failure. These can be automated and rerun forever.
- Qualified checks have a defensible reference answer: unique keys, units convert, a table
identifying assumption is plausible, whether a result deserves causal language — cannot be automated, and no volume of qualified checks substitutes for one.
- Scientific judgments — whether a field measures the intended construct, whether an
Confirm the check actually ran
A missing, substituted, or degraded check is missing evidence, not a pass. Verify the run happened (log, exit status, artifact timestamp — not an assumption); that it ran on the current object, not a cached one; that nothing was skipped, filtered, or swallowed into a default; and that the tolerance was fixed before the comparison. A tolerance loosened after a failed comparison converts evidence into decoration. If it must be loosened, record it as an approved divergence with a reason.
More skills from pedrohcgs/claude-code-my-workflow
- Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
- Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
- Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
- Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
- AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
- AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
- Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
- AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
- Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
- Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
- Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
- Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).