Mmcp.market

verify-claims skill

by pedrohcgs·pedrohcgs/claude-code-my-workflow·1.6k stars·MIT

Run Chain-of-Verification (CoVe) on a draft or a block of text with factual claims. Spawns the `claim-verifier` agent in a fresh context (never a conversation fork) so it never sees the draft — then reports which claims are supported, contradicted, or unverifiable. Use when user says "verify these citations", "check the claims in X", "did I hallucinate anything", "fact-check this draft", "run CoVe on this", or after any text generation that asserts facts about papers, datasets, or numerical results. NOT for style/grammar review (use `/proofread`) or substance review (use `/review-paper`).

A100/100content scan

Is the verify-claims skill safe?

Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.

No findings.

Install the verify-claims skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git /tmp/claude-code-my-workflow
mkdir -p ~/.claude/skills
cp -r /tmp/claude-code-my-workflow/.claude/skills/verify-claims ~/.claude/skills/verify-claims
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

/verify-claims — Chain-of-Verification on a Draft

Fact-check a draft using the Post-Flight Verification protocol (.claude/rules/post-flight-verification.md).

Input: $ARGUMENTS — path to a file containing the draft (markdown, .qmd, .tex, .md) or a shorthand pointer. Optional flags:

  • --source — one or more source-material pointers (repeat for multiple). If omitted, the skill infers from context (e.g., papers referenced, cited arXiv URLs).
  • --no-fail-closed — downgrade FAIL outcomes to warnings without regeneration. Use sparingly.

When to pick this skill

  • /verify-claims (this skill) — ad-hoc fact-checking on any draft or text block the user hands you. One-shot, user-invoked.
  • Other skills that auto-run Post-Flight internally (/lit-review, /research-ideation, /respond-to-referees, /review-paper --peer) — no need to call this separately; they already run it.
  • /proofread — grammar, typos, overflow. Different lens.
  • /review-paper (default mode) — full manuscript review, not just claim verification.
  • /validate-bib — checks citations exist and are well-formed (structural + DOI). This skill checks they hold (the cited paper supports the attributed claim). Complementary — run both before submission.

How it works

Implements the 4-step CoVe loop from Dhuliawala et al. 2023 (arXiv:2309.11495), with architectural enforcement of the fresh-context independence trick.

Phase 0 — Pre-Flight

Confirm:

  • Draft file exists and is readable
  • At least one source pointer available (either --source or auto-detected from draft)
  • claim-verifier agent file exists at .claude/agents/claim-verifier.md

If any fail → surface the failure, do NOT proceed.

Phase 1 — Extract claims

Read the draft. Identify factual assertions of these types:

Skip: opinions, forward-looking suggestions, definitions the draft introduces.

For citation-type claims, extract the claim↔citation PAIR — not just the citation. Capture what the draft attributes to which work, so the verifier checks appropriateness (does Smith 2019 actually show X?), not merely existence. "Smith (2019) shows a positive wage effect" becomes {cite: Smith2019, attributed: "positive wage effect"}. This is the layer /validate-bib explicitly defers here: validate-bib confirms the citation exists and is well-formed; this skill confirms it holds. A mis-citation (the paper exists but says something else, or the opposite) is exactly a numeric/directional contradiction → HIGH-WARN unless a concrete author_alternative is recorded (then EXPLAINED).

Output a claims table:

| ID | Claim | Source hint |
|----|-------|-------------|
| C1 | ... | ... |

Phase 2 — Generate verification questions

One question per claim. Make it specific and answerable from the source alone.

Phase 3 — Spawn claim-verifier (fresh context — never a conversation fork)

Agent: subagent_type=claim-verifier   # a named subagent starts fresh; a /fork copy would inherit the draft
Prompt: hand over claims table + verification questions + source material pointers.
        Do NOT include the draft text.

The forked agent runs the CoVe independent-answer step. It has never seen the draft and cannot confirm-bias. It returns a structured verification report.

Phase 4 — Reconcile

The verifier returns a per-claim verdict in one of these severity tiers:

  • HIGH-WARN — fabricated reference (the cited paper doesn't exist at the named venue/year), draft claim directly contradicted by the source, or notfound retrieval that the verifier interprets as a hallucinated citation. Fail closed** — surface these first and never present the draft as verified while one stands. This is a reporting rule, not a mechanical gate: nothing in /commit or the pre-commit hook reads these verdicts, so the author decides, and a HIGH-WARN left in place is stated in the report the user sees.
  • MED-WARN — transient infrastructure / retrieval failure (paywall the verifier can normally bypass via cached metadata; DOI resolver timeout; partial PDF read). Surface for the author; do not gate-refuse.
  • LOW-WARN — source genuinely inaccessible (paywalled and not in cache; private dataset; pre-print server transient). Surface with cannot-verify flag; do not gate-refuse.
  • EXPLAINED (v2.0) — a numeric/directional contradiction the author has pre-justified with a concrete named alternative (different defensible edition, specification, sample, or rounding convention), passed to the verifier via the claim's authoralternative field. Surfaced with the evidence and the recorded reason; non-gating. The hard floor holds: a fabricated* citation is never EXPLAINED, and a blank/vague alternative stays HIGH-WARN. This mirrors audit-reproducibility's EXPLAINED disposition for numeric claims — a mismatch is not always a failure when a defensible alternative is named.

Verdict aggregation by tier across all extracted claims (EXPLAINED counts as non-gating, like LOW):

--no-fail-closed reports HIGH-WARN verdicts as ordinary warnings instead of failing closed. Use sparingly — it's there for offline / hallucination-sensitive contexts where the user accepts the risk in writing.

If the draft is writeable and the user asked for auto-correction, regenerate the affected sections using the verifier's evidence. Otherwise return the report and let the user decide.

Example

/verify-claims quality_reports/lit-review_measurement-error.md --source master_supporting_docs/author_2021_method.pdf --source master_supporting_docs/coauthor_2020_survey.pdf

Expected output (abridged):

## Post-Flight Verification — lit-review_measurement-error.md

**Claims extracted:** 14
**Verified independently:** 14 (fresh-context claim-verifier)
**Outcome:** FAIL — 12 verified, 1 contradicted by its source (HIGH-WARN), 1 unverifiable (LOW-WARN); the draft is not reported as verified until C7 is corrected

### Discrepancies

- **C7** — draft claims "Coauthor (2020) *proposes* a bias-corrected estimator." Source Section 4 shows they propose a weighting estimator, not a bias-corrected one. Recommend correction.

### Unverifiable

- **C12** — draft cites "Third Author et al. 2024 (working paper)". No canonical URL in provided sources. Recommend user supply DOI or arXiv link.

### Verified

| ID | Claim | Evidence |
|----|-------|----------|
| C1 | "Author 2021 defines the calibration constant as a ratio of moments" | p. 5, eq. (3) |
| ... | ... | ... |

Fail modes and recovery

Verifier times out: surface a warning block, return draft as provisional. Do not silently ship.

Source material inaccessible (paywall, 404): report the specific claims that hinge on it, flag as cannot-verify, recommend user supply an alternative source.

Draft contains only opinions / forward-looking text: report "no verifiable factual claims extracted — nothing to check" and return.

Cross-references

  • .claude/agents/claim-verifier.md — the fresh-context verifier.
  • .claude/rules/post-flight-verification.md — the protocol.
  • MEMORY.md [LEARN:pattern] on Chain-of-Verification vs critic-fixer vs cross-artifact review.

More skills from pedrohcgs/claude-code-my-workflow

  • Aadjudicate-reviewTurn an incoming set of findings — from an AI reviewer, a referee report, a code review, a linter, or a second model — into verified fixes, without letting a confident misread damage correct work. Every finding is a CANDIDATE until checked against the actual source. Use whenever you receive review comments, audit findings, or a critique you did not write yourself, especially when the reviewer is a model or when the volume is too large to check by feel.
  • Aaudit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
  • Ablast-radiusBefore and after changing anything shared — a function's return value, a signature, a schema, a label set, a config default, a constant, a file format — find every consumer and actually run them. Catches the change that looks purely additive but silently breaks a contract in a file you never opened. Use when editing shared code, adding a field/column/return element, renaming, changing units or defaults, or touching a pipeline that produces reported numbers.
  • Acapture-environmentSnapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt / environment.yml / uv.lock, Stata version + ado package list), records seeds and RNG kind, optionally writes a pinning Dockerfile, and produces a paste-ready "Computational requirements" block. Use when user says "capture the environment", "snapshot my dependencies", "pin the versions", "make a renv.lock / requirements.txt", "make this byte-reproducible", or before releasing a replication package to openICPSR / the AEA Data Editor.
  • AchallengeStress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
  • AcheckpointSave a structured state snapshot before stopping or handing off. Captures the active plan, recent decisions, file pointers (with line numbers), open questions, and the next 1–3 actions into a checkpoint file under `quality_reports/checkpoints/`. Optionally proposes `[LEARN]` entries to add to MEMORY.md. Use when user says "checkpoint", "save state", "snapshot before I stop", "where am I", "wrap up the session for handoff", or before a long break / model switch / collaborator handoff. Companion to (NOT replacement for) the narrative session-log workflow.
  • Acoauthor-briefGenerate a co-author / collaborator handoff brief for a multi-author, multi-machine project — summarizing what changed since the last brief (git delta), the current state of each artifact (manuscript, analysis, slides), open questions, how to reproduce locally, and any restricted-data access steps. Use when user says "coauthor brief", "handoff brief", "bring my coauthor up to speed", "what changed since last week", "onboard a collaborator", "write a handoff for [name]", or before sending a co-author the repo. NOT a commit or a checkpoint — it is the cross-machine, cross-person summary `meta-governance.md` only partially covers.
  • AcommitCommit the current work — runs the quality, consistency and passport gates, branches off main if needed, stages specific files, and writes a commit whose subject states what is now true. Pushes and opens a pull request only with --pr or when the user asks; never merges — a merge happens only when the user explicitly says to merge. Use ONLY on explicit commit intent — user says "commit", "let's commit this", "open a PR", or prefixes with `/commit`. Do NOT auto-invoke on vague end-of-task phrases ("we're done", "wrap up") — those require explicit confirmation first. Never force-pushes or skips hooks.
  • Acompile-latexCompile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex). Use when user says "compile", "build the slides", "rebuild the PDF", "run latex", "render the tex", or asks why a `.tex` file isn't producing a PDF. Operates on `Slides/*.tex`.
  • Acompress-sessionDistill the current conversation into a structured note (decisions made, open questions, file pointers with line numbers, next 1–3 actions) and save to `quality_reports/session_logs/` before auto-compression. Differs from `/checkpoint` (explicit stop-point snapshot) and from auto-compaction (which truncates rather than distills). Use when context is approaching auto-compact threshold, when a long pipeline has accumulated many decisions, or when the user says "compress", "distil this session", "before we hit auto-compact", "structured handoff before context resets".
  • Acontext-statusShow current context status and session health. Use to check how much context has been used, whether auto-compact is approaching, and what state will be preserved.
  • Acreate-lectureCreate a new Beamer lecture `.tex` from source papers and materials, with notation consistency checks and the project's preamble wired in. Use when user says "create a lecture on X", "new lecture from these papers", "start a deck on topic Y", "scaffold a new Beamer file", "build me a lecture from these PDFs". Scaffolds the full deck — NOT for compiling existing `.tex` (use `/compile-latex`).

All agent skills → · MCP servers