test-evidence-review skill
Quality review of test files and evidence — goes beyond existence, evaluates assertion coverage. ADEQUATE/INCOMPLETE/MISSING/NOT ASSESSED per story.
Is the test-evidence-review skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the test-evidence-review skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git /tmp/Claude-Code-Game-Studios mkdir -p ~/.claude/skills cp -r /tmp/Claude-Code-Game-Studios/.claude/skills/test-evidence-review ~/.claude/skills/test-evidence-review
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
!bash "${CLAUDESKILLDIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation
Automation mode: Resolve modes.automation (project.local.yaml → project.yaml → default collaborative). Every AskUserQuestion call and every file write follows .claude/docs/automation-modes.md (collaborative asks always · guided major-only · autonomous logs and proceeds; automationalwaysask categories always prompt).
Test Evidence Review
/smoke-check verifies that test files exist and pass. This skill goes further — it reviews the quality of those tests and evidence documents. A test file that exists and passes may still leave critical behaviour uncovered. A manual evidence doc that exists may lack the sign-offs required for closure.
Output: Summary report (in conversation) + optional production/qa/evidence-review-[date].md
When to run:
- Before QA hand-off sign-off (/team-qa Phase 5)
- On any story where test quality is in question
- As part of milestone review for Logic and Integration story quality audit
1. Parse Arguments
Modes:
- /test-evidence-review [story-path] — review a single story's evidence
- /test-evidence-review sprint — review all stories in the current sprint
- /test-evidence-review [system-name] — review all stories in an epic/system
- No argument — ask which scope: "Single story", "Current sprint", "A system"
2. Load Stories in Scope
Based on the argument:
Single story: Read the story file directly. Extract: Story Type, Test Evidence section, story slug, system name.
Sprint: Read the most recently modified file in production/sprints/; extract the list of story file paths from the sprint plan.
System: Glob production/epics/[system-name]/story-*.md.
If the resolved scope contains ZERO stories, stop here. Report
NOT ASSESSED — no stories in scope, name which scope was searched and which
path was empty, and route: no sprint file → /sprint-plan new; a sprint plan
listing no stories → /create-stories [epic-slug]; a [system-name] glob that
matched nothing → name the glob. Do not continue to Section 3.
Guard the empty scope, not just the per-story unknown. The verdict
vocabulary here — ADEQUATE / INCOMPLETE / MISSING — needs a "could not check"
value, or an unverifiable story acquires a verdict claiming somebody verified
it; that is what NOT ASSESSED is for. **But giving the per-story unknown a
home does nothing for the empty-scope unknown.** With no stories
the Section 6 report renders an empty Summary table and ends
BLOCKING items: 0 / ADVISORY items: 0 — which reads as *everything reviewed,
all fine*. This skill gates story closure, and coding-standards.md marks Logic,
Integration, Visual/Feel and UI evidence BLOCKING, so a false-clean closes stories nobody
reviewed.
This is a recurring shape: the sophisticated inner rule present, the outer
boundary unguarded. Ask it of any skill that aggregates —
what does this emit when the set is empty?
For the resulting story set, collect the fields below with targeted section greps, not a full read of each story:
Grep pattern="## Test Evidence" glob="production/epics/**/story-*.md" output_mode="content" -A 8
Grep pattern="## Acceptance Criteria" glob="production/epics/**/story-*.md" output_mode="content" -A 15stated evidence path — both live under ## Test Evidence, so the first grep's -A 8 captures them.
- Story Type (Logic / Integration / Visual/Feel / UI / Config/Data) and the
Full-read a story only when its Test Evidence section is missing or ambiguous. (In Sprint mode, scope the globs to the sprint plan's story paths.)
- Acceptance Criteria list — the ## Acceptance Criteria block from the second grep.
- Story slug (from the file name) and System (from the directory path) — no read.
3. Locate Evidence Files
For each story, find the evidence:
Logic stories: Glob tests/unit/[system]/[story-slug]test.
containing the story slug
- If not found, also try: Grep in tests/unit/[system]/ for files
Integration stories: Glob tests/integration/[system]/[story-slug]test.
- Also check production/session-logs/ for playtest records mentioning the story
Visual/Feel and UI stories: Glob production/qa/evidence/[story-slug]-evidence.*
Config/Data stories: Glob production/qa/smoke-*.md (any smoke check report)
Note what was found (path) or not found (gap) for each story.
4. Review Automated Test Quality (Logic / Integration)
For each test file found, read it and evaluate:
Assertion coverage
Count the number of distinct assertions (lines containing assert, expect, check, verify, or engine-specific assertion patterns). Low assertion count is a quality signal — a test that makes only 1 assertion per test function may not cover the range of expected behaviour.
Thresholds:
test passes vacuously and proves nothing
- 3+ assertions per test function → normal
- 1-2 assertions per test function → note as potentially thin
- 0 assertions (test exists but no asserts) → flag as BLOCKING — the
Edge case coverage
For each acceptance criterion in the story that contains a number, threshold, or "when X happens" conditional: check whether a test function name or test body references that specific case.
Heuristics:
"boundary", "edge" — presence of any is a positive signal
More skills from Donchitos/Claude-Code-Game-Studios
- AadoptBrownfield audit — do existing artifacts actually work? Numbered migration plan. Unlike /project-stage-detect, checks compliance not existence.
- Aarchitecture-decisionCreate an ADR documenting a technical decision: context, alternatives considered, consequences.
- Aarchitecture-reviewTraceability matrix mapping GDD requirements to ADRs. Finds gaps, cross-ADR conflicts, engine compatibility. PASS/CONCERNS/NOT ASSESSED/FAIL.
- Aart-bibleAuthor the Art Bible — visual identity gating asset production. Run before /map-systems.
- Aasset-auditAudit assets against naming conventions, file size budgets, format standards. Finds orphaned assets, missing references.
- Aasset-specPer-asset visual specs plus AI generation prompts from GDDs and character profiles. After the art bible.
- Abalance-checkFind balance outliers, broken progressions, degenerate strategies, economy imbalances in formulas and data. 'Check game balance'.
- AbrainstormGuided concept ideation using professional studio techniques, player psychology, creative exploration.
- Abug-reportStructured bug report from a description, or analyze code for potential bugs. Reproduction steps, severity.
- Abug-triageRe-evaluate open bugs — priority vs severity, assign to sprints, surface systemic trends. Run when the count grows.
- AchangelogAuto-generate a changelog from git commits and sprint data. Internal and player-facing versions.
- Acode-reviewArchitectural code review — coding standards, SOLID, testability, performance concerns.