Mmcp.market

test-evidence-review skill

by Donchitos·Donchitos/Claude-Code-Game-Studios·25k stars·MIT

Quality review of test files and evidence — goes beyond existence, evaluates assertion coverage. ADEQUATE/INCOMPLETE/MISSING/NOT ASSESSED per story.

A100/100content scan

Is the test-evidence-review skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the test-evidence-review skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git /tmp/Claude-Code-Game-Studios
mkdir -p ~/.claude/skills
cp -r /tmp/Claude-Code-Game-Studios/.claude/skills/test-evidence-review ~/.claude/skills/test-evidence-review
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

!bash "${CLAUDESKILLDIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation

Automation mode: Resolve modes.automation (project.local.yaml → project.yaml → default collaborative). Every AskUserQuestion call and every file write follows .claude/docs/automation-modes.md (collaborative asks always · guided major-only · autonomous logs and proceeds; automationalwaysask categories always prompt).

Test Evidence Review

/smoke-check verifies that test files exist and pass. This skill goes further — it reviews the quality of those tests and evidence documents. A test file that exists and passes may still leave critical behaviour uncovered. A manual evidence doc that exists may lack the sign-offs required for closure.

Output: Summary report (in conversation) + optional production/qa/evidence-review-[date].md

When to run:

  • Before QA hand-off sign-off (/team-qa Phase 5)
  • On any story where test quality is in question
  • As part of milestone review for Logic and Integration story quality audit

1. Parse Arguments

Modes:

  • /test-evidence-review [story-path] — review a single story's evidence
  • /test-evidence-review sprint — review all stories in the current sprint
  • /test-evidence-review [system-name] — review all stories in an epic/system
  • No argument — ask which scope: "Single story", "Current sprint", "A system"

2. Load Stories in Scope

Based on the argument:

Single story: Read the story file directly. Extract: Story Type, Test Evidence section, story slug, system name.

Sprint: Read the most recently modified file in production/sprints/; extract the list of story file paths from the sprint plan.

System: Glob production/epics/[system-name]/story-*.md.

If the resolved scope contains ZERO stories, stop here. Report

NOT ASSESSED — no stories in scope, name which scope was searched and which

path was empty, and route: no sprint file → /sprint-plan new; a sprint plan

listing no stories → /create-stories [epic-slug]; a [system-name] glob that

matched nothing → name the glob. Do not continue to Section 3.

Guard the empty scope, not just the per-story unknown. The verdict

vocabulary here — ADEQUATE / INCOMPLETE / MISSING — needs a "could not check"

value, or an unverifiable story acquires a verdict claiming somebody verified

it; that is what NOT ASSESSED is for. **But giving the per-story unknown a

home does nothing for the empty-scope unknown.** With no stories

the Section 6 report renders an empty Summary table and ends

BLOCKING items: 0 / ADVISORY items: 0 — which reads as *everything reviewed,

all fine*. This skill gates story closure, and coding-standards.md marks Logic,

Integration, Visual/Feel and UI evidence BLOCKING, so a false-clean closes stories nobody

reviewed.

This is a recurring shape: the sophisticated inner rule present, the outer

boundary unguarded. Ask it of any skill that aggregates —

what does this emit when the set is empty?

For the resulting story set, collect the fields below with targeted section greps, not a full read of each story:

Grep pattern="## Test Evidence" glob="production/epics/**/story-*.md" output_mode="content" -A 8
Grep pattern="## Acceptance Criteria" glob="production/epics/**/story-*.md" output_mode="content" -A 15

stated evidence path — both live under ## Test Evidence, so the first grep's -A 8 captures them.

  • Story Type (Logic / Integration / Visual/Feel / UI / Config/Data) and the

Full-read a story only when its Test Evidence section is missing or ambiguous. (In Sprint mode, scope the globs to the sprint plan's story paths.)

  • Acceptance Criteria list — the ## Acceptance Criteria block from the second grep.
  • Story slug (from the file name) and System (from the directory path) — no read.

3. Locate Evidence Files

For each story, find the evidence:

Logic stories: Glob tests/unit/[system]/[story-slug]test.

containing the story slug

  • If not found, also try: Grep in tests/unit/[system]/ for files

Integration stories: Glob tests/integration/[system]/[story-slug]test.

  • Also check production/session-logs/ for playtest records mentioning the story

Visual/Feel and UI stories: Glob production/qa/evidence/[story-slug]-evidence.*

Config/Data stories: Glob production/qa/smoke-*.md (any smoke check report)

Note what was found (path) or not found (gap) for each story.

4. Review Automated Test Quality (Logic / Integration)

For each test file found, read it and evaluate:

Assertion coverage

Count the number of distinct assertions (lines containing assert, expect, check, verify, or engine-specific assertion patterns). Low assertion count is a quality signal — a test that makes only 1 assertion per test function may not cover the range of expected behaviour.

Thresholds:

test passes vacuously and proves nothing

  • 3+ assertions per test function → normal
  • 1-2 assertions per test function → note as potentially thin
  • 0 assertions (test exists but no asserts) → flag as BLOCKING — the

Edge case coverage

For each acceptance criterion in the story that contains a number, threshold, or "when X happens" conditional: check whether a test function name or test body references that specific case.

Heuristics:

"boundary", "edge" — presence of any is a positive signal

More skills from Donchitos/Claude-Code-Game-Studios

  • AadoptBrownfield audit — do existing artifacts actually work? Numbered migration plan. Unlike /project-stage-detect, checks compliance not existence.
  • Aarchitecture-decisionCreate an ADR documenting a technical decision: context, alternatives considered, consequences.
  • Aarchitecture-reviewTraceability matrix mapping GDD requirements to ADRs. Finds gaps, cross-ADR conflicts, engine compatibility. PASS/CONCERNS/NOT ASSESSED/FAIL.
  • Aart-bibleAuthor the Art Bible — visual identity gating asset production. Run before /map-systems.
  • Aasset-auditAudit assets against naming conventions, file size budgets, format standards. Finds orphaned assets, missing references.
  • Aasset-specPer-asset visual specs plus AI generation prompts from GDDs and character profiles. After the art bible.
  • Abalance-checkFind balance outliers, broken progressions, degenerate strategies, economy imbalances in formulas and data. 'Check game balance'.
  • AbrainstormGuided concept ideation using professional studio techniques, player psychology, creative exploration.
  • Abug-reportStructured bug report from a description, or analyze code for potential bugs. Reproduction steps, severity.
  • Abug-triageRe-evaluate open bugs — priority vs severity, assign to sprints, surface systemic trends. Run when the count grows.
  • AchangelogAuto-generate a changelog from git commits and sprint data. Internal and player-facing versions.
  • Acode-reviewArchitectural code review — coding standards, SOLID, testability, performance concerns.

All agent skills → · MCP servers