Mmcp.market

skill-test skill

by Donchitos·Donchitos/Claude-Code-Game-Studios·25k stars·MIT

Validate skill files for structural compliance and behavioral correctness. Four modes: static linter, spec, category rubric, audit.

A100/100content scan

Is the skill-test skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the skill-test skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git /tmp/Claude-Code-Game-Studios
mkdir -p ~/.claude/skills
cp -r /tmp/Claude-Code-Game-Studios/.claude/skills/skill-test ~/.claude/skills/skill-test
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

!bash "${CLAUDESKILLDIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation

Automation mode: Resolve modes.automation (project.local.yaml → project.yaml → default collaborative). Every AskUserQuestion call and every file write follows .claude/docs/automation-modes.md (collaborative asks always · guided major-only · autonomous logs and proceeds; automationalwaysask categories always prompt).

Skill Test

Validates .claude/skills/*/SKILL.md files for structural compliance and behavioral correctness. No external dependencies — runs entirely within the existing skill/hook/template architecture.

Four modes:

Phase 1: Parse Arguments

Determine mode from the first argument:

  • static [name] → run 7 structural checks on one skill
  • static all → run 7 structural checks on all skills (Glob .claude/skills/*/SKILL.md)
  • spec [name] → read skill + test spec, evaluate assertions
  • category [name] → run category-specific rubric from CCGS Skill Testing Framework/quality-rubric.md
  • category all → run category rubric for every skill that has a category: in catalog
  • audit (or no argument) → read catalog, list all skills and agents, show coverage

If argument is missing or unrecognized, output usage and stop.

Phase 2A: Static Mode — Structural Linter

For each skill being tested, read its SKILL.md fully and run all 7 checks:

Check 1 — Required Frontmatter Fields

The file must contain all of these in the YAML frontmatter block:

  • name:
  • description:
  • argument-hint:
  • user-invocable:
  • allowed-tools:

FAIL if any are absent.

Check 2 — Multiple Phases

The skill must have ≥2 numbered phase headings. Look for patterns like:

  • ## Phase N or ## Phase N:
  • ## N. (numbered top-level sections)
  • At least 2 distinct ## headings if phases aren't explicitly numbered

FAIL if fewer than 2 phase-like headings are found.

Check 3 — Verdict Keywords

The skill must communicate a clear outcome. Accept any of:

BLOCKED, COMPLETE, READY, COMPLIANT, NON-COMPLIANT

  • Gate / review verdicts — PASS, FAIL, CONCERNS, APPROVED,

findings by severity instead of issuing one verdict for the whole run.

  • Go / no-go verdicts — PROCEED, PIVOT, KILL, GO, NO-GO
  • Severity scales — CRITICAL, HIGH, MEDIUM, LOW. Audit skills rank

FAIL if none are present and the skill produces an assessment — its description or body promises a review, audit, check, gate, or readiness judgement.

WARN (never FAIL) if none are present and the skill's output is an artifact or a value rather than a judgement. /settings is the reference case: it prints and writes configuration and has no verdict to give. Do not invent one to satisfy this check.

The narrow earlier list (gate verdicts only, hard FAIL) failed 5 of 74 skills

for reasons that were not their fault — /prototype and /vertical-slice

advertise PROCEED/PIVOT/KILL in their own descriptions, /adopt and

/security-audit rank by severity, and /settings has no verdict by design.

A linter that cries wolf on 7% of the corpus stops being read.

Check 4 — Collaborative Protocol Language

The skill must contain ask-before-write language. Look for:

  • "May I write" (canonical form)
  • "before writing" or "approval" near file-write instructions
  • "ask" + "write" in close proximity (within same section)

WARN if absent (some read-only skills legitimately skip this). FAIL if allowed-tools includes Write or Edit but no ask-before-write language is found.

Check 5 — Next-Step Handoff

The skill must end with a recommended next action or follow-up path. Look for:

  • A final section mentioning another skill (e.g., /story-done, /gate-check)
  • "Recommended next" or "next step" phrasing
  • A "Follow-Up" or "After this" section

WARN if absent.

Check 6 — Fork Context Complexity

If frontmatter contains context: fork, the skill should have ≥5 phase headings (## level or numbered Phase N headers). Fork context is for complex multi-phase skills; simple skills should not use it.

WARN if context: fork is set but fewer than 5 phases found.

Check 7 — Argument Hint Plausibility

argument-hint must be non-empty. If the skill body mentions multiple modes (e.g., "Mode A | Mode B"), the hint should reflect them. Cross-reference the hint against the first phase's "Parse Arguments" section.

WARN if hint is "" or if documented modes don't match hint.

Static Mode Output Format

For a single skill:

=== Skill Static Check: /[name] ===

Check 1 — Frontmatter Fields:    PASS
Check 2 — Multiple Phases:       PASS (7 phases found)
Check 3 — Verdict Keywords:      PASS (PASS, FAIL, CONCERNS)
Check 4 — Collaborative Protocol: PASS ("May I write" found)
Check 5 — Next-Step Handoff:     WARN (no follow-up section found)
Check 6 — Fork Context Complexity: PASS (8 phases, context: fork set)
Check 7 — Argument Hint:         PASS

Verdict: WARNINGS (1 warning, 0 failures)
Recommended: Add a "Follow-Up Actions" section at the end of the skill.

For static all, produce a summary table then list any non-compliant skills:

=== Skill Static Check: All 74 Skills ===

Skill                  | Result       | Issues
-----------------------|--------------|-------
gate-check             | COMPLIANT    |
design-review          | COMPLIANT    |
story-readiness        | WARNINGS     | Check 5: no handoff
...

Summary: 48 COMPLIANT, 3 WARNINGS, 1 NON-COMPLIANT, 1 NOT ASSESSED
Aggregate Verdict: N WARNINGS / N FAILURES / N NOT ASSESSED

NOT ASSESSED is a per-skill result here, not only an aggregate line. A skill whose file could not be read or parsed, or whose checks could not run, is reported as NOT ASSESSED with the reason — never omitted from the table and never counted as COMPLIANT. Ranked above COMPLIANT, below WARNINGS and NON-COMPLIANT.

And state the denominator. All 74 Skills in the header must be the number actually examined, not the number that exist: report [N] of [M] skills checked whenever they differ. A summary whose counts silently sum to less than its own title is the failure this skill is supposed to catch in others.

Phase 2B: Spec Mode — Behavioral Verifier

Step 1 — Locate Files

Find skill at .claude/skills/[name]/SKILL.md. Look up the spec path from CCGS Skill Testing Framework/catalog.yaml — use the spec: field for the matching skill entry.

If either is missing:

to see coverage gaps."

  • Missing skill: "Skill '[name]' not found in .claude/skills/."
  • Missing spec path in catalog: "No spec path set for '[name]' in catalog.yaml."
  • Spec file not found at path: "Spec file missing at [path]. Run /skill-test audit

Step 2 — Read Both Files

More skills from Donchitos/Claude-Code-Game-Studios

  • AadoptBrownfield audit — do existing artifacts actually work? Numbered migration plan. Unlike /project-stage-detect, checks compliance not existence.
  • Aarchitecture-decisionCreate an ADR documenting a technical decision: context, alternatives considered, consequences.
  • Aarchitecture-reviewTraceability matrix mapping GDD requirements to ADRs. Finds gaps, cross-ADR conflicts, engine compatibility. PASS/CONCERNS/NOT ASSESSED/FAIL.
  • Aart-bibleAuthor the Art Bible — visual identity gating asset production. Run before /map-systems.
  • Aasset-auditAudit assets against naming conventions, file size budgets, format standards. Finds orphaned assets, missing references.
  • Aasset-specPer-asset visual specs plus AI generation prompts from GDDs and character profiles. After the art bible.
  • Abalance-checkFind balance outliers, broken progressions, degenerate strategies, economy imbalances in formulas and data. 'Check game balance'.
  • AbrainstormGuided concept ideation using professional studio techniques, player psychology, creative exploration.
  • Abug-reportStructured bug report from a description, or analyze code for potential bugs. Reproduction steps, severity.
  • Abug-triageRe-evaluate open bugs — priority vs severity, assign to sprints, surface systemic trends. Run when the count grows.
  • AchangelogAuto-generate a changelog from git commits and sprint data. Internal and player-facing versions.
  • Acode-reviewArchitectural code review — coding standards, SOLID, testability, performance concerns.

All agent skills → · MCP servers