Mmcp.market

test-flakiness skill

by Donchitos·Donchitos/Claude-Code-Game-Studios·25k stars·MIT

Find flaky tests from CI logs — aggregates pass rates, spots intermittent failures, recommends quarantine. After multiple runs.

A100/100content scan

Is the test-flakiness skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the test-flakiness skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git /tmp/Claude-Code-Game-Studios
mkdir -p ~/.claude/skills
cp -r /tmp/Claude-Code-Game-Studios/.claude/skills/test-flakiness ~/.claude/skills/test-flakiness
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

!bash "${CLAUDESKILLDIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation

Automation mode: Resolve modes.automation (project.local.yaml → project.yaml → default collaborative). Every AskUserQuestion call and every file write follows .claude/docs/automation-modes.md (collaborative asks always · guided major-only · autonomous logs and proceeds; automationalwaysask categories always prompt).

Test Flakiness Detection

A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.

Output: Updated tests/regression-suite.md quarantine section + optional production/qa/flakiness-report-[date].md

When to run:

  • Polish phase (tests have had many runs; statistical signal is reliable)
  • When developers start dismissing CI failures as "probably flaky"
  • After /regression-suite identifies quarantined tests that need diagnosis

1. Parse Arguments

Modes:

standard log output directories

  • /test-flakiness [ci-log-path] — analyse a specific CI run log file
  • /test-flakiness scan — scan all available CI logs in .github/ or

section and provide remediation guidance for already-known flaky tests

  • /test-flakiness registry — read existing regression-suite.md quarantine

registry

  • No argument — auto-detect: run scan if CI logs are accessible, else

2. Locate CI Log Data

Option A — GitHub Actions (preferred)

Check for test result artifacts:

ls -t .github/ 2>/dev/null
ls -t test-results/ 2>/dev/null

For Godot projects: GdUnit4 outputs XML results compatible with JUnit format. Check test-results/ for .xml files.

For Unity projects: game-ci test runner outputs NUnit XML to test-results/ by default.

For Unreal projects: automation logs go to Saved/Logs/. Grep for Result: Success and Result: Fail patterns.

Option B — Local log files

If a path argument is provided, read that file directly.

Option C — No log data available

If no logs found:

"No CI log data found. To detect flaky tests, this skill needs test result

history from multiple runs. Options:

1. Run the test suite at least 3 times and collect the output logs

2. Check CI pipeline output and save a log to test-results/

3. Run /test-flakiness registry to review tests already flagged as flaky

in tests/regression-suite.md"

Stop and ask the user which option to pursue.

3. Parse Test Results

For each CI log or result file found, parse:

JUnit XML format (GdUnit4 / Unity):

  • Grep for <testcase name= to get test names
  • Grep for <failure or <error to identify failures
  • Parse classname and name attributes for full test identifiers

Plain text logs:

  • Grep for pass/fail patterns:
  • Godot: PASSED / FAILED adjacent to test names
  • Unreal: Result: Success / Result: Fail
  • Unity: Test passed / Test failed

Build a table: testid → [run1result, run2result, run3result, ...]

4. Identify Flaky Tests

A test is flaky if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.

Flakiness thresholds:

genuinely rare failure

  • High flakiness: Fails in >25% of runs — quarantine immediately
  • Moderate flakiness: Fails in 5–25% of runs — investigate and fix soon
  • Low/suspected flakiness: Fails in 1–5% of runs — monitor; may be

For each flaky test, classify the likely cause:

Cause classification

Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.

5. Recommend Action

For each flaky test:

Quarantine (High flakiness):

"Quarantine this test immediately. Disable it in CI by adding

@pytest.mark.skip / [Ignore] / GdUnitSkip annotation. Log it in

tests/regression-suite.md quarantine section. The test is now opt-in only.

Fix the root cause before removing quarantine."

Investigate and fix soon (Moderate):

"This test is intermittently unreliable. Root cause appears to be [cause].

Suggested fix: [specific fix based on cause classification]. Do not quarantine

yet — fix the test directly."

Monitor (Low/suspected):

More skills from Donchitos/Claude-Code-Game-Studios

  • AadoptBrownfield audit — do existing artifacts actually work? Numbered migration plan. Unlike /project-stage-detect, checks compliance not existence.
  • Aarchitecture-decisionCreate an ADR documenting a technical decision: context, alternatives considered, consequences.
  • Aarchitecture-reviewTraceability matrix mapping GDD requirements to ADRs. Finds gaps, cross-ADR conflicts, engine compatibility. PASS/CONCERNS/NOT ASSESSED/FAIL.
  • Aart-bibleAuthor the Art Bible — visual identity gating asset production. Run before /map-systems.
  • Aasset-auditAudit assets against naming conventions, file size budgets, format standards. Finds orphaned assets, missing references.
  • Aasset-specPer-asset visual specs plus AI generation prompts from GDDs and character profiles. After the art bible.
  • Abalance-checkFind balance outliers, broken progressions, degenerate strategies, economy imbalances in formulas and data. 'Check game balance'.
  • AbrainstormGuided concept ideation using professional studio techniques, player psychology, creative exploration.
  • Abug-reportStructured bug report from a description, or analyze code for potential bugs. Reproduction steps, severity.
  • Abug-triageRe-evaluate open bugs — priority vs severity, assign to sprints, surface systemic trends. Run when the count grows.
  • AchangelogAuto-generate a changelog from git commits and sprint data. Internal and player-facing versions.
  • Acode-reviewArchitectural code review — coding standards, SOLID, testability, performance concerns.

All agent skills → · MCP servers