soak-test skill
Soak test protocol for extended play — what to observe and log for slow leaks, fatigue, late-appearing edge cases.
Is the soak-test skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the soak-test skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git /tmp/Claude-Code-Game-Studios mkdir -p ~/.claude/skills cp -r /tmp/Claude-Code-Game-Studios/.claude/skills/soak-test ~/.claude/skills/soak-test
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
!bash "${CLAUDESKILLDIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation
Automation mode: Resolve modes.automation (project.local.yaml → project.yaml → default collaborative). Every AskUserQuestion call and every file write follows .claude/docs/automation-modes.md (collaborative asks always · guided major-only · autonomous logs and proceeds; automationalwaysask categories always prompt).
Soak Test
A soak test (also called an endurance test) is an extended play session run with specific observation goals. Unlike a smoke check (broad critical path, ~10 min) or a single-feature playtest (~30 min), a soak test runs for 30 minutes to several hours to surface:
of a mechanic (inventory full, score overflow, AI state corruption)
- Memory leaks — gradual heap growth that only appears after scene transitions
- Performance drift — frame time degradation that worsens over time
- State accumulation bugs — issues that only appear after N repetitions
repetitive over extended play
- Fun fatigue — mechanics that feel good in a first session but grow
- Content exhaustion — the point where players run out of novel content
This skill generates the observation protocol and analysis harness — the human does the actual playing.
Output: production/qa/soak-test-[date]-[duration].md
When to run:
- Polish phase — before /gate-check release
- After fixing a memory or stability issue (regression soak)
- When extended play has not been formally tracked
1. Parse Arguments
Duration (default: 1h):
- 30m — short soak; suitable for testing a single mechanic or scene
- 1h — standard soak; covers most common leak categories
- 2h — extended soak; recommended for first full Polish soak
- 4h — deep soak; required for games with long session design (RPGs, sims)
Focus (default: all):
- memory — focus on heap size, object count, leak patterns
- stability — focus on crash/freeze/hang detection
- balance — focus on fun fatigue, content exhaustion, difficulty perception
- all — all of the above
2. Load Context
Read:
guidance) and performance.* budgets (memory ceiling, target FPS); for any key absent or empty (including when project.yaml has no performance or engine block), fall back to .claude/docs/technical-preferences.md
- project.yaml — engine.name (for engine-specific memory monitoring
soak duration), core loop description
- design/gdd/game-concept.md — intended session length (for comparison against
(to avoid re-documenting known issues)
- Most recent file in production/qa/playtests/ — prior playtest findings
(to understand what has been formally tested vs. what the soak covers)
- Most recent file in production/qa/qa-plan-*.md — current sprint test coverage
Note any performance budget targets (performance.* from project.yaml, else .claude/docs/technical-preferences.md):
- Memory ceiling: [N MB, or "not set"]
- Target FPS: [N, or "not set"]
- Frame budget: [N ms, or "not set"]
3. Define Observation Checkpoints
Based on duration, generate timed checkpoints:
30m soak: T+0, T+10, T+20, T+30 1h soak: T+0, T+15, T+30, T+45, T+60 2h soak: T+0, T+20, T+40, T+60, T+80, T+100, T+120 4h soak: T+0, T+30, T+60, T+90, T+120, T+180, T+240
At each checkpoint, the observer records the observation items defined in Phase 4.
4. Generate the Soak Test Protocol
Memory / Stability observation items (if focus = memory or all)
Engine-specific monitoring guidance.
Record the unit the tool shows; never convert, and never assume one.
A soak test looks for growth, so every threshold below is a *ratio or a
delta against this session's own T+0 baseline* — which is unit-agnostic and
stays correct however the editor reports the number. Write the unit down at
T+0 exactly as displayed and use it consistently for the rest of the run.
This replaces a note asserting the return units of
Performance.getmonitor — NOT SOURCEABLE from docs/engine-reference/**,
the identifier appears nowhere in it — a claim that sat one line under a row asking
the tester to record "Static Memory (KB)". A wrong units claim in a leak
detector is off by 1024× in the one measurement the protocol exists to take,
and it would read as a plausible instruction throughout. Deltas need no such
claim, so the safest fix was to stop needing it.
Godot 4:
Object Count → Objects across checkpoints
- Open Debugger → Monitors tab; track Memory → Static Memory and
(some growth on load is expected; sustained growth indicates a leak)
- Record: Static Memory (unit as displayed), Object Count, Orphan Nodes count
- Alert threshold: Memory growth > 20% from T+0 after the first 15 minutes
its T+0 value after a scene unload. A ratio hides that; a non-zero floor that keeps rising is a leak regardless of units
- Orphan Nodes is the one absolute number worth watching: it should return to
Unity:
(units as displayed)
- Open Memory Profiler (Window → Analysis → Memory Profiler)
- Record: Total Reserved Memory, GC Allocated, Object Count at each checkpoint
a monotonicity check, deliberately unit-free
More skills from Donchitos/Claude-Code-Game-Studios
- AadoptBrownfield audit — do existing artifacts actually work? Numbered migration plan. Unlike /project-stage-detect, checks compliance not existence.
- Aarchitecture-decisionCreate an ADR documenting a technical decision: context, alternatives considered, consequences.
- Aarchitecture-reviewTraceability matrix mapping GDD requirements to ADRs. Finds gaps, cross-ADR conflicts, engine compatibility. PASS/CONCERNS/NOT ASSESSED/FAIL.
- Aart-bibleAuthor the Art Bible — visual identity gating asset production. Run before /map-systems.
- Aasset-auditAudit assets against naming conventions, file size budgets, format standards. Finds orphaned assets, missing references.
- Aasset-specPer-asset visual specs plus AI generation prompts from GDDs and character profiles. After the art bible.
- Abalance-checkFind balance outliers, broken progressions, degenerate strategies, economy imbalances in formulas and data. 'Check game balance'.
- AbrainstormGuided concept ideation using professional studio techniques, player psychology, creative exploration.
- Abug-reportStructured bug report from a description, or analyze code for potential bugs. Reproduction steps, severity.
- Abug-triageRe-evaluate open bugs — priority vs severity, assign to sprints, surface systemic trends. Run when the count grows.
- AchangelogAuto-generate a changelog from git commits and sprint data. Internal and player-facing versions.
- Acode-reviewArchitectural code review — coding standards, SOLID, testability, performance concerns.