test-suite-architect skill
This skill should be used when establishing comprehensive QA testing processes for any software project. Use when creating test strategies, writing test cases following Google Testing Standards, executing test plans, tracking bugs with P0-P4 classification, calculating quality metrics, or generating progress reports. Includes autonomous execution capability via master prompts and complete documentation templates for third-party QA team handoffs. Implements OWASP security testing and achieves 90% coverage targets.
Is the test-suite-architect skill safe?
Serious findings: read the flagged lines first. We read 10 files in the folder on 2026-09-28.
- high
assets/templates/TEST-CASE-TEMPLATE.md:75Reads credential files (SSH keys, cloud or package-manager tokens) that a skill has no normal reason to touch.
- Path traversal vulnerability (test with `../../../etc/passwd`) - high
references/master_qa_prompt.md:251Reads credential files (SSH keys, cloud or package-manager tokens) that a skill has no normal reason to touch.
Issue: Path traversal vulnerability allows reading /etc/passwd
Install the test-suite-architect skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description. Read the findings above first.
git clone --depth 1 https://github.com/zebbern/claude-code-guide.git /tmp/claude-code-guide mkdir -p ~/.claude/skills cp -r /tmp/claude-code-guide/skills/test-suite-architect ~/.claude/skills/test-suite-architect
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
QA Expert
Establish world-class QA testing processes for any software project using proven methodologies from Google Testing Standards and OWASP security best practices.
When to Use This Skill
Trigger this skill when:
- Setting up QA infrastructure for a new or existing project
- Writing standardized test cases (AAA pattern compliance)
- Executing comprehensive test plans with progress tracking
- Implementing security testing (OWASP Top 10)
- Filing bugs with proper severity classification (P0-P4)
- Generating QA reports (daily summaries, weekly progress)
- Calculating quality metrics (pass rate, coverage, gates)
- Preparing QA documentation for third-party team handoffs
- Enabling autonomous LLM-driven test execution
Quick Start
One-command initialization:
python scripts/init_qa_project.py <project-name> [output-directory]What gets created:
- Directory structure (tests/docs/, tests/e2e/, tests/fixtures/)
- Tracking CSVs (TEST-EXECUTION-TRACKING.csv, BUG-TRACKING-TEMPLATE.csv)
- Documentation templates (BASELINE-METRICS.md, WEEKLY-PROGRESS-REPORT.md)
- Master QA Prompt for autonomous execution
- README with complete quickstart guide
For autonomous execution (recommended): See references/masterqaprompt.md - single copy-paste command for 100x speedup.
Core Capabilities
1. QA Project Initialization
Initialize complete QA infrastructure with all templates:
python scripts/init_qa_project.py <project-name> [output-directory]Creates directory structure, tracking CSVs, documentation templates, and master prompt for autonomous execution.
Use when: Starting QA from scratch or migrating to structured QA process.
2. Test Case Writing
Write standardized, reproducible test cases following AAA pattern (Arrange-Act-Assert):
- Read template: assets/templates/TEST-CASE-TEMPLATE.md
- Follow structure: Prerequisites (Arrange) → Test Steps (Act) → Expected Results (Assert)
- Assign priority: P0 (blocker) → P4 (low)
- Include edge cases and potential bugs
Test case format: TC-[CATEGORY]-[NUMBER] (e.g., TC-CLI-001, TC-WEB-042, TC-SEC-007)
Reference: See references/googletestingstandards.md for complete AAA pattern guidelines and coverage thresholds.
3. Test Execution & Tracking
Ground Truth Principle (critical):
- Test case documents (e.g., 02-CLI-TEST-CASES.md) = authoritative source for test steps
- Tracking CSV = execution status only (do NOT trust CSV for test specifications)
- See references/groundtruthprinciple.md for preventing doc/CSV sync issues
Manual execution:
- Read test case from category document (e.g., 02-CLI-TEST-CASES.md) ← always start here
- Execute test steps exactly as documented
- Update TEST-EXECUTION-TRACKING.csv immediately after EACH test (never batch)
- File bug in BUG-TRACKING-TEMPLATE.csv if test fails
Autonomous execution (recommended):
- Copy master prompt from references/masterqaprompt.md
- Paste to LLM session
- LLM auto-executes, auto-tracks, auto-files bugs, auto-generates reports
Innovation: 100x faster vs manual + zero human error in tracking + auto-resume capability.
4. Bug Reporting
File bugs with proper severity classification:
Required fields:
- Bug ID: Sequential (BUG-001, BUG-002, ...)
- Severity: P0 (24h fix) → P4 (optional)
- Steps to Reproduce: Numbered, specific
- Environment: OS, versions, configuration
Severity classification:
- P0 (Blocker): Security vulnerability, core functionality broken, data loss
- P1 (Critical): Major feature broken with workaround
- P2 (High): Minor feature issue, edge case
- P3 (Medium): Cosmetic issue
- P4 (Low): Documentation typo
Reference: See BUG-TRACKING-TEMPLATE.csv for complete template with examples.
5. Quality Metrics Calculation
Calculate comprehensive QA metrics and quality gates status:
python scripts/calculate_metrics.py <path/to/TEST-EXECUTION-TRACKING.csv>Metrics dashboard includes:
- Test execution progress (X/Y tests, Z% complete)
- Pass rate (passed/executed %)
- Bug analysis (unique bugs, P0/P1/P2 breakdown)
- Quality gates status (✅/❌ for each gate)
Quality gates (all must pass for release):
6. Progress Reporting
Generate QA reports for stakeholders:
Daily summary (end-of-day):
- Tests executed, pass rate, bugs filed
- Blockers (or None)
- Tomorrow's plan
Weekly report (every Friday):
- Use template: WEEKLY-PROGRESS-REPORT.md (created by init script)
- Compare against baseline: BASELINE-METRICS.md
- Assess quality gates and trends
Reference: See references/llmpromptslibrary.md for 30+ ready-to-use reporting prompts.
7. Security Testing (OWASP)
Implement OWASP Top 10 security testing:
Coverage targets:
- A01: Broken Access Control - RLS bypass, privilege escalation
- A02: Cryptographic Failures - Token encryption, password hashing
- A03: Injection - SQL injection, XSS, command injection
- A04: Insecure Design - Rate limiting, anomaly detection
- A05: Security Misconfiguration - Verbose errors, default credentials
- A07: Authentication Failures - Session hijacking, CSRF
- Others: Data integrity, logging, SSRF
Target: 90% OWASP coverage (9/10 threats mitigated).
Each security test follows AAA pattern with specific attack vectors documented.
Day 1 Onboarding
For new QA engineers joining a project, complete 5-hour onboarding guide:
Read: references/day1_onboarding.md
Timeline:
More skills from zebbern/claude-code-guide
- Aacademic-paper-reviewerSimulates academic peer review, evaluating papers across Originality, Methodology, Results, and Writing to provide Major/Minor Revision recommendations with actionable feedback. Triggers when a user asks to \"review my paper,\" \"simulate peer review,\" or \"give my paper a peer review.
- Aactive-directory-attacksThis skill should be used when the user asks to "attack Active Directory", "exploit AD", "Kerberoasting", "DCSync", "pass-the-hash", "BloodHound enumeration", "Golden Ticket", "Silver Ticket", "AS-REP roasting", "NTLM relay", or needs guidance on Windows domain penetration testing.
- Capi-fuzzing-bug-bountyThis skill should be used when the user asks to "test API security", "fuzz APIs", "find IDOR vulnerabilities", "test REST API", "test GraphQL", "API penetration testing", "bug bounty API testing", or needs guidance on API security assessment techniques.
- Aapi-shape-explorerGenerate multiple radically different interface designs for a module using parallel sub-agents. Use when user wants to design an API, explore interface options, compare module shapes, or mentions "design it twice".
- Aaudit-flowInteractive system flow tracing across CODE, API, AUTH, DATA, NETWORK layers with SQLite persistence and Mermaid export. Use for security audits, compliance documentation, flow tracing, feature ideation, brainstorming, debugging, architecture reviews, or incident post-mortems. Triggers on audit, trace flow, document flow, security review, debug flow, brainstorm, architecture review, post-mortem, incident review.
- Aauthentication-patternsAuthentication patterns: session vs JWT vs OAuth comparison, provider selection (NextAuth, Clerk, Supabase Auth), security checklist, and common mistakes. Use when implementing auth, reviewing auth flows, or choosing auth providers.
- Aaws-penetration-testingThis skill should be used when the user asks to "pentest AWS", "test AWS security", "enumerate IAM", "exploit cloud infrastructure", "AWS privilege escalation", "S3 bucket testing", "metadata SSRF", "Lambda exploitation", or needs guidance on Amazon Web Services security assessment.
- Abroken-authenticationThis skill should be used when the user asks to "test for broken authentication vulnerabilities", "assess session management security", "perform credential stuffing tests", "evaluate password policies", "test for session fixation", or "identify authentication bypass flaws". It provides comprehensive techniques for identifying authentication and session management weaknesses in web applications.
- Cburp-suite-testingThis skill should be used when the user asks to "intercept HTTP traffic", "modify web requests", "use Burp Suite for testing", "perform web vulnerability scanning", "test with Burp Repeater", "analyze HTTP history", or "configure proxy for web testing". It provides comprehensive guidance for using Burp Suite's core features for web application security testing.
- AcachingCaching strategies — invalidation, TTL guidelines, cache keys, cache layers, and when not to cache. Use when implementing or reviewing caching logic.
- Achart-imageGenerate publication-quality PNG chart images from data, supporting line, bar, area, candlestick, pie, and heatmap charts. Triggers when the user asks to visualize data, create a graph, plot a time series, or generate a chart for a report, alert, or dashboard. Runs as a lightweight, headless Node.js process without a browser.
- Dcloud-penetration-testingThis skill should be used when the user asks to "perform cloud penetration testing", "assess Azure or AWS or GCP security", "enumerate cloud resources", "exploit cloud misconfigurations", "test O365 security", "extract secrets from cloud environments", or "audit cloud infrastructure". It provides comprehensive techniques for security assessment across major cloud platforms.