Mmcp.market

agent-harness-fault-injection skill

by sickn33·sickn33/agentic-awesome-skills·47k stars·MIT

Use when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.

A100/100content scan

Is the agent-harness-fault-injection skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the agent-harness-fault-injection skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git /tmp/agentic-awesome-skills
mkdir -p ~/.claude/skills
cp -r /tmp/agentic-awesome-skills/plugins/agentic-awesome-skills-claude/skills/agent-harness-fault-injection ~/.claude/skills/agent-harness-fault-injection
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Agent Harness Fault Injection

Overview

Use a deterministic, non-production fault schedule to test whether an agent workflow preserves state, budgets, safety boundaries, and evidence when a dependency fails. The output is a small fault matrix, an event timeline, and a verdict that distinguishes recovered, contained, unrecoverable, and inconclusive runs.

When to Use This Skill

  • Use when a multi-step agent, state machine, loop, or multi-agent workflow has a new recovery path.
  • Use when sandbox execution, an MCP/tool call, a worker, a checkpoint store, or memory can time out or disappear.
  • Use before claiming retry, resume, deadline, isolation, or partial-failure behavior is production-ready.
  • Use when a regression needs reproducible failure evidence instead of a random chaos run.

Do not use this skill against a production target, real user data, live credentials, or an unbounded external service. Convert those cases to a local simulator or an authorized staging harness first.

Safety and Boundary Preconditions

input fixture, timeout, retry budget, deadline, and expected terminal states.

  1. Freeze the workflow revision, model/prompt configuration, tool schemas, seed,

network disabled unless the test explicitly needs a local test server.

  1. Run in a disposable sandbox with synthetic inputs and stubbed tools. Keep

delete real data, revoke real credentials, kill an unrelated process, or mutate a live service to create a failure.

  1. Make every injected failure an in-memory or fixture-controlled event. Never

fixture, or recovery contract makes the verdict inconclusive.

  1. Record the test scope and a run identifier before starting. A missing scope,

Recovery Contract

Write the invariant before injecting a fault. A useful contract names the state that must survive and the side effects that must not repeat:

After recovery, resume from the latest durable checkpoint, preserve the task
identity and safety policy, spend no more than the remaining retry/deadline
budget, and commit each externally visible effect at most once.

Model the workflow with explicit states. For example:

created -> running -> checkpointed -> waiting_for_tool
                       |                |
                       v                v
                    failed <--------- recovering -> resumed -> completed

For each transition, define the owner, durable fields, allowed retry count, and terminal behavior. In-memory values are not checkpoints unless the harness proves they survive the simulated restart.

Fault Matrix

Select the smallest set of faults that covers the new recovery logic. Do not randomize the schedule until a deterministic schedule has passed.

Deterministic Injection Schedule

Use event numbers rather than wall-clock randomness. A schedule should be portable across harnesses:

{
  "seed": "harness-fixture-07",
  "faults": [
    {"event": "tool.call", "ordinal": 2, "kind": "timeout", "tool": "search"},
    {"event": "worker.start", "ordinal": 2, "kind": "restart"},
    {"event": "branch.join", "ordinal": 1, "kind": "partial_failure", "branch": "summarize"}
  ]
}

The harness should emit the schedule, not merely the seed. Keep fault identity separate from the observed error so a wrapper cannot accidentally turn a timeout into a generic failure. Run the same schedule twice and compare the normalized timeline before trying a different schedule.

Recovery Rules by Boundary

Sandbox and MCP/tool failures

attempt count in the evidence.

  • Assign a request id and idempotency key before the call.
  • Distinguish timeout, explicit tool error, invalid output, and policy denial.
  • Retry only the declared retryable classes; preserve the original error and

replay. A read timeout is not proof that a write did not happen.

  • Do not retry a side effect unless the tool contract says the key is safe to

stop scheduling work.

  • When the deadline or retry budget is exhausted, emit one terminal event and

Worker restart and checkpoints

budgets, and the checkpoint sequence before a restart test.

  • Persist task id, workflow version, state name, completed effects, remaining

checkpoint instead of guessing.

  • Reload the newest valid checkpoint and reject a future-version or corrupted

cannot be proven, downgrade the verdict and require reconciliation.

  • Verify that resumption does not replay a committed effect. If exactly-once

Parallel branches

Represent each branch as its own child attempt. The join record must retain success, failure, timeout, and not-started states. Choose one predeclared join policy:

  • all_required: any required branch failure stops the join;
  • best_effort: continue with an explicit degraded marker;
  • compensate: run a bounded compensating action and then stop or resume.

Never let a successful sibling erase a failed branch from the final ledger.

Memory loss

Clear only the ephemeral context named in the schedule. Rebuild from the checkpoint and durable evidence, then check that the agent does not fabricate missing user intent, tool output, or approval. If a required fact is absent, the safe result is inconclusive or a human clarification state.

Budgets and Terminal Verdicts

Track remaining attempts and remaining time after every event. Do not reset a budget on a worker restart or branch retry. Use these verdicts:

contained_failure is not autonomous success. Report it separately from completed work and include the terminal reason.

Evidence Output

Produce one machine-readable record and one concise human summary. Every event should include runid, monotonic seq, logical time, statebefore, stateafter, actor, event, faultid (when injected), attempt, checkpointseq, retryremaining, deadlineremainingms, and a redacted evidence_ref.

{
  "run_id": "fi-2026-08-19-07",
  "verdict": "recovered",
  "invariants": {"resume_from_checkpoint": "pass", "effect_at_most_once": "pass", "budget": "pass"},
  "faults": [{"id": "f1", "kind": "tool_timeout", "at": "tool.call#2", "handled": true}],
  "timeline": [
    {"seq": 4, "event": "checkpoint.write", "checkpoint_seq": 3},
    {"seq": 5, "event": "tool.timeout", "fault_id": "f1", "retry_remaining": 1},
    {"seq": 8, "event": "workflow.completed", "checkpoint_seq": 4}
  ],
  "limitations": ["Tool output was synthetic; no deployed MCP was exercised."]
}

The human summary should state the frozen contract, injected schedule, verdict, failed invariants, budget consumption, and the narrowest next verification. Redact prompts, tokens, private records, and tool payloads; stable references are enough for replay.

Example: Local Harness Run

Fixture: checkout planner / seed harness-fixture-07
Schedule: search timeout on call 2; worker restart after checkpoint 3
Policy: one retry, 2s deadline, all_required branch join

Result: recovered
Proof: checkpoint 3 reloaded, search request key replayed once, no duplicate
commit, deadline remaining 640ms, final ledger contains both branch outcomes.

Best Practices

  • Freeze inputs and schedules so a failure can be replayed from the evidence.
  • Test one boundary at a time, then add a combined schedule for interaction risk.
  • Assert invariants after every recovery transition, not only at final output.
  • Keep attempt-level faults and task-level outcomes in separate ledgers.
  • Treat missing evidence as inconclusive, never as a passing recovery.

Limitations

More skills from sickn33/agentic-awesome-skills

  • A00-andruia-consultantArquitecto de Soluciones Principal y Consultor Tecnológico de Andru.ia. Diagnostica y traza la hoja de ruta óptima para proyectos de IA en español.
  • F007Security audit, hardening, threat modeling (STRIDE/PASTA), Red/Blue Team, OWASP checks, code review, incident response, and infrastructure security for any project.
  • A10-andruia-skill-smithIngeniero de Sistemas de Andru.ia. Diseña, redacta y despliega nuevas habilidades (skills) dentro del repositorio siguiendo el Estándar de Diamante.
  • A20-andruia-niche-intelligenceEstratega de Inteligencia de Dominio de Andru.ia. Analiza el nicho específico de un proyecto para inyectar conocimientos, regulaciones y estándares únicos del sector. Actívalo tras definir el nicho.
  • A2slides-ppt-generatorAI-powered presentation generation via the 2slides API — create slides from text, match a reference image style, summarize documents into decks, add AI voice narration, and export pages/audio. Use for any \"make slides\", \"create a deck\", or \"slides from this document\" request.
  • A3d-web-experienceExpert in building 3D experiences for the web - Three.js, React
  • Aab-test-setupUse when designing an A/B or split test: define the hypothesis, control and variants, estimate sample size, verify tracking, and predeclare metrics and stopping rules.
  • Aab-testingWhen the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
  • Aacceptance-orchestratorUse when a coding task should be driven end-to-end from issue intake through implementation, review, deployment, and acceptance verification with minimal human re-intervention.
  • Aaccess-reviewConduct periodic access reviews and certifications. Implement access
  • Aaccessibility-compliance-accessibility-auditYou are an accessibility expert specializing in WCAG compliance, inclusive design, and assistive technology compatibility. Conduct audits, identify barriers, and provide remediation guidance.
  • Aaccesslint-auditFind and fix WCAG 2.2 accessibility issues. Two modes — report (sweep a codebase or page, produce a prioritized written report, no edits) and fix (audit→edit→verify loop on a target). Prefers direct-CDP live-DOM auditing; falls back to a browser-MCP composition or HTML-string audits.

All agent skills → · MCP servers