Mmcp.market

eval-harness skill

by a5c-ai·a5c-ai/babysitter·1.8k stars·MIT

Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.

A100/100content scan

Is the eval-harness skill safe?

Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.

No findings.

Install the eval-harness skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/a5c-ai/babysitter.git /tmp/babysitter
mkdir -p ~/.claude/skills
cp -r /tmp/babysitter/library/methodologies/everything-claude-code/skills/eval-harness ~/.claude/skills/eval-harness
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

  • Define test cases with known-correct outputs
  • Run agent against each test case
  • Score: accuracy, completeness, relevance
  • Compare against baseline performance
  • Track performance over time

2. Skill Quality Testing

  • Verify skill instructions produce expected outcomes
  • Test edge cases and boundary conditions
  • Measure consistency across multiple runs
  • Check for harmful or incorrect outputs
  • Validate against ground truth

3. Regression Suite

  • Collection of previously-passing test cases
  • Run after any agent/skill modification
  • Flag regressions with before/after comparison
  • Maintain pass rate threshold (>= 95%)

4. Process Verification

  • End-to-end process execution with known inputs
  • Verify each phase produces expected outputs
  • Check task ordering and dependency satisfaction
  • Measure total execution time

Quality Scoring

Accuracy Score (0-100)

  • Correctness of output vs expected
  • Partial credit for partially correct outputs
  • Penalty for hallucinated or fabricated content

Completeness Score (0-100)

  • Coverage of required output elements
  • Missing sections flagged and scored
  • Bonus for useful additional context

Consistency Score (0-100)

  • Run same input 3 times
  • Compare outputs for semantic similarity
  • Flag inconsistencies

Composite Score

  • (accuracy 0.4 + completeness 0.3 + consistency * 0.3)
  • Threshold: 80 to pass

When to Use

  • After creating new agents or skills
  • After modifying existing agents or skills
  • Periodic quality audits
  • Before promoting skills to production

Agents Used

  • Used by process-level evaluation orchestrators
  • No specific agent dependency (evaluates other agents)

More skills from a5c-ai/babysitter

  • Aadversarial-reviewFresh adversarial code review with binary PASS/FAIL verdicts, evidence citations, and anchoring bias prevention via fresh reviewer spawning.
  • Aagent-boosterWASM-based instant code transforms for simple tasks, achieving 352x speedup over LLM inference with zero cost.
  • Aagent-coordinationCoordinate Crew (persistent) and Polecat (transient) agents using Gas Town's hook-based work distribution and GUPP principle.
  • Aagent-dispatch
  • Aanti-driftHierarchical coordination and drift detection with frequent checkpoints, shared memory coherence validation, role specialization enforcement, and short task cycles.
  • Aarchitecture-design
  • Aarchitecture-patternsSystem and API design guidance covering component boundaries, data flow, integration patterns, and scalability considerations.
  • Aassimilate-popular-workflowsThis skill should be used when the user asks to "find skills in the wild", "assimilate popular workflows", "discover SKILL.md files in repos", "research external skills", "find workflow patterns", "survey the skill landscape", "what skills exist out there", or wants to investigate public repositories for extractable processes, babysitter plugins, and reusable procedural insights. Searches GitHub for SKILL.md files, classifies repos by archetype, and maintains structured research under docs/reference-repos/.
  • Aaudit-trail
  • AbabysitOrchestrate via @babysitter. Use this skill when asked to babysit a run, orchestrate a process or whenever it is called explicitly. (babysit, babysitter, orchestrate, orchestrate a run, workflow, etc.)
  • Ababysit-babysitter-issuesThis skill should be used when the user asks to "babysit issues", "work on assigned issues", "check a5c-agent issues", "process babysitter issues", or wants to find and work on open GitHub issues assigned to a5c-agent in the babysitter repo.
  • Abehavior-contractBug condition/postcondition formalization as testable Behavior Contracts. Defines invariants that must be preserved across fixes.

All agent skills → · MCP servers