Mmcp.market

eval-harness-first skill

by wshobson·wshobson/agents·40k stars·MIT

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.

A100/100content scan

Is the eval-harness-first skill safe?

Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.

No findings.

Install the eval-harness-first skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents
mkdir -p ~/.claude/skills
cp -r /tmp/agents/plugins/llm-finetuning/skills/eval-harness-first ~/.claude/skills/eval-harness-first
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Eval Harness First

The Phase 0 gate for the whole plugin: finetuning-method-selection and every downstream skill assume this harness exists before a training config gets written. The harness is not a run-end side artifact — it is the data-curation engine. The same labeled traces that build the goldens feed training data, minus an explicit holdout.

Input: production/agent traces if they exist, or a task spec if they don't, plus labelers willing to grade ≥100 examples. Output format: the eval/ directory below — goldens, graders, drift suite, and the base-model baseline that later phases gate on.

The Gate

No eval harness, no fine-tune. Skip to a training config and there is nothing to measure against, nothing to catch regressions, and no labeled data to train on. The flywheel:

synthetic tasks if none exist yet.

  1. Collect traces — production/agent spans, or

axial coding into 4–8 failure buckets.

  1. Error analysis — open coding on ≥100 traces,

calibrated LLM-judge only for genuinely subjective criteria.

  1. One grader per bucket — deterministic first;

an explicit holdout.** Every eval/goldens.jsonl ID stays excluded from training data by ID.

  1. Prioritize by frequency × severity × value.
  2. **The labeled traces feed dataset curation, minus

not a different, looser one.

  1. Train.
  2. Re-run the same harness on the checkpoint —

production failure modes re-open error analysis.

  1. Drift detection feeds back to step 2 — new

Steps 2–4 build the harness; steps 5–8 are why it must exist first — it is both the training data source and the checkpoint's exit gate.

Building Goldens

analysis — open coding on ≥100 real traces (read them, tag failures in your own words, no fixed taxonomy yet), then axial coding to collapse those tags into 4–8 named failure buckets. Fewer than 4 means the coding pass was too shallow; more than 8 means buckets need merging. Exception: single-failure-surface tasks (e.g. strict-schema extraction) may land at 1–2 buckets with per-field sub-metrics inside one grader — don't invent artificial splits with no evidence behind them.

  • From traces, when they exist: run error

dimension-based generation — enumerate the axes that matter (task type, difficulty, edge case, persona) and sample the cross-product; free- generated prompts cluster around whatever's easiest to write.

  • Synthetic, when traces don't exist yet:

eval/goldens.jsonl, diff it in review, tag it per release. It doubles as the CI regression suite.

  • Goldens are versioned like code — commit

Graders

One grader per failure bucket from error analysis — not one for the whole eval set. A single blended score hides which bucket regressed.

or execution checks are cheaper, reproducible, and need no calibration.

  • Deterministic first. Regex, schema validation,

criteria** — tone, faithfulness, "which response is better" — where no deterministic check can express it.

  • **LLM-judge only for genuinely subjective

scale is noisier to calibrate and harder to apply consistently; collapse to pass/fail.

  • Binary pass/fail over Likert. A 1–5 or 1–10

over generate-and-extract** — a tight token budget makes generate-and-extract parse-brittle for models that preamble, conflating format compliance with the knowledge being measured. Templates for all four grader shapes and this scoring note: references/grader-templates.md.

  • **Drift-suite MMLU-style scoring: prefer logprob

Judge Calibration Is a Prerequisite

Any bucket routed to an LLM-judge needs calibration before its verdicts count for anything beyond exploration — a hard prerequisite, not a nice-to-have. N/A when no bucket routes to a judge — an all-deterministic harness has nothing to calibrate; state that rather than leaving this section unaddressed.

test** (report once, no re-touching after).

  • Label ≥100 items, split train/dev/**sealed

number — a judge can hit 90% by always saying "pass" on a skewed set.

  • Report TPR and TNR, not one blended accuracy

recalibrate on judge-model change, quarterly regardless.

  • Pin the judge to a fixed model snapshot and

than the model under test.**

  • **The judge must come from a different model family

advisory-only — flags for human review, never gates a promotion. Full protocol, bias correction, and recalibration checklist: references/judge-calibration.md.

  • A judge that misses the agreed TPR/TNR bar ships

The Baseline

Before Phase 1 (method selection) starts, run the full harness — goldens plus the capability-drift suite — against the unmodified base model. This is the number every later checkpoint gets compared against.

eval/baseline-.json is the gate token. No baseline file, no comparison basis for checkpoint-promotion — a checkpoint that "looks better" against nothing measured isn't a finding.

Directory Contract

eval/
├── goldens.jsonl          # labeled traces + synthetic goldens, versioned
├── graders/                # one module per failure bucket
│   ├── schema_compliance.py
│   ├── exact_match.py
│   └── rubric_judge.py
├── drift-suite.yaml        # frozen benchmarks + 200-500 domain-adjacent items
└── baseline-<model>.json   # gate token: harness + drift suite vs the base model
runs/
└── <run-id>/
    └── results.json         # per-run harness output, one per checkpoint

eval/ persists across runs and lives outside runs/ — the fixed measuring stick, not a run artifact. runs/ is disposable; eval/ is not. Never let a run script write into eval/. Canonical location: every per-trace results.json — the Phase 0 baseline included — lives at runs//results.json, never under eval/runs/...; an instruction requesting the latter is wrong, not this contract.

Phase 0 Exit Checklist

Before finetuning-method-selection, confirm:

floor for synthetic goldens on a single-failure- surface task — see the Building Goldens exception; bucket count then comes from post-baseline error analysis instead).

  1. ≥100 traces open-coded; 4–8 failure buckets (N/A

different family (N/A when no bucket routes to an LLM-judge; state that explicitly).

  1. eval/goldens.jsonl committed and versioned.
  2. One grader per bucket, deterministic first.
  3. Judges calibrated — TPR/TNR, snapshot pinned,
  1. eval/drift-suite.yaml frozen.
  2. eval/baseline-.json written.

More skills from wshobson/agents

  • Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
  • Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
  • Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
  • Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
  • Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
  • Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
  • Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
  • Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
  • Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
  • Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
  • Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
  • Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.

All agent skills → · MCP servers