eval-harness-first skill
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.
Is the eval-harness-first skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the eval-harness-first skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents mkdir -p ~/.claude/skills cp -r /tmp/agents/plugins/llm-finetuning/skills/eval-harness-first ~/.claude/skills/eval-harness-first
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Eval Harness First
The Phase 0 gate for the whole plugin: finetuning-method-selection and every downstream skill assume this harness exists before a training config gets written. The harness is not a run-end side artifact — it is the data-curation engine. The same labeled traces that build the goldens feed training data, minus an explicit holdout.
Input: production/agent traces if they exist, or a task spec if they don't, plus labelers willing to grade ≥100 examples. Output format: the eval/ directory below — goldens, graders, drift suite, and the base-model baseline that later phases gate on.
The Gate
No eval harness, no fine-tune. Skip to a training config and there is nothing to measure against, nothing to catch regressions, and no labeled data to train on. The flywheel:
synthetic tasks if none exist yet.
- Collect traces — production/agent spans, or
axial coding into 4–8 failure buckets.
- Error analysis — open coding on ≥100 traces,
calibrated LLM-judge only for genuinely subjective criteria.
- One grader per bucket — deterministic first;
an explicit holdout.** Every eval/goldens.jsonl ID stays excluded from training data by ID.
- Prioritize by frequency × severity × value.
- **The labeled traces feed dataset curation, minus
not a different, looser one.
- Train.
- Re-run the same harness on the checkpoint —
production failure modes re-open error analysis.
- Drift detection feeds back to step 2 — new
Steps 2–4 build the harness; steps 5–8 are why it must exist first — it is both the training data source and the checkpoint's exit gate.
Building Goldens
analysis — open coding on ≥100 real traces (read them, tag failures in your own words, no fixed taxonomy yet), then axial coding to collapse those tags into 4–8 named failure buckets. Fewer than 4 means the coding pass was too shallow; more than 8 means buckets need merging. Exception: single-failure-surface tasks (e.g. strict-schema extraction) may land at 1–2 buckets with per-field sub-metrics inside one grader — don't invent artificial splits with no evidence behind them.
- From traces, when they exist: run error
dimension-based generation — enumerate the axes that matter (task type, difficulty, edge case, persona) and sample the cross-product; free- generated prompts cluster around whatever's easiest to write.
- Synthetic, when traces don't exist yet:
eval/goldens.jsonl, diff it in review, tag it per release. It doubles as the CI regression suite.
- Goldens are versioned like code — commit
Graders
One grader per failure bucket from error analysis — not one for the whole eval set. A single blended score hides which bucket regressed.
or execution checks are cheaper, reproducible, and need no calibration.
- Deterministic first. Regex, schema validation,
criteria** — tone, faithfulness, "which response is better" — where no deterministic check can express it.
- **LLM-judge only for genuinely subjective
scale is noisier to calibrate and harder to apply consistently; collapse to pass/fail.
- Binary pass/fail over Likert. A 1–5 or 1–10
over generate-and-extract** — a tight token budget makes generate-and-extract parse-brittle for models that preamble, conflating format compliance with the knowledge being measured. Templates for all four grader shapes and this scoring note: references/grader-templates.md.
- **Drift-suite MMLU-style scoring: prefer logprob
Judge Calibration Is a Prerequisite
Any bucket routed to an LLM-judge needs calibration before its verdicts count for anything beyond exploration — a hard prerequisite, not a nice-to-have. N/A when no bucket routes to a judge — an all-deterministic harness has nothing to calibrate; state that rather than leaving this section unaddressed.
test** (report once, no re-touching after).
- Label ≥100 items, split train/dev/**sealed
number — a judge can hit 90% by always saying "pass" on a skewed set.
- Report TPR and TNR, not one blended accuracy
recalibrate on judge-model change, quarterly regardless.
- Pin the judge to a fixed model snapshot and
than the model under test.**
- **The judge must come from a different model family
advisory-only — flags for human review, never gates a promotion. Full protocol, bias correction, and recalibration checklist: references/judge-calibration.md.
- A judge that misses the agreed TPR/TNR bar ships
The Baseline
Before Phase 1 (method selection) starts, run the full harness — goldens plus the capability-drift suite — against the unmodified base model. This is the number every later checkpoint gets compared against.
eval/baseline-.json is the gate token. No baseline file, no comparison basis for checkpoint-promotion — a checkpoint that "looks better" against nothing measured isn't a finding.
Directory Contract
eval/
├── goldens.jsonl # labeled traces + synthetic goldens, versioned
├── graders/ # one module per failure bucket
│ ├── schema_compliance.py
│ ├── exact_match.py
│ └── rubric_judge.py
├── drift-suite.yaml # frozen benchmarks + 200-500 domain-adjacent items
└── baseline-<model>.json # gate token: harness + drift suite vs the base model
runs/
└── <run-id>/
└── results.json # per-run harness output, one per checkpointeval/ persists across runs and lives outside runs/ — the fixed measuring stick, not a run artifact. runs/ is disposable; eval/ is not. Never let a run script write into eval/. Canonical location: every per-trace results.json — the Phase 0 baseline included — lives at runs//results.json, never under eval/runs/...; an instruction requesting the latter is wrong, not this contract.
Phase 0 Exit Checklist
Before finetuning-method-selection, confirm:
floor for synthetic goldens on a single-failure- surface task — see the Building Goldens exception; bucket count then comes from post-baseline error analysis instead).
- ≥100 traces open-coded; 4–8 failure buckets (N/A
different family (N/A when no bucket routes to an LLM-judge; state that explicitly).
- eval/goldens.jsonl committed and versioned.
- One grader per bucket, deterministic first.
- Judges calibrated — TPR/TNR, snapshot pinned,
- eval/drift-suite.yaml frozen.
- eval/baseline-.json written.
More skills from wshobson/agents
- Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
- Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
- Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
- Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
- Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
- Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
- Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
- Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
- Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
- Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
- Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
- Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.