Mmcp.market

trace-to-training-data skill

by wshobson·wshobson/agents·40k stars·MIT

Convert evaluation traces and production logs into SFT examples and preference pairs. Use when graded traces or failure examples exist and need to become training data, when applying rejection sampling to model outputs, or when building DPO pairs from passing and failing runs.

A100/100content scan

Is the trace-to-training-data skill safe?

Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.

No findings.

Install the trace-to-training-data skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents
mkdir -p ~/.claude/skills
cp -r /tmp/agents/plugins/llm-finetuning/skills/trace-to-training-data ~/.claude/skills/trace-to-training-data
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Trace To Training Data

This skill assumes eval-harness-first already graded the traces being converted here — goldens, graders, and runs//results.json all exist before conversion starts. This is the flywheel edge that skill names in its own flow: "the same labeled traces become the training set." Conversion happens here; grading already happened upstream.

Input: graded traces — eval/goldens.jsonl plus runs//results.json, each row carrying a task_id, a verdict from the grader, and a reward when the task supports a scalar score (judge score, execution partial-credit, or an RLVR verifier):

{"task_id": "t-042", "trace_id": "t-042-a3",
 "messages": [{"role": "user", "content": "..."}],
 "verdict": "pass", "reward": 0.91,
 "grader": "exact_match"}

Output format: rows shaped exactly like dataset-curation's Format Selection table — SFT messages rows or DPO prompt/chosen/rejected pairs — so this skill's output is that skill's input with no reshaping step in between.

The Principle

The eval harness already did the labeling work: every trace in results.json carries a verdict, and often a reward, before this skill ever touches it. Converting a graded trace into a training row is mechanical — pick a shape from dataset-curation's table, map fields, write JSONL. Curation is the work that remains — which traces clear a quality bar, which pairs are informative, and which rows must never enter the training set at all.

Treat any conversion step that requires re-judging a trace as a sign the harness is missing a grader, not a gap this skill should paper over. A trace with no verdict or reward isn't convertible yet — route it back to eval-harness-first first, don't hand-label it here to unblock conversion.

SFT From Traces

of successful trajectories**, not every passing one. Rank passing traces by reward and take a fraction (the Agent-lightning pattern) rather than every trace that merely cleared the pass bar — a trace that barely passed is a weaker SFT signal than one that scored well above threshold.

  • **Keep the top-reward fraction

become gold SFT examples directly** (the Langfuse pattern) — when a human edits a failing trace's output into a correct one, that correction needs no reward threshold; a human already validated it. Route corrections straight into the SFT set.

  • **Expert-corrected failures

whole-trajectory discard for multi-step traces.** When only some steps in a multi-step trajectory are bad, mask the loss on the bad steps and keep the good ones, rather than discarding the whole trajectory. SRFT reports 32.2% vs. 30.9% on SWE-bench for step-level critic masking over trajectory discard — a real, if modest, gap from the finer-grained cut.

  • **Step-level masking beats

Preference Pairs From Traces

passing-vs-failing trajectories on the SAME task**, never from unrelated best- and worst-scoring traces pulled across different tasks — cross-task pairs teach the model to prefer one task over another, not one response over another.

  • **Build pairs from

μ−2σ of the reward distribution for that task, never the absolute minimum.** preference-optimization's Pair Construction section owns the full selection formula; this skill supplies the graded trajectories it consumes.

  • **Select the rejected member at

cuts pair volume without cutting signal.** Score each candidate pair by chosen-minus-rejected judge delta and keep only the highest-delta subset — the top 5k of a 16.5k candidate pool matched the full pool's downstream result. Build the full candidate set first, then filter by delta; don't cap generation at 5k up front.

  • **Judge-scored delta selection

Hygiene

and redact what's found.** Traces sourced from production logs can carry credentials, API keys, tokens, or customer data — run a secret/PII scan over every SFT and DPO row and redact matches; conversion fails closed (the row is dropped, not shipped with the raw content) if sensitive fields remain after redaction. Never commit secrets.

  • **Scan for secrets and PII before any row ships,

into training data.** Hold every eval/goldens.jsonl ID out of every converted SFT and DPO set — a trace that also appears as a golden trains on the exact item the checkpoint gets graded against later, silently inflating every subsequent eval run.

  • **Eval goldens must never leak

set**, not just within the newly converted rows — exact-match or embedding-similarity, matching dataset-curation's dedup method field, run against whatever training data already exists before this batch merges in.

  • **Dedup against the training

dataset card. Every converted row must trace back to its source runid and trace_id — dataset-curation's Provenance field checks for exactly this link back to trace-to-training-data output; a row with no traceable source isn't ready to merge.

  • **Provenance goes into the

Related Skills

the graded traces this skill converts; a trace with no verdict or reward isn't convertible yet, route it back there before conversion.

  • eval-harness-first — produces

target formats and the dataset card this skill's provenance data feeds; converted rows must match its Format Selection table field names exactly, not an approximation of them.

  • dataset-curation — owns the

consumes the DPO pairs this skill builds and owns the full μ−2σ rejection-selection formula referenced above.

  • preference-optimization —

Worked JSONL-to-JSONL conversions — graded trace to SFT row, trace pair to DPO pair, correction to SFT row, the rejection-sampling loop, and the goldens-holdout check — live in references/conversion-recipes.md.

More skills from wshobson/agents

  • Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
  • Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
  • Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
  • Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
  • Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
  • Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
  • Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
  • Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
  • Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
  • Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
  • Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
  • Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.

All agent skills → · MCP servers