checkpoint-promotion skill
Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.
Is the checkpoint-promotion skill safe?
Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.
No findings.
Install the checkpoint-promotion skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents mkdir -p ~/.claude/skills cp -r /tmp/agents/plugins/llm-finetuning/skills/checkpoint-promotion ~/.claude/skills/checkpoint-promotion
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Checkpoint Promotion
The Phase 5 gate for the whole plugin: a checkpoint that trains cleanly and beats its task metric still doesn't ship without clearing all four stages below. eval-harness-first built the suite re-run here — this skill is where that suite's baseline decides something.
Input: a trained checkpoint, eval/baseline-.json from eval-harness-first, and the frozen eval/drift-suite.yaml. Output format: promotion-report.md — the four-stage evidence plus a terminal PROMOTE or REJECT verdict that /finetune Phase 5 and /promote-checkpoint consume directly.
The Four-Stage Gate
Each stage gates the next — a failure at stage 2 means stage 3 doesn't run. Stages 2 and 3 share one expensive inference pass, so running them concurrently and applying gate order at verdict time is licensed on a deterministic arena (nothing saved by serializing); a judge-based arena should still wait for stage 2 first — that's where the real savings are.
any eval touches the checkpoint: dedup the training set, check for eval-goldens leakage (the exact failure trace-to-training-data's Hygiene section exists to prevent), and scan for label noise. A checkpoint trained on leaked goldens invalidates every later stage.
- Data-quality gate. Before
capability-drift suite.** Re-run eval-harness-first's eval/drift-suite.yaml — MMLU/GSM8K/IFEval plus 200–500 domain-adjacent items — against the checkpoint and diff against baseline-.json per benchmark against the Drift Budget table below.
- **Held-out + frozen
Position-randomized judge, checkpoint vs. base model, same prompts — or the deterministic paired-comparison variant in references/gate-templates.md when every grader in the harness is deterministic (no LLM-judge; position randomization N/A there). A holdout win that loses the live arena does not ship — stage-2 numbers and stage-3 judgments must agree; a win on frozen goldens and a loss in paired comparison is a real signal, not a discrepancy to explain away.
- Paired arena vs. base.
rollout with auto-rollback for any checkpoint reaching production traffic. Local-only users stop at stage 3 — skipping stage 4 for a local deployment is the correct stopping point, not a shortcut.
- Canary. 5–10% stratified
Drift Budget
The >5pt row governs regardless of the others: a checkpoint that gained 8 points on the target task and lost 6 points of general capability still fails here — task improvement never buys back a drift-budget breach.
Item count derives from the budget, not convenience: the strict n for a half-width under half the 5pt hard-fail threshold is ~1,300 at typical accuracy (p≈0.7); n=200 is a pragmatic floor (±6pt half-width at that same p, n=50 ±13pt) — report the half-width with every verdict, and treat a margin smaller than it as REJECT (uncertain), not PASS/HARD FAIL. Full math and a 5-run cautionary example: references/gate-templates.md.
RERUN is not a verdict. A 2–5pt drift only ever produces a PROMOTE or REJECT after the seed-variation rerun completes — PROMOTE requires landing back at ≤1pt (noise); any rerun still
1pt — 2–5pt band or >5pt breach
alike — resolves stage 2 to a hard REJECT. No report may reach the Verdict section with stage 2 still showing RERUN.
Catastrophic Forgetting
Unmanaged LoRA fine-tuning loses real general capability, and stage 2 is what catches it:
unmanaged** — no replay, no regularization.
- **~43% knowledge loss
— some replay or a conservative LR.
- ~10% with basic management
disciplined case.
- ~3% with replay + EWC — the
mix is the standard mitigation** — blend general- domain data into training rather than target-task data alone.
- **10–30% general-data replay
If a checkpoint hits the >5pt hard fail in stage 2, work this escalation ladder in order — the one canonical order this skill and references/gate-templates.md both point to:
fraction — swap rows, don't add them** (adding confounds fraction with total optimizer steps). Dose is not monotonic at small-run scale (<~100 steps) — re-check drift after any swap.
- **Adjust the replay-mix
same rank/LR levers lora-qlora-recipes and preference-optimization tune for the training run, applied here in reverse.
- Lower the learning rate.
- Fewer epochs.
- A smaller LoRA rank — the
This order is a default, not a law: remediation guidance from a single before/after run pair is a hypothesis — label it low-confidence once any lever produces a reversal, and prefer a seed-variation repeat over trusting the next rung blindly. A lever that clears the drift breach but drops a success-criterion metric below target is a two-sided tradeoff for a human, not a reason to keep descending the ladder. Full reasoning and the 5-run trajectory behind both caveats: references/gate-templates.md.
Disclose drift-suite instruction reuse. A replay row copying the drift harness's exact instruction phrasing (not just disjoint source items) makes that benchmark's post-replay score an upper bound — flag it instruction-familiar, or re-probe with a paraphrase, before treating a near-budget pass as clean.
The Verdict
promotion-report.md covers all four stages as sections and must end with a terminal verdict: PROMOTE or REJECT, the evidence that produced it, and exactly one top remediation when the verdict is REJECT. Template: references/gate-templates.md. The terminal contract other skills parse:
## Verdict
REJECT
Evidence: domain-adjacent drift
suite dropped 6.2pt (threshold:
>5pt hard fail) despite +8pt on
the target task.
Top remediation: swap the
replay-mix fraction from 10%
toward 20%, holding step count
constant.error.** A checkpoint that fails stage 2's drift budget or stage 3's arena comparison did its job. Don't treat a REJECT as a failed run needing a rerun of this skill; it's the correct output of a working gate.
- **REJECT is a result, not an
menu.** Evidence sections may list everything observed; the verdict section names the single highest-leverage fix per the escalation ladder above. A report that hedges across three possible fixes hasn't done the prioritization this skill exists to do.
- **One remediation, not a
skill produces a verdict and a report, not a re-triggered training run. A REJECT hands the remediation back to a human decision at finetuning-method-selection or the relevant training skill.
- No auto-retraining. This
Related Skills
drift suite and baseline this skill re-runs and diffs against; no baseline-.json means nothing to gate against.
- eval-harness-first — owns the
valid next step after a PROMOTE verdict.
- quantized-export — the only
lora-qlora-recipes — own the LR and rank levers in the Catastrophic Forgetting escalation path; this skill diagnoses the breach, those skills own the config that caused it.
- preference-optimization and
replay-mix construction recipe the escalation ladder's first rung applies.
- dataset-curation — owns the
Complete promotion-report.md template with all four stages, the drift-suite scoring table, the paired-arena protocol (item count, position randomization, win-rate threshold), and a replay-mix configuration example: references/gate-templates.md.
More skills from wshobson/agents
- Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
- Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
- Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
- Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
- Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
- Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
- Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
- Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
- Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
- Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
- Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
- Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.