finetuning-method-selection skill
Decide whether to fine-tune at all, and route to the right method (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) and base model. Use when starting any fine-tuning effort, when unsure whether RAG or prompting would suffice, or when choosing between preference-optimization and reinforcement methods.
Is the finetuning-method-selection skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the finetuning-method-selection skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents mkdir -p ~/.claude/skills cp -r /tmp/agents/plugins/llm-finetuning/skills/finetuning-method-selection ~/.claude/skills/finetuning-method-selection
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Fine-Tuning Method Selection
This is the router skill for the fine-tuning lifecycle: it decides whether fine-tuning is the right tool at all, and if so, which method and which base-model size class. Every other skill in this plugin assumes this routing already happened — start here before opening lora-qlora-recipes, preference-optimization, or grpo-rlvr-training.
When to Use This Skill
framework or base model has been chosen.
- Starting any fine-tuning effort, before a
solve the problem more cheaply than training.
- Unsure whether RAG or prompt engineering would
family) and a reinforcement method (GRPO/RLVR) for the same underlying task.
- Choosing between preference optimization (DPO
before committing to a run.
- Sizing a candidate model/method combination
Quick Reference
Off-Ramps First
Most requests that sound like "fine-tune this" are served better and cheaper elsewhere. Check these off-ramps before opening a training run:
facts that change — prices, docs, current events): route to RAG, not fine-tuning. A fine-tuned model bakes in a snapshot; volatile facts go stale immediately.
- Knowledge-bound and volatile (the gap is
behavior is still being figured out, or changes per request): route to prompt engineering. Fine-tuning locks in a behavior; don't lock in one that hasn't stabilized yet.
- Behavior-bound and shifting (the desired
where continued pretraining (CPT) enters, sized by how much domain text exists:
- Stable, dense domain knowledge: this is
CPT learning rate ≈ 10% of the pretraining LR. CPT is guidance-only in this plugin — sizing and LR guidance live here, but this plugin does not execute a CPT run.
Method Router
Once the off-ramps are ruled out, this is the full decision tree (verbatim from the research this plugin is built on):
New FACTS? volatile → RAG | stable+dense → CPT (LR ~10% of pretrain) → SFT
New BEHAVIOR? shifting → prompt-engineering | stable:
demos → SFT (LoRA/QLoRA, all-linear, α=2r)
preference pairs → DPO (SimPO if length-bias, ORPO if memory-bound)
unpaired 👍/👎 → KTO
verifiable success → RLVR + GRPO (DAPO/GSPO/Dr.GRPO per failure mode)
Deploy: FP8 (Hopper+) | NVFP4 (Blackwell scale) | AWQ (older) | GGUF+imatrix (edge)
BEFORE ANY OF THIS: the eval harness must exist first.Read the tree top-down: answer "new facts or new behavior," then follow the branch that matches the data shape in hand (demos, preference pairs, thumbs up/down, or verifiable success/failure). The data shape picks the method — not the other way around.
Worked Routing Examples
support macros exactly." Behavior is stable and demonstrable from transcripts → demos → SFT**.
- *"Users want the assistant to follow our
reviewer thumbs-up/down, unpaired." → unpaired signal → KTO**, not DPO (DPO needs paired preferences).
- *"We have pairs of good/bad responses from
math problems and we can grade correctness automatically." → verifiable success signal → GRPO+RLVR**, and only after confirming the model succeeds at least sometimes (see Key Routing Facts below).
- *"The model can already solve some of these
pricing page." → volatile facts → RAG**, no training run at all.
- *"We want the model to know this week's
Key Routing Facts
240-H100-run study found method choice worth ~1 percentage point versus ~50 points for model scale, and zero of 20 DPO variants beat vanilla DPO. Don't spend a routing decision agonizing over DPO-variant selection — spend it on getting the data shape and scale right.
- Loss-function choice is low-leverage. A
reasoning.** Preference pairs that encode a subjective judgment (tone, style, "which answer is better") route to DPO. Tasks with a verifiable pass/fail signal (math, code, tool calls) route to GRPO+RLVR instead.
- **DPO is for taste, GRPO+RLVR is for
succeeds.** GRPO and other RL methods sharpen an existing capability — they don't teach one from zero. If the model doesn't yet understand the task or output format, run SFT first; only bring in RL once the model succeeds at least sometimes.
- **RL is not the fix for a model that never
Common Routing Mistakes
change weekly — that's a RAG problem, and fine-tuning will just go stale faster than the source data does.
- Reaching for fine-tuning to fix facts that
the actual bottleneck is data quality or model scale — variant choice is the ~1pp lever, not the ~50pp one.
- Picking a DPO variant before checking whether
rollout — route to SFT first so RL has something to sharpen.
- Starting an RL run on a model that fails every
doesn't know our domain" — check the data volume thresholds first; under 500MB, RAG or RAG+fine-tune iterates faster than a CPT run.
- Treating CPT as the default for "the model
Model Selection
Base-model choice is size-class first, family second, and it goes stale fast — so it lives in exactly one place: references/model-catalog.md. That file is the only place in this plugin (and in the DGX Spark ops plugin) that names a base model family. Neither this skill nor references/memory-math.md names one; both describe models by size class only (for example, "8B-class LoRA," not a model name).
The catalog is dated on purpose — model rankings turn over quarterly. It carries a "last verified" date and a refresh checklist. Before trusting a row, check that date; if stale, work the refresh checklist in the catalog before recommending a model from it.
Precedence when the catalog and a method skill disagree: the catalog's per-row Notes column states hardware/size-class feasibility, not a method recommendation — lora-qlora-recipes's LoRA vs QLoRA vs Full FT table (routed by task shape) governs the actual method choice.
Memory Feasibility
Before committing to a method, size it: total memory ≈ params × dtype bytes + optimizer state + gradients + activations. Work each term for the chosen dtype and method (full fine-tune, LoRA, or QLoRA) — worked worksheets and size-class examples live in references/memory-math.md.
On DGX Spark specifically, unified-memory behavior breaks the naive estimate (transient load peaks, nvidia-smi underreporting, thermal throttling on long runs). Once the dgx-spark-ops plugin is installed, defer Spark-specific feasibility calls to its spark-memory-thermal-ops skill rather than re-deriving them here.
Related Skills
Once this skill has picked a method, hand off to the skill that executes it:
rewards
More skills from wshobson/agents
- Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
- Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
- Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
- Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
- Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
- Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
- Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
- Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
- Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
- Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
- Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
- Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.