preference-optimization skill
Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.
Is the preference-optimization skill safe?
Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.
No findings.
Install the preference-optimization skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents mkdir -p ~/.claude/skills cp -r /tmp/agents/plugins/llm-finetuning/skills/preference-optimization ~/.claude/skills/preference-optimization
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Preference Optimization
This skill assumes finetuning-method-selection already routed here because the data shape is preference pairs or unpaired thumbs-up/down feedback, not demonstrations (that's lora-qlora-recipes) or a verifiable reward signal (that's grpo-rlvr-training). What follows is method selection among the DPO family, the evidence for how much that selection actually matters, the production training pattern, and how to build the pairs in the first place.
Input: a routing decision (preference optimization) plus preference pairs or unpaired feedback, usually from an SFT checkpoint. Output format: a validated method choice plus a config — the kwarg values in references/method-configs.md, not free-form advice — that llm-finetuning-training-engineer consumes directly.
Method Selection
learning rate of 5e-7 to 1e-6 for 1–2 epochs. This LR is lower than the SFT LR that produced the checkpoint being aligned — porting an SFT- scale LR into a DPO run is the most common misconfiguration here, not an edge case.
- DPO is the safe default. Use β=0.1 and a
constraint, or when there's no separate SFT checkpoint to start from — it's reference-free and fuses the SFT and preference objectives into one loss, skipping the separate SFT pass and the reference-model memory cost DPO carries.
- ORPO routes in when memory is the
binary signal (thumbs-up/down) rather than matched preference pairs — don't force unpaired feedback into synthetic pairs to use DPO instead.
- KTO routes in when feedback is unpaired
off with disciplined sweeping — its published gains are a ceiling reported under a tuned sweep, not a baseline any single config will reproduce. Route here only when there's sweep budget; use DPO instead if there isn't.
- SimPO fixes DPO's length bias but only pays
outside frontier labs. Don't reach for it in a production pipeline — every method above is cheaper and better-supported for the same data shapes.
- Classic RLHF (reward model + PPO) is retired
Worked Examples
preference data, no length-bias complaints yet." → default case → DPO** at β=0.1.
- *"We have an SFT checkpoint and clean paired
nothing is paired." → unpaired signal → KTO**, not DPO — don't synthesize pairs to force DPO onto unpaired data.
- *"Reviewers click thumbs-up/down per response;
plus a DPO reference model." → memory-bound, no separate checkpoint → ORPO**.
- *"GPU budget doesn't cover a separate SFT pass
quality, and there's time to run a sweep." → length bias plus sweep budget → SimPO**. Skip it if the sweep budget isn't actually there.
- *"DPO output favors longer answers regardless of
The Low-Leverage Truth
A 2026 240-H100-run study (arXiv 2603.19335) is the load-bearing evidence behind the table above: loss-function choice is worth roughly 1 percentage point of leverage, model scale is worth roughly 50. Zero of 20 DPO variants tested beat vanilla DPO. Rankings also invert with scale — a variant that wins in a small pilot can lose at deployment size.
Two practical consequences:
DPO-variant bake-offs. The table above is sufficient; deeper variant selection is low-leverage compared to data quality and scale.
- Don't spend a routing decision agonizing over
ranking.** A method comparison run on a small pilot model doesn't transfer to the production size class — re-check the winner once scale changes.
- **Validate at deployment scale before trusting a
This is also why the Method Selection table above is deliberately short: it encodes the ~1pp lever, not a ranking of DPO variants that the same study shows doesn't hold up across scale. Treat any variant-selection advice that isn't in that table — including advice that claims a specific variant "wins" — as unproven until it's been validated at the target deployment size.
Production Pattern: Iterative On-Policy DPO
A single offline DPO pass on a static preference dataset is a starting point, not the production pattern. The policy drifts away from the distribution the pairs were sampled from as training proceeds, and a static dataset goes stale against that drift. Production pipelines run DPO iteratively and on-policy instead:
checkpoint.
- Sample completions from the current policy
judge, or task grader).
- Score or rank the completions (reward model,
the reference model.
- Run a DPO pass using the current checkpoint as
policy and the new reference for the next round.
- The resulting checkpoint becomes both the new
Repeat. Each round's reference model is the prior round's output, not a fixed initial checkpoint — that's what keeps the preference signal on-policy instead of scoring against an increasingly stale distribution.
A single-pass DPO run is still a reasonable first iteration — it just isn't the whole pipeline. Plan for at least one more round once the first checkpoint exists, rather than treating pass one as the finished artifact.
Pair Construction
Build DPO/ORPO pairs from same-task passing-vs-failing trajectories — two attempts at the same underlying task, not unrelated best-and-worst examples pulled from different tasks. Within that trajectory set, select the rejected member at μ−2σ of the reward distribution, never the minimum. Naive best-vs-worst pair construction (max reward vs. absolute minimum) degrades as scale increases; the μ−2σ selection is more robust to the same scale sensitivity the low-leverage study surfaced above.
sorted_by_reward = sort(trajectories, key=reward)
chosen = sorted_by_reward[-1] # highest reward
mu, sigma = mean(rewards), stdev(rewards)
rejected = closest(sorted_by_reward, mu - 2 * sigma)
# NOT sorted_by_reward[0] — the absolute minimum
# is the naive best-vs-worst construction that
# degrades as scale increases.For the mechanics of turning graded traces into these pairs — including rejection sampling and judge-scored delta selection — see trace-to-training-data.
References
Complete TRL config blocks per method — DPOConfig, ORPOConfig, KTOConfig, and the SimPO sweep grid — plus Unsloth wrappers and a catastrophic-forgetting note live in references/method-configs.md. Those configs use the same current-TRL API conventions established in lora-qlora-recipes's references/unsloth-trl-mapping.md (processing_class, not tokenizer=).
references/method-configs.md also carries the catastrophic-forgetting note: a too-high learning rate is the usual cause when a preference-tuned checkpoint loses general capability, and the fix is almost always to drop the LR toward the low end of the range in the Method Selection table above before reaching for any other remediation.
Related skills: finetuning-method-selection routes here once preference pairs or unpaired feedback exist; lora-qlora-recipes produces the SFT checkpoint DPO/KTO/SimPO align (ORPO's fused path can skip it); trace-to-training-data converts passing/failing trajectories into the pairs this skill's Pair Construction section consumes.
More skills from wshobson/agents
- Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
- Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
- Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
- Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
- Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
- Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
- Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
- Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
- Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
- Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
- Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
- Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.