Mmcp.market

lora-qlora-recipes skill

by wshobson·wshobson/agents·40k stars·MIT

Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Use when writing or reviewing a LoRA/QLoRA training configuration, choosing rank/alpha/target modules, or deciding between LoRA, QLoRA, and full fine-tuning.

A100/100content scan

Is the lora-qlora-recipes skill safe?

Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.

No findings.

Install the lora-qlora-recipes skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents
mkdir -p ~/.claude/skills
cp -r /tmp/agents/plugins/llm-finetuning/skills/lora-qlora-recipes ~/.claude/skills/lora-qlora-recipes
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

LoRA & QLoRA Recipes

This skill assumes the routing decision already happened — finetuning-method-selection should have already pointed here because the data shape is demonstrations (SFT), not preference pairs or a verifiable reward signal. What follows is the current best-practice recipe for configuring the adapter itself: which modules to target, how to size rank and alpha, what learning rate to use, and when QLoRA buys real headroom versus when it just adds risk. Dataset preparation and quality checks are a separate concern — see dataset-curation.

Input: a routing decision (SFT via LoRA/ QLoRA) plus a target size class. Output format: a validated adapter config — the kwarg values below, not free-form advice — that llm-finetuning-training-engineer consumes directly when it generates a runnable script.

The Reference Recipe

The reference recipe is "LoRA Without Regret" (Thinking Machines/Schulman, 2025-09), now the settled convention for LoRA/QLoRA SFT.

Target Modules

Target all-linear modules, not just attention:

target_modules = [
    "q_proj", "k_proj", "v_proj", "o_proj",   # attention
    "gate_proj", "up_proj", "down_proj",      # MLP — matters most
]

The MLP layers (gateproj, upproj, down_proj) matter most — attention-only targeting was the older, weaker convention. Dropping modules to save memory is a Failure Mode below, not a valid optimization.

Alpha and Learning Rate

convention (NeurIPS 2025 "intruder dimensions" result). Don't hand-tune alpha independently of rank — derive it from rank every time.

  • loraalpha = 2 r is the settled

full-fine-tune LR. For QLoRA specifically, 2e-4** is the standard starting point. Full hyperparameter tables and worked examples: references/hyperparameters.md.

  • **LoRA learning rate ≈ 10x the equivalent

Rank by Task

Rank is task-shaped, not a single global default:

Higher rank isn't automatically better — it raises capacity to memorize as fast as it raises capacity to generalize. Start at the row matching the task, and only move up a row if the lower rank measurably underfits on held-out eval, not as a default hedge.

Effective Batch Size

Keep effective batch size under 32. This recipe was validated at that scale — pushing effective batch higher is an untested extrapolation, not a free throughput win.

Unsloth Defaults

Unsloth is the reference implementation this plugin assumes as the default fast path — except for messages-shaped conversational SFT with assistantonlyloss=True, where Unsloth 2026.7.x's compiled trainer has no messages-shaped path at all and the plain-TRL escape hatch (references/unsloth-trl-mapping.md) is the default for that combination, not a rare-regression fallback. Its out-of-the-box defaults, and why each one is set that way:

path assumes zero dropout; setting a nonzero value forfeits the fused-kernel speedup.

  • loradropout=0** — the optimized kernel

parameters for negligible quality gain at this rank range.

  • bias="none" — bias terms add adapter

Unsloth's checkpointing variant, not vanilla HF checkpointing; saves roughly 30% VRAM over no checkpointing.

  • usegradientcheckpointing="unsloth" —

optimizer-state memory with negligible quality impact at LoRA/QLoRA adapter scale.

  • optim="adamw8bit"** — 8-bit AdamW cuts

initialization for reproducibility across runs; treat it like any other seed, not a tunable.

  • randomstate** fixed — pins LoRA

These show up together on the getpeftmodel call:

model = FastLanguageModel.get_peft_model(
    model,
    r=32,
    target_modules=target_modules,
    lora_alpha=64,               # 2 * r
    lora_dropout=0,
    bias="none",
    use_gradient_checkpointing="unsloth",
    random_state=3407,
)

Exact kwarg names and their plain-TRL/PEFT equivalents, plus a full worked config including SFTConfig: references/unsloth-trl-mapping.md and references/hyperparameters.md.

LoRA vs QLoRA vs Full FT

BF16 adapters. This is what makes a 65B-class model trainable on 48GB — the quantized base is the memory win, not the adapter itself.

  • QLoRA = NF4-quantized frozen base weights +

it for dense knowledge injection where the goal is changing what the model knows at the weight level, not adapting a behavior. For everything else in this skill's scope, LoRA or QLoRA is the starting assumption.

  • Full fine-tuning is not a default. Reserve

equivalent bf16 LoRA run would**, even though QLoRA's steady-state footprint is smaller — bitsandbytes dequantization buffers are transient CUDA-side allocations that spike during load. A QLoRA OOM is not proof the model doesn't fit; the dgx-spark-ops plugin's spark-memory-thermal-ops skill covers the full OOM remediation ladder (bf16 LoRA is the next thing to try, not a further QLoRA shrink).

  • **On DGX Spark, QLoRA can OOM before an

Failure Modes

in fp16 on hardware that doesn't have solid BF16 support is a known source of loss spikes and silent divergence. Force bf16=True wherever the hardware supports it; don't fall back to fp16 as if it were equivalent. Check hardware support before picking a dtype:

  • fp16 divergence on non-BF16 GPUs. Training
python -c "import torch; print(torch.cuda.is_bf16_supported())"

A rank picked for "SFT at scale" (up to ~256) on a dataset that doesn't have scale behind it memorizes rather than generalizes. Match rank to the Rank by Task table above, not to the largest number available.

  • Rank too high on a small dataset overfits.

quality for negligible savings. The adapter parameters on gateproj/upproj/downproj are a small fraction of total model size — cutting them barely moves memory but measurably hurts quality. If memory is tight, move to QLoRA or reduce rank/batch/pack length before trimming target modules.

  • **Removing target modules to save memory costs

All three failure modes share a pattern: they look like a training-loop bug (loss spikes, plateaus, memorization) but are actually a config choice that contradicts the reference recipe above. Check configuration against this skill before debugging the training loop itself.

References

alpha/LR tables by task type, rsLoRA notes, batch/packing interactions, and a complete worked Unsloth config block.

  • references/hyperparameters.md — full rank/

Unsloth kwarg mapped to its TRL/PEFT equivalent, current TRL API notes, and the escape-hatch rule for when to drop back to plain TRL.

  • references/unsloth-trl-mapping.md — every

Related skills: finetuning-method-selection routes here; dataset-curation covers the data side this skill doesn't; llm-finetuning-training-engineer is the downstream consumer of the config this skill produces.

More skills from wshobson/agents

  • Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
  • Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
  • Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
  • Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
  • Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
  • Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
  • Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
  • Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
  • Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
  • Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
  • Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
  • Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.

All agent skills → · MCP servers