quantized-export skill
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Use after a checkpoint passes promotion, when choosing a quantization format for a target device, or when an exported model fails its smoke test.
Is the quantized-export skill safe?
Clean: nothing in its files matched our rules. We read 2 files in the folder on 2026-09-28.
No findings.
Install the quantized-export skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents mkdir -p ~/.claude/skills cp -r /tmp/agents/plugins/llm-finetuning/skills/quantized-export ~/.claude/skills/quantized-export
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Quantized Export
The last stop after checkpoint-promotion hands off a PROMOTE verdict: a checkpoint that cleared the four-stage gate still isn't deployed until it's exported in the right format for its target runtime and proven to still work post-export. A REJECT verdict never reaches this skill — export starts only from a promoted checkpoint.
Input: a promoted checkpoint (or LoRA adapter) plus the target deployment surface — GPU class, serving stack, and whether long-context/code/math workloads are in scope. Output format: an exported artifact in the chosen format plus a smoke-test diff report comparing 3–5 golden outputs pre-export and post-export.
Format Map
Pick format by hardware and deployment shape, not by habit — the wrong pick either wastes throughput headroom or breaks silently on specific workloads (see Workload Overrides).
and newer.** It preserves near-bf16 quality at roughly half the memory, and it's the safe first choice whenever the target GPU supports it and no edge-device constraint applies.
- **FP8 is the default on Hopper-class GPUs
FP8 hardware support. GPTQ is superseded for new deployments — don't reach for it on a fresh export; AWQ has better accuracy retention at the same bit width and wider current tooling support.
- AWQ INT4 targets older GPUs that predate
an imatrix, is the edge/llama.cpp format.** Use it for local or CPU-adjacent deployment, not for GPU-serving throughput — it optimizes for footprint, not tokens/sec on a datacenter GPU.
- GGUF with Q4K_M quantization, built from
deployments only — and explicitly NOT on GB10. NVFP4 on SM121 (GB10) runs ~32% slower than FP8 because the hardware lacks a native cvt.e2m1x2 path unless the kernel is compiled sm121a. Choosing NVFP4 on a GB10 target is a regression, not an upgrade — pick FP8 there instead.
- **NVFP4 is for Blackwell-at-scale
from quant format.** A merged export folds the adapter into the base weights: larger artifact, no base-model dependency at serve time. LoRA-only keeps the adapter separate: much smaller artifact, but the serving stack must load the exact same base model alongside it — a mismatched or wrong-revision base silently changes outputs. Pick merged when artifact portability matters more than storage; pick LoRA-only when disk footprint or multi-adapter serving matters more.
- **Merged vs. LoRA-only is a separate axis
Worked Picks
The core format-selection tradeoff, read as a lookup table for common scenarios:
# quick decision snippet — see the table above for the full map
hopper_or_newer: fp8
older_gpu: awq-int4
edge_llama_cpp: gguf-q4_k_m+imatrix
gb10_any_workload: fp8-vllm-nightly # never nvfp4 on GB10Workload Overrides
The Format Map above is a default, not a rule that survives every workload. Long-context, code, and math workloads break at INT4 — quantization error compounds across long sequences and precise token-level reasoning in ways that don't show up on short, generic prompts. For any of these three workload classes, stay on FP8 or W8A8 even if the target hardware would otherwise justify INT4 on cost grounds.
similar broad-knowledge benchmarks — they don't stress the failure mode. Measure with the actual task evals — the goldens and graders from eval-harness-first, run through the exported artifact — because INT4 degradation on long-context, code, or math shows up as task-specific failures (dropped context, broken syntax, arithmetic errors) well before it moves a knowledge benchmark.
- Don't validate this override with MMLU or
export on one of these three workload classes, the fix is switching format, not re-tuning the quantization recipe — AWQ and GPTQ variants at the same bit width share the same compounding-error failure mode on these workloads.
- If a task eval regresses after an INT4
The Smoke Test
Export bugs are silent at the file level — a malformed export still produces a loadable artifact, so file-existence checks prove nothing. The smoke test is mandatory for every export, with no exception for a format that "should just work":
target runtime** — vLLM for FP8/AWQ, llama.cpp for GGUF, not a quick sanity load in a different framework than the one that will serve it in production.
- **Load the exported artifact in its actual
pull these from the same eval/goldens.jsonl eval-harness-first maintains, not a fresh ad hoc set.
- Run 3–5 golden prompts through it —
pre-export generation for the same prompt, same deterministic sampling settings — greedy decoding (temperature 0) and a fixed seed, persisted and reused between the pre- and post-export runs, not just nominally identical config. For a lossless export, byte match is the gate — any diff is a bug. For a lossy** (quantized) export, byte match is expected to fail; the gate is task-grader verdict agreement instead — see references/export-commands.md's Smoke-Test Script Skeleton.
- **Compare each output against the
Run this as a gate, not a manual check:
python smoke_test.py "$EXPORT_PATH" \
eval/goldens.jsonl pre-export-outputs.jsonl
# non-zero exit on any pre/post mismatchFailure Signatures
What export bugs actually look like, not a clean pass/fail flag:
run-on output — the chat template baked into the export doesn't match the one the checkpoint was trained and evaluated against, so turn boundaries or special tokens land in the wrong place.
- Template mismatch presents as garbled or
presents as off-template or semantically nonsensical output that still looks fluent — the output head lost precision it needed even though the rest of the network quantized cleanly.
- Wrong quantization applied to lmhead**
Never ship an export that skipped this step — a checkpoint's PROMOTE verdict says the un-exported checkpoint is good; it says nothing about the export pipeline. Re-run on any quant-method or runtime version bump, not only after the first export. Runnable command sequences for every format plus the smoke-test script skeleton: references/export-commands.md.
Related Skills
upstream source for this skill. A checkpoint without a PROMOTE verdict doesn't reach export.
- checkpoint-promotion — the only valid
eval/goldens.jsonl this skill's smoke test draws its 3–5 prompts from, and the task evals the Workload Overrides section requires for long-context/code/math validation.
- eval-harness-first — owns the
references/model-catalog.md is the place to check hardware-class assumptions (which GPU generations a base model targets) before picking a format off the Format Map above.
- finetuning-method-selection — its
Spark users: on GB10, GGUF via llama.cpp works well for local serving, and FP8 serving via vLLM nightly builds is the other proven path — NVFP4 is the one format to avoid there (see the Format Map exception above). Once the dgx-spark-ops plugin is installed, defer Spark-specific serving and thermal questions to its skills rather than re-deriving them here.
More skills from wshobson/agents
- Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
- Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
- Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
- Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
- Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
- Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
- Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
- Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
- Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
- Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
- Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
- Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.