spark-training-gotchas skill
Preflight and diagnose the ten known failure modes for ML training on NVIDIA DGX Spark. Use when a training run on DGX Spark fails to start, OOMs below the 128GB limit, slows down mid-run, or before any multi-hour training job on GB10.
Is the spark-training-gotchas skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the spark-training-gotchas skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents mkdir -p ~/.claude/skills cp -r /tmp/agents/plugins/dgx-spark-ops/skills/spark-training-gotchas ~/.claude/skills/spark-training-gotchas
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Spark Training Gotchas
DGX Spark's GB10 chip (Grace Blackwell, SM121, 128GB unified memory, aarch64) has ten recurring failure modes across launch, memory, thermals, bandwidth, and precision. Each is named G1–G10 so it can be checked by number — the numbering is load-bearing for tooling that runs these checks. Read this before a long run, not after hour six.
When to Use This Skill
segfault that doesn't point at the real cause.
- A training run fails to start, with an import error or a
strategy.
- A run OOMs while nvidia-smi still shows headroom.
- Throughput degrades partway through a run that started fine.
- Before any multi-hour or multi-epoch job on GB10.
- Wiring two Sparks together, before picking a parallelism
- Choosing between FP8 and NVFP4 for a Spark-hosted run.
Common Issues Quick Reference
The Ten Gotchas
G1: CUDA 12/13 ABI Mismatch
function, or a segfault on the first .cuda() call.
- SYMPTOM: ImportError: undefined symbol naming a CUDA
ships CUDA 13. pip never checks CUDA ABI, so it surfaces only at import or first kernel launch.
- CAUSE: most PyPI wheels link libcudart.so.12; Spark
CUDA build tag.
- CHECK: references/gotcha-checks.md G1 — the wheel's
use a matched container.
- FIX: reinstall from download.pytorch.org/whl/cu130 or
G2: flash-attn — Skip the pip Build, Watch Unsloth's Auto-Detect
Unsloth may also silently train flash-attn over an explicitly requested SDPA.
- SYMPTOM: pip install flash-attn still fails/hangs.
containers ship a working SM121 flash-attn, and Unsloth auto-prefers it, dropping attn_implementation="sdpa".
- CAUSE: no aarch64/sm_121 wheel for bare pip — but NGC
already present and working.
- CHECK: references/gotcha-checks.md G2 — is flash-attn
NGC — the only reliable override is the monkeypatch in references/gotcha-checks.md G2.
- FIX: bare pip — skip flash-attn, use SDPA (unchanged). On
G3: UMA OOM Below 128GB
nvidia-smi still reports free memory under the 128GB cap — or, on some setups, [N/A] outright instead of a number.
- SYMPTOM: OOM during model load/training while
during safetensors load; QLoRA can OOM earlier than bf16 since dequantization adds transient allocs.
- CAUSE: mmap and the CUDA allocator double-count pages
and /proc/meminfo, not nvidia-smi.
- CHECK: references/gotcha-checks.md G3 — read free -g
sync; echo 3 > /proc/sys/vm/drop_caches — needs root, a between-run reset, not a mid-training step.
- FIX: drop the page cache with
G4: Thermal Throttling
run, or the box spontaneously reboots under sustained load.
- SYMPTOM: throughput drops partway through a multi-hour
240W rated figure; long runs push into that ceiling and throttle or, sometimes, reboot.
- CAUSE: sustained power draw caps around 100W versus the
nvidia-smi --query-gpu=temperature.gpu,power.draw.
- CHECK: references/gotcha-checks.md G4 — sample
climbs, treat throttling as the cause; improve cooling or cap run length.
- FIX: if power plateaus under 240W while temperature
G5: Bandwidth Ceiling
especially, plateau well below expected throughput.
- SYMPTOM: memory-bound workloads, decode-heavy RL loops
measured bandwidth runs 180–192 GB/s.
- CAUSE: 273 GB/s is a spec ceiling, not sustained;
time vs. the measured range, not spec.
- CHECK: references/gotcha-checks.md G5 — observed step
built on the 273 GB/s figure.
- FIX: budget throughput from 180–192 GB/s; revise a plan
G6: Global UMA Resource Contention
mid-run silently, no OOM in its own logs.
- SYMPTOM: a process's KV cache/weights get evicted
one global pool; an uncapped or near-capacity process competes with anything else and can evict it. A small, bounded workload doesn't — a <4GB LoRA coexists fine alongside vLLM capped at gpu-memory-utilization<=0.5.
- CAUSE: unified memory is
More skills from wshobson/agents
- Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
- Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
- Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
- Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
- Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
- Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
- Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
- Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
- Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
- Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
- Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
- Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.