Mmcp.market

spark-memory-thermal-ops skill

by wshobson·wshobson/agents·40k stars·MIT

Manage unified memory and thermals during long-running ML jobs on NVIDIA DGX Spark. Use when planning memory headroom for a training run on GB10, when a job OOMs on unified memory, or when monitoring temperature and power during multi-hour training.

A100/100content scan

Is the spark-memory-thermal-ops skill safe?

Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.

No findings.

Install the spark-memory-thermal-ops skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/wshobson/agents.git /tmp/agents
mkdir -p ~/.claude/skills
cp -r /tmp/agents/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops ~/.claude/skills/spark-memory-thermal-ops
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Spark Memory & Thermal Ops

DGX Spark's GB10 chip has one 128GB unified memory (UMA) pool shared by CPU and GPU, and a sustained power ceiling well below its rated figure. Both break discrete-GPU assumptions: headroom isn't what nvidia-smi reports, and a run that starts fast will slow down mid-job with nothing misconfigured. This skill covers planning memory headroom, working an actual OOM, and watching thermals across a long job. For launch-time failure modes (ABI mismatches, flash-attn, playbook breakage), see spark-training-gotchas — this skill assumes the job starts.

Common Issues Quick Reference

When to Use This Skill

before launch — will this model, method, and batch/pack combination fit.

  • Sizing a training run against the 128GB pool

remediation order matters — what to try first, second, third.

  • A run OOMs mid-load or mid-step and the

multi-hour job, deciding whether a slowdown is thermal throttling or something else.

  • Watching temperature and power during a

inference server (vLLM, Ollama) on the same box.

  • Planning to run a trainer alongside an

UMA Memory Model

Spark has no separate GPU VRAM — the GPU and CPU share one 128GB pool. Two consequences:

underreport pressure — or report nothing at all.** Both report CUDA-allocator-visible memory, not the pool's actual state — a box can show headroom in nvidia-smi and still OOM, because page-cache and mmap'd pages the allocator doesn't see consume the same pool. On some driver/setups, the memory query returns [N/A], [N/A] outright instead of a number — a script grepping for a numeric value there gets nothing, not a misleading undercount (see spark-training-gotchas gotcha G3).

  • **nvidia-smi and cudaMemGetInfo

steady state.** Loading safetensors weights mmaps the file, then copies into CUDA tensors — for a window during load, both the mmap'd pages and the CUDA copy count against the pool at once. A model that fits while training can still OOM during load if headroom was sized for the post-load footprint instead of this doubled transient.

  • **Model load is a transient peak, not the

Plan and diagnose with free -g, not nvidia-smi:

free -g | awk 'NR==2 {print "free:", $4, "GB"}'

Rule of thumb: take that free figure, subtract a few GB for OS/driver overhead, and budget against the result — not the 128GB spec number. The worksheet in references/uma-accounting.md accepts parameter count, dtype, and method as input, and returns a memory estimate to compare against known anchors.

Planning Sequence

Before launch, work through these in order:

for the budget.

  1. Read free -g; subtract OS/driver overhead

activations from references/uma-accounting.md.

  1. Estimate weights + optimizer + gradients +

QLoRA, 27B LoRA, 9B full FT), not the estimate alone.

  1. Compare against the closest anchor (70B

with shorter packing or a smaller batch — cheaper than hitting the OOM Ladder mid-run.

  1. If the estimate is close to the budget, start

Example: Sizing a 70B QLoRA Run

A sanity check of the worksheet formula against the ≈40GB anchor:

params = 70e9
weights_gb = params * 0.5 / 1e9      # NF4, step 1
adapter_gb = 0.5                     # step 5, negligible
total_gb = weights_gb + adapter_gb   # + activations
print(f"{total_gb:.0f}GB before activations")

Weights alone land near the ≈40GB anchor — a plan estimating far above that for the same model class is a signal to recheck dtype and method.

The OOM Ladder

When a job OOMs on unified memory, work this ladder in order. Each step is more disruptive than the last — don't skip ahead: reducing batch size is never step 1.

previous run or a large dataset read often accounts for GB of the "missing" headroom. This costs nothing but a rerun and doesn't touch the job's configuration:

  1. Flush the buffer cache. Page cache from a
sync; echo 3 > /proc/sys/vm/drop_caches

Needs root; a between-run reset, not a mid-training step. See spark-training-gotchas (gotcha G3) for the full diagnostic behind this step.

after a flush fails to free enough headroom, cut batch size or packing length — the first step that changes what the run does. Prefer packing length first; it drives activation footprint more directly at long context.

  1. Reduce batch size or packing length. Only

QLoRA.** If flushing and shrinking batch/pack still OOM, drop the method a tier — bf16 LoRA is next, not the reverse. QLoRA's bitsandbytes dequantization buffers are transient CUDA-side allocations that can OOM before an equivalent bf16 LoRA run would, even though QLoRA's steady-state footprint is smaller. A QLoRA OOM is not proof the model doesn't fit.

  1. **Downgrade the method: bf16 LoRA before

Fall back further (smaller model, multi-Spark) only after all three steps and the job still won't fit.

Thermal Monitoring

Multi-hour runs push into Spark's sustained power ceiling, well under the rated figure — expected platform behavior, not a symptom to explain away:

training logs, not after a slowdown is noticed — every 30-60 seconds correlates a throughput drop with a thermal event. Keep the CSV output format assets/thermal-sample.sh writes, so timestamps line up against the log:

  • Sample temperature and power alongside the
bash assets/thermal-sample.sh 30 thermal.log

cap, not a configuration bug.** Don't re-tune batch size or precision to "fix" a plateau that's the box behaving normally under load. If temperature climbs while power stays flat under the rated 240W figure, that's the signature to recognize.

  • **A sustained ~100W power draw is the platform

letting a run silently slow down unrecorded. A run whose per-step time doubles two hours in should show that in the log, correlated against the thermal sample at that timestamp. Full throttling diagnostics: spark-training-gotchas (gotcha G4).

  • Log throttle events explicitly instead of

Concurrent Workloads

Because the 128GB pool is global, eviction happens without either process's logs showing an OOM:

near-capacity** workloads — an uncapped trainer and inference server (vLLM, Ollama) compete for the same pool. A small, capped workload doesn't: a <4GB LoRA fine-tune coexists fine alongside vLLM capped at gpu-memory-utilization<=0.5 — check the other process's cap, not just its presence, before stopping it.

  • The one-heavy-job rule applies to **uncapped or

under uncapped/near-capacity contention, and vice versa — neither logs an error, so a slow run or lost KV cache is a contention symptom to check for. Stop unrelated uncapped servers before a long or full-pool run.

More skills from wshobson/agents

  • Aairflow-dag-patternsBuild production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
  • Aangular-migrationMigrate from AngularJS to Angular using hybrid mode, incremental component rewriting, and dependency injection updates. Use when upgrading AngularJS applications, planning framework migrations, or modernizing legacy Angular code.
  • Aanti-reversing-techniquesUnderstand anti-reversing, obfuscation, and protection techniques encountered during software analysis. Use this skill when analyzing malware evasion techniques, when implementing anti-debugging protections for CTF challenges, when reverse engineering packed binaries, or when building security research tools that need to detect virtualized environments.
  • Aapi-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers. Use when designing new APIs, reviewing API specifications, or establishing API design standards.
  • Aarchitecture-decision-recordsWrite and maintain Architecture Decision Records (ADRs) following best practices for technical decision documentation. Use when documenting significant technical decisions, reviewing past architectural choices, or establishing decision processes.
  • Aarchitecture-patternsImplement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use this skill when designing clean architecture for a new microservice, when refactoring a monolith to use bounded contexts, when implementing hexagonal or onion architecture patterns, or when debugging dependency cycles between application layers.
  • Aasync-python-patternsMaster Python asyncio, concurrent programming, and async/await patterns for high-performance applications. Use when building async APIs, concurrent systems, or I/O-bound applications requiring non-blocking operations.
  • Aauth-implementation-patternsMaster authentication and authorization patterns including JWT, OAuth2, session management, and RBAC to build secure, scalable access control systems. Use when implementing auth systems, securing APIs, or debugging security issues.
  • Aavoid-ai-writingAudit and rewrite prose so it stops reading as machine-generated. Use this skill when asked to remove AI-isms, clean up AI writing, edit a draft for AI tells, audit a README, changelog, release note, PR description, or blog post for machine-sounding prose, or make text sound less like AI. Supports a detect-only mode, a rewrite mode, and an edit-in-place mode, with optional voice and context profiles.
  • Abacktesting-frameworksBuild robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
  • Abazel-build-optimizationOptimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.
  • Abefore-you-buildPre-build product and feature risk review for founders, product managers, and AI-assisted builders. Use this skill when the user is about to build a landing page, MVP, SaaS product, internal tool, agent workflow, or major feature and needs to check demand, positioning, monetization, retention, trust, distribution, and adoption risk before implementation starts.

All agent skills → · MCP servers