deepspeed skill
Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention
Is the deepspeed skill safe?
Clean: nothing in its files matched our rules. We read 9 files in the folder on 2026-09-28.
No findings.
Install the deepspeed skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git /tmp/AI-Research-SKILLs mkdir -p ~/.claude/skills cp -r /tmp/AI-Research-SKILLs/08-distributed-training/deepspeed ~/.claude/skills/deepspeed
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Deepspeed Skill
Comprehensive assistance with deepspeed development, generated from official documentation.
When to Use This Skill
This skill should be triggered when:
- Working with deepspeed
- Asking about deepspeed features or APIs
- Implementing deepspeed solutions
- Debugging deepspeed code
- Learning deepspeed best practices
Quick Reference
Common Patterns
Pattern 1: DeepNVMe Contents Requirements Creating DeepNVMe Handles Using DeepNVMe Handles Blocking File Write Non-Blocking File Write Parallel File Write Pinned Tensors Putting it together Acknowledgements Appendix Advanced Handle Creation Performance Tuning DeepNVMe APIs General I/O APIs GDS-specific APIs Handle Settings APIs This tutorial will show how to use DeepNVMe for data transfers between persistent storage and tensors residing in host or device memory. DeepNVMe improves the performance and efficiency of I/O operations in Deep Learning applications through powerful optimizations built on Non-Volatile Memory Express (NVMe) Solid State Drives (SSDs), Linux Asynchronous I/O (libaio), and NVIDIA Magnum IOTM GPUDirect® Storage (GDS). Requirements Ensure your environment is properly configured to use DeepNVMe. First, you need to install DeepSpeed version >= 0.15.0. Next, ensure that the DeepNVMe operators are available in the DeepSpeed installation. The asyncio operator is required for any DeepNVMe functionality, while the gds operator is required only for GDS functionality. You can confirm availability of each operator by inspecting the output of dsreport to check that compatible status is [OKAY]. Below is a snippet of dsreport output confirming the availability of both asyncio and gds operators. If asyncio operator is unavailable, you will need to install the appropriate libaio library binaries for your Linux flavor. For example, Ubuntu users will need to run apt install libaio-dev. In general, you should carefully inspect dsreport output for helpful tips such as the following: [WARNING] asyncio requires the dev libaio .so object and headers but these were not found. [WARNING] asyncio: please install the libaio-dev package with apt [WARNING] If libaio is already installed (perhaps from source), try setting the CFLAGS and LDFLAGS environment variables to where it can be found. To enable gds operator, you will need to install NVIDIA GDS by consulting the appropriate guide for bare-metal systems or Azure VMs (coming soon). Creating DeepNVMe Handles DeepNVMe functionality can be accessed through two abstractions: aiohandle and gdshandle. The aiohandle is usable on both host and device tensors. while gdshandle works only on CUDA tensors, but is more efficient. The first step to use DeepNVMe is to create a desired handle. aiohandle requires asyncio operator, while gdshandle requires both asyncio and gds operators. The following snippets illustrate aiohandle and gdshandle creation respectively. ### Create aiohandle from deepspeed.ops.opbuilder import AsyncIOBuilder aiohandle = AsyncIOBuilder().load().aiohandle() ### Create gdshandle from deepspeed.ops.opbuilder import GDSBuilder gdshandle = GDSBuilder().load().gdshandle() For simplicity, the above examples illustrate handle creation using default parameters. We expect that handles created with default parameters to provide good performance in most environments. However, you can see below for advanced handle creation. Using DeepNVMe Handles aiohandle and gdshandle provide identical APIs for storing tensors to files or loading tensors from files. A common feature of these APIs is that they take a tensor and a file path as arguments for the desired I/O operation. For best performance, pinned device or host tensors should be used for I/O operations (see here for details). For brevity, this tutorial will use aiohandle for illustration, but keep in mind that gdshandle works similarly. You can see the available APIs in a Python shell via tab completion on an aiohandle object . This is illustrated using tab completion of h.. >python Python 3.10.12 (main, Jul 29 2024, 16:56:48) [GCC 11.4.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> h = AsyncIOBuilder().load().aiohandle() >>> h. h.asyncpread( h.freecpulockedtensor( h.getoverlapevents( h.getsinglesubmit( h.newcpulockedtensor( h.pwrite( h.syncpread( h.wait( h.asyncpwrite( h.getblocksize( h.getqueuedepth( h.getintraopparallelism( h.pread( h.read( h.syncpwrite( h.write( The APIs of interest for performing I/O operations are those named with pread and pwrite substrings. For brevity, we will focus on the file write APIs, namely syncpwrite, asyncpwrite, and pwrite. We will discuss only syncpwrite and asyncpwrite below because they are specializations of pwrite. Blocking File Write syncpwrite provides the standard blocking semantics of Python file write. The example below illustrates using syncpwrite to store a 1GB CUDA tensor to a local NVMe file. >>> import os >>> os.path.isfile('/localnvme/test1GB.pt') False >>> import torch >>> t=torch.empty(10243, dtype=torch.uint8).cuda() >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> h = AsyncIOBuilder().load().aiohandle() >>> h.syncpwrite(t,'/localnvme/test1GB.pt') >>> os.path.isfile('/localnvme/test1GB.pt') True >>> os.path.getsize('/localnvme/test1GB.pt') 1073741824 Non-Blocking File Write An important DeepNVMe optimization is the non-blocking I/O semantics which enables Python threads to overlap computations with I/O operations. asyncpwrite provides the non-blocking semantics for file writes. The Python thread can later use wait() to synchronize with the I/O operation. asyncwrite can also be used to submit multiple back-to-back non-blocking I/O operations, of which can then be later blocked on using a single wait(). The example below illustrates using asyncpwrite to store a 1GB CUDA tensor to a local NVMe file. >>> import os >>> os.path.isfile('/localnvme/test1GB.pt') False >>> import torch >>> t=torch.empty(10243, dtype=torch.uint8).cuda() >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> h = AsyncIOBuilder().load().aiohandle() >>> h.asyncpwrite(t,'/localnvme/test1GB.pt') >>> h.wait() 1 >>> os.path.isfile('/localnvme/test1GB.pt') True >>> os.path.getsize('/localnvme/test1GB.pt') 1073741824 Warning for non-blocking I/O operations: To avoid data races and corruptions, .wait() must be carefully used to serialize the writing of source tensors, and the reading of destination tensors. For example, the following update of t during a non-blocking file write is unsafe and could corrupt /localnvme/test1GB.pt. >>> t=torch.empty(10243, dtype=torch.uint8).cuda() >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> h = AsyncIOBuilder().load().aiohandle() >>> h.asyncpwrite(t,'/localnvme/test1GB.pt') >>> t += 1 # <--- Data race; avoid by preceding with h.wait() Similar safety problems apply to reading the destination tensor of a non-blocking file read without .wait() synchronization. Parallel File Write An important DeepNVMe optimization is the ability to parallelize individual I/O operations. This optimization is enabled by specifying the desired parallelism degree when constructing a DeepNVMe handle. Subsequent I/O operations with that handle are automatically parallelized over the requested number of host or device threads, as appropriate. I/O parallelism is composable with either the blocking or non-blocking I/O APIs. The example below illustrates 4-way parallelism of a file write using asyncpwrite. Note the use of intraopparallelism argument to specify the desired parallelism degree in handle creation. >>> import os >>> os.path.isfile('/localnvme/test1GB.pt') False >>> import torch >>> t=torch.empty(10243, dtype=torch.uint8).cuda() >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> h = AsyncIOBuilder().load().aiohandle(intraopparallelism=4) >>> h.asyncpwrite(t,'/localnvme/test1GB.pt') >>> h.wait() 1 >>> os.path.isfile('/localnvme/test1GB.pt') True >>> os.path.getsize('/localnvme/test1GB.pt') 1073741824 Pinned Tensors A key part of DeepNVMe optimizations is using direct memory access (DMA) for I/O operations, which requires that the host or device tensor be pinned. To pin host tensors, you can use mechanisms provided by Pytorch or DeepSpeed Accelerators. The following example illustrates writing a pinned CPU tensor to a local NVMe file. >>> import os >>> os.path.isfile('/localnvme/test1GB.pt') False >>> import torch >>> t=torch.empty(10243, dtype=torch.uint8).pinmemory() >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> h = AsyncIOBuilder().load().aiohandle() >>> h.asyncpwrite(t,'/localnvme/test1GB.pt') >>> h.wait() 1 >>> os.path.isfile('/localnvme/test1GB.pt') True >>> os.path.getsize('/localnvme/test1GB.pt') 1073741824 On the other hand,gdshandle provides newpinneddevicetensor() and pindevicetensor() functions for pinning CUDA tensors. The following example illustrates writing a pinned CUDA tensor to a local NVMe file. >>> import os >>> os.path.isfile('/localnvme/test1GB.pt') False >>> import torch >>> t=torch.empty(10243, dtype=torch.uint8).cuda() >>> from deepspeed.ops.opbuilder import GDSBuilder >>> h = GDSBuilder().load().gdshandle() >>> h.pindevicetensor(t) >>> h.asyncpwrite(t,'/localnvme/test1GB.pt') >>> h.wait() 1 >>> os.path.isfile('/localnvme/test1GB.pt') True >>> os.path.getsize('/localnvme/test1GB.pt') 1073741824 >>> h.unpindevicetensor(t) Putting it together We hope that the above material helps you to get started with DeepNVMe. You can also use the following links to see DeepNVMe usage in real-world Deep Learning applications. Parameter swapper in ZeRO-Inference and ZeRO-Infinity. Optimizer swapper in ZeRO-Infinity. Gradient swapper in ZeRO-Infinity. Simple file read and write operations. Acknowledgements This tutorial has been significantly improved by feedback from Guanhua Wang, Masahiro Tanaka, and Stas Bekman. Appendix Advanced Handle Creation Achieving peak I/O performance with DeepNVMe requires careful configuration of handle creation. In particular, the parameters of aiohandle and gdshandle constructors are performance-critical because they determine how efficiently DeepNVMe interacts with the underlying storage subsystem (i.e., libaio, GDS, PCIe, and SSD). For convenience we make it possible to create handles using default parameter values which will provide decent performance in most scenarios. However, squeezing out every available performance in your environment will likely require tuning the constructor parameters, namely blocksize, queuedepth, singlesubmit, overlapevents, and intraopparallelism. The aiohandle constructor parameters and default values are illustrated below: >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> help(AsyncIOBuilder().load().aiohandle()) Help on aiohandle in module asyncio object: class aiohandle(pybind11builtins.pybind11object) | Method resolution order: | aiohandle | pybind11builtins.pybind11object | builtins.object | | Methods defined here: | | init(...) | init(self: asyncio.aiohandle, blocksize: int = 1048576, queuedepth: int = 128, singlesubmit: bool = False, overlapevents: bool = False, intraopparallelism: int = 1) -> None | | AIO handle constructor Performance Tuning As discussed earlier, achieving peak DeepNVMe performance for a target workload or environment requires using optimally configured aiohandle or gdshandle handles. For configuration convenience, we provide a utility called dsnvmetune to automate the discovery of optimal DeepNVMe configurations. dsnvmetune automatically explores a user-specified or default configuration space and recommends the option that provides the best read and write performance. Below is an example usage of dsnvmetune to tune aiohandle data transfers between GPU memory and a local NVVMe SSD mounted on /localnvme. This example used the default configuration space of dsnvmetune for tuning. $ dsnvmetune --nvmedir /localnvme --gpu Running DeepNVMe performance tuning on ['/localnvme/'] Best performance (GB/sec): read = 3.69, write = 3.18 { "aio": { "singlesubmit": "false", "overlapevents": "true", "intraopparallelism": 8, "queuedepth": 32, "blocksize": 1048576 } } The above tuning was executed on a Lambda workstation equipped with two NVIDIA A6000-48GB GPUs, 252GB of DRAM, and a CS3040 NVMe 2TB SDD with peak read and write speeds of 5.6 GB/s and 4.3 GB/s respectively. The tuning required about four and half minutes. Based on the results, one can expect to achieve read and write transfer speeds of 3.69 GB/sec and 3.18 GB/sec respectively by using an aiohandle configured as below. >>> from deepspeed.ops.opbuilder import AsyncIOBuilder >>> h = AsyncIOBuilder().load().aiohandle(blocksize=1048576, queuedepth=32, singlesubmit=False, overlapevents=True, intraopparallelism=8) The full command line options of dsnvmetune can be obtained via the normal -h or --help. usage: dsnvmetune [-h] --nvmedir NVMEDIR [NVMEDIR ...] [--sweepconfig SWEEPCONFIG] [--noread] [--nowrite] [--iosize IOSIZE] [--gpu] [--gds] [--flushpagecache] [--logdir LOGDIR] [--loops LOOPS] [--verbose] options: -h, --help show this help message and exit --nvmedir NVMEDIR [NVMEDIR ...] Directory in which to perform I/O tests. A writeable directory on a NVMe device. --sweepconfig SWEEPCONFIG Performance sweep configuration json file. --noread Disable read performance measurements. --nowrite Disable write performance measurements. --iosize IOSIZE Number of I/O bytes to read/write for performance measurements. --gpu Test tensor transfers between GPU device and NVME device. --gds Run the sweep over NVIDIA GPUDirectStorage operator --flushpagecache Page cache will not be flushed and reported read speeds may be higher than actual Requires sudo access. --logdir LOGDIR Output directory for performance log files. Default is ./aiobenchlogs --loops LOOPS Count of operation repetitions --verbose Print debugging information. DeepNVMe APIs For convenience, we provide listing and brief descriptions of the DeepNVMe APIs. General I/O APIs The following functions are used for I/O operations with both aiohandle and gdshandle. Function Description asyncpread Non-blocking file read into tensor syncpread Blocking file read into tensor pread File read with blocking and non-blocking options asyncpwrite Non-blocking file write from tensor syncpwrite Blocking file write from tensor pwrite File write with blocking and non-blocking options wait Wait for non-blocking I/O operations to complete GDS-specific APIs The following functions are available only for gdshandle Function Description newpinneddevicetensor Allocate and pin a device tensor freepinneddevicetensor Unpin and free a device tensor pindevicetensor Pin a device tensor unpindevicetensor unpin a device tensor Handle Settings APIs The following APIs can be used to probe handle configuration. Function Description getqueuedepth Return queue depth setting getsinglesubmit Return whether singlesubmit is enabled getintraopparallelism Return I/O parallelism degree getblocksize Return I/O block size setting getoverlapevents Return whether overlap_event is enabled Updated: November 5, 2025 Previous Next
libaioPattern 2: Mixture of Experts for NLG models Contents 1. Installation 2. Training NLG+MoE models 2.1. Changes to the model 2.2. Pre-training the Standard MoE model 2.3. Pre-training the PR-MoE model 2.4. Training MoS with reduced model size In this tutorial, we introduce how to apply DeepSpeed Mixture of Experts (MoE) to NLG models, which reduces the training cost by 5 times and reduce the MoE model size by 3 times (details in our Blog). We use the GPT-3 like models in Megatron-LM framework as the example. Before reading this tutorial, we recommend to first read the tutorials about Mixture of Experts and Megatron-LM GPT pre-training. 1. Installation You would need to install DeepSpeed v0.6.0 or higher to use the MoE feature. The MoE for NLG model examples are in the Megatron-DeepSpeed repo under the MoE folder. 2. Training NLG+MoE models 2.1. Changes to the model To apply MoE to the GPT-style model, we made several changes in Megatron framework, mostly in megatron/model/ where we add the MoE layers into the model. 2.2. Pre-training the Standard MoE model We provide example training scripts under examplesdeepspeed/MoE which we used to perform the experiments in our Blog. There are a few new hyperparameters for standard MoE model: --num-experts: the number of experts per MoE layer. In our experiments we set it to 128. Larger number of experts tend to provide better convergence, but it’s a diminishing return. --moe-expert-parallel-size: degree of the MoE expert parallelism. In other words, there will be num-experts/moe-expert-parallel-size experts on each GPU. Thus --moe-expert-parallel-size should be no more than both number of GPUs, and --num-experts. --moe-loss-coeff: scaling coefficient for adding MoE loss to model loss. In our experiments we find that 0.01 is a good setting. --moe-train-capacity-factor, --moe-eval-capacity-factor, --moe-min-capacity: these configs determine how many tokens can a single expert handle. Larger numbers could lead to better convergence, but would also lead to slower training since the load would be more unbalanced on different experts. --disable-moe-token-dropping: this will completely remove the limitation of how many tokens can a single expert handle. For the same reason as above, we only recommend using this during inference/eval. 2.3. Pre-training the PR-MoE model PR-MoE is a new designed MoE models, standing for Pyramid-Residual-MoE, which improves the parameter efficiency up to 3x as compared to standard MoE. Please see our Blog for more details. We provide example training scripts under examplesdeepspeed/MoE. There are a few different hyperparameters for PR-MoE model compared to standard MoE: --num-experts: Instead of providing a single number, to enable Pyramid-MoE, you need to provide a list, whose length is the same as the number of MoE layers. We suggest to use more experts in the latter stage (close to output) of the model. --mlp-type: chosen from [standard, residual]. When it is residual, Residual-MoE is enabled. In addition to the new hyperparameters above for standard MoE and PR-MoE, for NLG+MoE models we found that it’s helpful to lower the learning rate and increase the learning rate decay duration compared to the base dense model. Details of our tuning can be found in the example training scripts. Regarding training data, we are not able to release our internal data but any public data for Megatron-LM pre-training can be directly used to train MoE models (with the caveat that it might not provide the exact same model quality as in our experiments). For example, we evaluated The Pile dataset (pile.eleuther.ai, github.com/EleutherAI/the-pile) for both dense and MoE models. Table 1 below shows that this public data provides similar evaluation results as our internal data. Model size LAMBADA: completion prediction PIQA: commonsense reasoning BoolQ: reading comprehension RACE-h: reading comprehension TriviaQA: question answering WebQs: question answering Dense NLG: 350M, internal data 0.5203 0.6931 0.5364 0.3177 0.0321 0.0157 350M, public Pile 0.5106 0.6589 0.5933 0.3196 0.0257 0.0064 Standard MoE NLG: 350M+MoE-128, internal data 0.6270 0.7459 0.6046 0.3560 0.1658 0.0517 350M+MoE-128, public Pile 0.6128 0.7323 0.6040 0.3349 0.1111 0.0335 PR-MoE NLG: 350M+MoE-128, internal data 0.6365 0.7399 0.5988 0.3569 0.1630 0.0473 PR-MoE + MoS NLG: 350M+MoE-128, internal data 0.6346 0.7334 0.5807 0.3483 0.1369 0.0522 Table 1: Zero-shot evaluation results (last six columns)
More skills from Orchestra-Research/AI-Research-SKILLs
- Aacademic-plottingGenerates publication-quality figures for ML papers from research context. Given a paper section or description, extracts system components and relationships to generate architecture diagrams via Gemini. Given experiment results or data, auto-selects chart type and generates data-driven figures via matplotlib/seaborn. Use when creating any figure for a conference paper.
- Aara-compilerCompiles any research input — PDF papers, GitHub repositories, experiment logs, code directories, or raw notes — into a complete Agent-Native Research Artifact (ARA) with cognitive layer (claims, concepts, heuristics), physical layer (configs, code stubs), exploration graph, and grounded evidence. Use when ingesting a paper or codebase into a structured, machine-executable knowledge package, building an ARA from scratch, or converting research outputs into a falsifiable, agent-traversable form.
- Aara-research-managerRecords research provenance as a post-task epilogue, scanning conversation history at the end of a coding or research session to extract decisions, experiments, dead ends, claims, heuristics, and pivots, and writing them into the ara/ directory with user-vs-AI provenance tags. Use as a session epilogue — never during execution — to maintain a faithful, auditable trace of how a research project actually evolved.
- Aara-rigor-reviewerPerforms ARA Seal Level 2 semantic epistemic review on Agent-Native Research Artifacts, scoring six dimensions (evidence relevance, falsifiability, scope calibration, argument coherence, exploration integrity, methodological rigor) and producing a constructive, severity-ranked report with a Strong Accept-to-Reject recommendation. Use after Level 1 structural validation passes, when an ARA needs an objective epistemic critique before publication or release.
- Aaudiocraft-audio-generationPyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation.
- Aautogpt-agentsAutonomous AI agent platform for building and deploying continuous agents. Use when creating visual workflow agents, deploying persistent autonomous agents, or building complex multi-step AI automation systems.
- AautoresearchOrchestrates end-to-end autonomous AI research projects using a two-loop architecture. The inner loop runs rapid experiment iterations with clear optimization targets. The outer loop synthesizes results, identifies patterns, and steers research direction. Routes to domain-specific skills for execution, supports continuous agent operation via Claude Code /loop and OpenClaw heartbeat, and produces research presentations and papers. Use when starting a research project, running autonomous experiments, or managing a multi-hypothesis research effort.
- Aawq-quantizationActivation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.
- CaxolotlExpert guidance for fine-tuning LLMs with Axolotl - YAML configs, 100+ models, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, multimodal support
- Ablip-2-vision-languageVision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
- Abrainstorming-research-ideasGuides researchers through structured ideation frameworks to discover high-impact research directions. Use when exploring new problem spaces, pivoting between projects, or seeking novel angles on existing work.
- AchromaOpen-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects.