v0.55.8

Try our Chrome extension

Chrome store icon Chrome Webstore

Easily add the current web-page from your browser directly into your changedetection.io tool, more great features coming soon!

Changedetection.io needs your support!

You can help us by supporting changedetection.io on these platforms;

The more popular changedetection.io is, the more time we can dedicate to adding amazing features!

Many thanks :)

changedetection.io team

  • Cannot set language without session cookie
Not yet seconds ago
            False
        
Not yet seconds ago
Current erroring screenshot from most recent request

Triggered text Ignored text Blocked text

3 hours ago
    Skip to content

      Navigation Menu

            Sign in Appearance settings
                * Platform
                        + AI CODE CREATION
                            o GitHub Copilot Write better code with AI
                            o GitHub Copilot app Direct agents from issue to merge
                            o MCP Registry Integrate external tools
                        + DEVELOPER WORKFLOWS
                            o Actions Automate any workflow
                            o Codespaces Instant dev environments
                            o Issues Plan and track work
                            o Code Review Manage code changes
                            o Code Quality Enforce quality at merge
                        + APPLICATION SECURITY
                            o GitHub Advanced Security Find and fix vulnerabilities
                            o Code security Secure your code as you build
                            o Secret protection Stop leaks before they start
                        + EXPLORE
                            o Why GitHub
                            o Documentation
                            o Blog
                            o Changelog
                            o Marketplace
                      View all features
                * Solutions
                        + BY COMPANY SIZE
                            o Enterprises
                            o Small and medium teams
                            o Startups
                            o Nonprofits
                        + BY USE CASE
                            o App Modernization
                            o DevSecOps
                            o DevOps
                            o CI/CD
                            o View all use cases
                        + BY INDUSTRY
                            o Healthcare
                            o Financial services
                            o Manufacturing
                            o Government
                            o View all industries
                      View all solutions
                * Resources
                        + EXPLORE BY TOPIC
                            o AI
                            o Software Development
                            o DevOps
                            o Security
                            o View all topics
                        + EXPLORE BY TYPE
                            o Customer stories
                            o Events & webinars
                            o Ebooks & reports
                            o Business insights
                            o GitHub Skills
                        + SUPPORT & SERVICES
                            o Documentation
                            o Customer support
                            o Community forum
                            o Trust center
                            o Partners
                      View all resources
                * Open Source
                        + COMMUNITY
                            o GitHub Sponsors Fund open source developers
                        + PROGRAMS
                            o Security Lab
                            o Maintainer Community
                            o Accelerator
                            o GitHub Stars
                            o Archive Program
                        + REPOSITORIES
                            o Topics
                            o Trending
                            o Collections
                * Enterprise
                        + ENTERPRISE SOLUTIONS
                            o Enterprise platform AI-powered developer platform
                        + AVAILABLE ADD-ONS
                            o GitHub Advanced Security Enterprise-grade security features
                            o Copilot for Business Enterprise-grade AI features
                            o Premium Support Enterprise-grade 24/7 support
              * Pricing
                Type / to search
                Sign in
              Sign up Appearance settings
      You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert

              Uh oh!

              There was an error while loading. Please reload this page.

              deepspeedai / DeepSpeed Public
              * Notifications You must be signed in to change notification settings
              * Fork 4.9k
                * Star 42.9k
          * Code
          * Issues 1.2k
          * Pull requests 146
          * Discussions
          * Actions
          * Projects
          * Security and quality 0
          * Insights
          Additional navigation options
                  * Code
                  * Issues
                  * Pull requests
                  * Discussions
                  * Actions
                  * Projects
                  * Security and quality
                  * Insights
                                              master
                                          Branches Tags
                                            Go to file
                                        Code
                                          Open more actions menu

                                        Folders and files

                                          Name                             Name                               Last commit message    Last commit date
                                            Latest commit                
                                                                         
                                                History                  
                                                                         
                                                3,295 Commits            
                                                  3,295 Commits          
                                                  .github                          .github                                                           
                                                  accelerator                      accelerator                                                       
                                                  azure                            azure                                                             
                                                  benchmarks                       benchmarks                                                        
                                                  bin                              bin                                                               
                                                  blogs                            blogs                                                             
                                                  ci                               ci                                                                
                                                  csrc                             csrc                                                              
                                                  deepspeed                        deepspeed                                                         
                                                  docker                           docker                                                            
                                                  docs                             docs                                                              
                                                  examples                         examples                                                          
                                                  op_builder                       op_builder                                                        
                                                  release                          release                                                           
                                                  requirements                     requirements                                                      
                                                  scripts                          scripts                                                           
                                                  tests                            tests                                                             
                                                  .clang-format                    .clang-format                                                     
                                                  .flake8                          .flake8                                                           
                                                  .gitignore                       .gitignore                                                        
                                                  .gitmodules                      .gitmodules                                                       
                                                  .pre-commit-config.yaml          .pre-commit-config.yaml                                           
                                                  .pylintrc                        .pylintrc                                                         
                                                  .readthedocs.yml                 .readthedocs.yml                                                  
                                                  .style.yapf                      .style.yapf                                                       
                                                  AGENTS.md                        AGENTS.md                                                         
                                                  CLAUDE.md                        CLAUDE.md                                                         
                                                  CODEOWNERS                       CODEOWNERS                                                        
                                                  CODE_OF_CONDUCT.md               CODE_OF_CONDUCT.md                                                
                                                  COMMITTERS.md                    COMMITTERS.md                                                     
                                                  CONTRIBUTING.md                  CONTRIBUTING.md                                                   
                                                  GOVERNANCE.md                    GOVERNANCE.md                                                     
                                                  LICENSE                          LICENSE                                                           
                                                  MANIFEST.in                      MANIFEST.in                                                       
                                                  MANIFEST_win.in                  MANIFEST_win.in                                                   
                                                  Makefile                         Makefile                                                          
                                                  README.md                        README.md                                                         
                                                  SECURITY.md                      SECURITY.md                                                       
                                                  THIRD_PARTY_NOTICES.md           THIRD_PARTY_NOTICES.md                                            
                                                  build_win.bat                    build_win.bat                                                     
                                                  environment.yml                  environment.yml                                                   
                                                  install.sh                       install.sh                                                        
                                                  setup.cfg                        setup.cfg                                                         
                                                  setup.py                         setup.py                                                          
                                                  version.txt                      version.txt                                                       
                                            View all files               
                                          

                                              Repository files navigation

                                                * 
                                                * README
                                                * Code of conduct
                                                * Contributing
                                                * Apache-2.0 license
                                                * Security
                                                More items

                                                Office Hours

                                              DeepSpeed hosts regular office hours on the last Tuesday of each month at 12:00 America/New_York to discuss development plans, features, etc. This meeting is public for anyone to join and ask questions. The meeting is hosted on Zoom and can be joined here.

                                                Latest News

                                                * [2026/05] Using Muon Optimizer with DeepSpeed

                                                * [2026/05] System DMA (SDMA) for ZeRO-3: offload collectives off compute units on AMD GPUs for better overlap

                                                * [2026/03] DeepSpeed Team gave a tutorial at ASPLOS 2026 titled "Building Efficient Large-Scale Model Systems with DeepSpeed: From Open-Source Foundations to Emerging Research"

                                                * [2026/03] Our SuperOffload work received an Honorable Mention for the ASPLOS 2026 Best Paper Award

                                                * [2025/12] DeepSpeed Core API updates: PyTorch-style backward and low-precision master states

                                                * [2025/11] DeepSpeed ZeRO++ powers large-scale distillation training of LLMs for Recommendation Systems at LinkedIn

                                                * [2025/10] We hosted the Ray x DeepSpeed Meetup at Anyscale. We shared our most recent work on SuperOffload, ZenFlow, Muon Optimizer Support, Arctic Long Sequence Training and DeepCompile. Please find the meetup slides here.

                                                * [2025/10] SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips

                                                * [2025/10] Study of ZenFlow and ZeRO offload performance with DeepSpeed CPU core binding

                                                * [2025/08] ZenFlow: Stall-Free Offloading Engine for LLM Training

                                                * [2025/06] Arctic Long Sequence Training (ALST) with DeepSpeed: Scalable And Efficient Training For Multi-Million Token Sequences

                                                * [2025/06] DeepNVMe: Affordable I/O scaling for Deep Learning Applications

                                              More news
                                                * [2025/04] DeepCompile: Unlocking Compiler Optimization for Distributed Training
                                                * [2025/03] DeepSpeed AutoTP: Automatic Tensor Parallel Training of Hugging Face models
                                                * [2024/12] Ulysses-Offload: Democratizing Long Context LLM Training

                                                Extreme Speed and Scale for DL Training

                                              DeepSpeed enabled the world's most powerful language models (at the time of this writing) such as MT-530B and BLOOM. DeepSpeed offers a confluence of system innovations, that has made large scale DL training effective, and efficient, greatly improved ease of use, and redefined the DL training landscape in terms of scale that is possible. These innovations include ZeRO, ZeRO-Infinity, 3D-Parallelism, Ulysses Sequence Parallelism, DeepSpeed-MoE, etc.

                                                DeepSpeed Adoption

                                              DeepSpeed was an important part of Microsoft’s AI at Scale initiative to enable next-generation AI capabilities at scale, where you can find more information here.

                                              DeepSpeed has been used to train many different large-scale models, below is a list of several examples that we are aware of (if you'd like to include your model please submit a PR):

                                                * Megatron-Turing NLG (530B)
                                                * Jurassic-1 (178B)
                                                * BLOOM (176B)
                                                * GLM (130B)
                                                * xTrimoPGLM (100B)
                                                * YaLM (100B)
                                                * GPT-NeoX (20B)
                                                * AlexaTM (20B)
                                                * Turing NLG (17B)
                                                * METRO-LM (5.4B)

                                              DeepSpeed has been integrated with several different popular open-source DL frameworks such as:

                                                Documentation              
                                                Transformers with DeepSpeed
                                                Accelerate with DeepSpeed  
                                                Lightning with DeepSpeed   
                                                MosaicML with DeepSpeed    
                                                Determined with DeepSpeed  
                                                MMEngine with DeepSpeed    
                                              

                                                Build Pipeline Status

                                              Description        Status
                                              NVIDIA                   
                                              AMD                      
                                              CPU                      
                                              Intel Gaudi              
                                              Intel XPU                
                                              Integrations             
                                              Misc                     
                                              Huawei Ascend NPU        
                                              

                                                Installation

                                              The quickest way to get started with DeepSpeed is via pip, this will install the latest release of DeepSpeed which is not tied to specific PyTorch or CUDA versions. DeepSpeed includes several C++/CUDA extensions that we commonly refer to as our 'ops'. By default, all of these extensions/ops will be built just-in-time (JIT) using torch's JIT C++ extension loader that relies on ninja to build and dynamically link them at runtime.

                                                Requirements

                                                * PyTorch must be installed before installing DeepSpeed.
                                                * For full feature support we recommend a version of PyTorch that is >= 2.0 and ideally the latest PyTorch stable release.
                                                * A CUDA or ROCm compiler such as nvcc or hipcc used to compile C++/CUDA/HIP extensions.
                                                * Specific GPUs we develop and test against are listed below, this doesn't mean your GPU will not work if it doesn't fall into this category it's just DeepSpeed is most well tested on the following:
                                                    + NVIDIA: Pascal, Volta, Ampere, and Hopper architectures
                                                    + AMD: MI100 and MI200

                                                Contributed HW support

                                                * DeepSpeed now support various HW accelerators.
                                              Contributor  Hardware                             Accelerator Name  Contributor validated  Upstream validated
                                              Huawei       Huawei Ascend NPU                    npu               Yes                    No                
                                              Intel        Intel(R) Gaudi(R) 2 AI accelerator   hpu               Yes                    Yes               
                                              Intel        Intel(R) Xeon(R) Processors          cpu               Yes                    Yes               
                                              Intel        Intel(R) Data Center GPU Max series  xpu               Yes                    Yes               
                                              Tecorigin    Scalable Data Analytics Accelerator  sdaa              Yes                    No                
                                              

                                                PyPI

                                              We regularly push releases to PyPI and encourage users to install from there in most cases.

                                                pip install deepspeed

                                              After installation, you can validate your install and see which extensions/ops your machine is compatible with via the DeepSpeed environment report.

                                                ds_report

                                              If you would like to pre-install any of the DeepSpeed extensions/ops (instead of JIT compiling) or install pre-compiled ops via PyPI please see our advanced installation instructions.

                                                Windows

                                              Many DeepSpeed features are supported on Windows for both training and inference. You can read more about this in the original blog post here. Among features that are currently not supported are async io (AIO) and GDS (which does not support Windows).

                                               1. Install PyTorch, such as pytorch 2.3+cu121.
                                               2. Install Visual C++ build tools, such as VS2022 C++ x64/x86 build tools.
                                               3. Launch Cmd console with Administrator permissions for creating required symlink folders and ensure MSVC tools are added to your PATH or launch the Developer Command Prompt for Visual Studio 2022 with administrator permissions.
                                               4. Run build_win.bat to build wheel in dist folder.

                                                Further Reading

                                              All DeepSpeed documentation, tutorials, and blogs can be found on our website: deepspeed.ai

                                                                            Description                          
                                              Getting Started               First steps with DeepSpeed           
                                              DeepSpeed JSON Configuration  Configuring DeepSpeed                
                                              API Documentation             Generated DeepSpeed API documentation
                                              Tutorials                     Tutorials                            
                                              Blogs                         Blogs                                
                                              

                                                CI funding

                                              This being an open source project we rely on others to provide us resources for CI hardware. At this moment Modal is kindly supporting our GPU CI runs by funding the hardware for us. Modal is an AI infrastructure platform for inference, fine-tuning, batch jobs and more. Get started with $30/mo in free credits today at https://modal.com. We have been getting an amazing support from Modal's team and will surely recommend them to your business.

                                                Contributing

                                              DeepSpeed welcomes your contributions! Please see our contributing guide for more details on formatting, testing, etc.
                                              Thanks so much to all of our amazing contributors!

                                                Developer Certificate of Origin

                                              This project welcomes contributions and suggestions. Most contributions require you to agree to a Developer Certificate of Origin DCO stating that they agree to the terms published at https://developercertificate.org for that particular contribution.

                                              DCOs are per-commit, so each commit needs to be signed off. These can be signed in the commit by adding the -s flag. DCO enforcement can also be signed off in the PR itself by clicking on the DCO enforcement check.

                                                Code of Conduct

                                              This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

                                                Publications

                                               1. Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He. (2019) ZeRO: memory optimizations toward training trillion parameter models. arXiv:1910.02054 and In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC '20).

                                               2. Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. (2020) DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD '20, Tutorial).

                                               3. Minjia Zhang, Yuxiong He. (2020) Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping. arXiv:2010.13369 and NeurIPS 2020.

                                               4. Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, Yuxiong He. (2021) ZeRO-Offload: Democratizing Billion-Scale Model Training. arXiv:2101.06840 and USENIX ATC 2021. [paper] [slides] [blog]

                                               5. Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, Yuxiong He. (2021) 1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed. arXiv:2102.02888 and ICML 2021.

                                               6. Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He. (2021) ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning. arXiv:2104.07857 and SC 2021. [paper] [slides] [blog]

                                               7. Conglong Li, Ammar Ahmad Awan, Hanlin Tang, Samyam Rajbhandari, Yuxiong He. (2021) 1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed. arXiv:2104.06069 and HiPC 2022.

                                               8. Conglong Li, Minjia Zhang, Yuxiong He. (2021) The Stability-Efficiency Dilemma: Investigating Sequence Length Warmup for Training GPT Models. arXiv:2108.06084 and NeurIPS 2022.

                                               9. Yucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa, Yuxiong He. (2022) Maximizing Communication Efficiency for Large-scale Training via 0/1 Adam. arXiv:2202.06009.

                                              10. Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, Yuxiong He. (2022) DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale arXiv:2201.05596 and ICML 2022. [pdf] [slides] [blog]

                                              11. Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, Bryan Catanzaro. (2022) Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model arXiv:2201.11990.

                                              12. Xiaoxia Wu, Zhewei Yao, Minjia Zhang, Conglong Li, Yuxiong He. (2022) Extreme Compression for Pre-trained Transformers Made Simple and Efficient. arXiv:2206.01859 and NeurIPS 2022.

                                              13. Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, Yuxiong He. (2022) ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers. arXiv:2206.01861 and NeurIPS 2022 [slides] [blog]

                                              14. Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, Yuxiong He. (2022) DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale. arXiv:2207.00032 and SC 2022. [paper] [slides] [blog]

                                              15. Zhewei Yao, Xiaoxia Wu, Conglong Li, Connor Holmes, Minjia Zhang, Cheng Li, Yuxiong He. (2022) Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers. arXiv:2211.11586.

                                              16. Conglong Li, Zhewei Yao, Xiaoxia Wu, Minjia Zhang, Yuxiong He. (2022) DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and Routing. arXiv:2212.03597 ENLSP2023 Workshop at NeurIPS2023

                                              17. Xiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao, Yuxiong He. (2023) Understanding INT4 Quantization for Transformer Models: Latency Speedup, Composability, and Failure Cases. arXiv:2301.12017 and ICML2023.

                                              18. Syed Zawad, Cheng Li, Zhewei Yao, Elton Zheng, Yuxiong He, Feng Yan. (2023) DySR: Adaptive Super-Resolution via Algorithm and System Co-design. ICLR:2023.

                                              19. Sheng Shen, Zhewei Yao, Chunyuan Li, Trevor Darrell, Kurt Keutzer, Yuxiong He. (2023) Scaling Vision-Language Models with Sparse Mixture of Experts. arXiv:2303.07226 and Finding at EMNLP2023.

                                              20. Quentin Anthony, Ammar Ahmad Awan, Jeff Rasley, Yuxiong He, Aamir Shafi, Mustafa Abduljabbar, Hari Subramoni, Dhabaleswar Panda. (2023) MCR-DL: Mix-and-Match Communication Runtime for Deep Learning arXiv:2303.08374 and will appear at IPDPS 2023.

                                              21. Siddharth Singh, Olatunji Ruwase, Ammar Ahmad Awan, Samyam Rajbhandari, Yuxiong He, Abhinav Bhatele. (2023) A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training arXiv:2303.06318 and ICS 2023.

                                              22. Guanhua Wang, Heyang Qin, Sam Ade Jacobs, Xiaoxia Wu, Connor Holmes, Zhewei Yao, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, Yuxiong He. (2023) ZeRO++: Extremely Efficient Collective Communication for Giant Model Training arXiv:2306.10209 and ML for Sys Workshop at NeurIPS2023 [blog]

                                              23. Zhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn, Yuxiong He. (2023) ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation arXiv:2303.08302 and ENLSP2023 Workshop at NeurIPS2023 [slides]

                                              24. Pareesa Ameneh Golnari, Zhewei Yao, Yuxiong He. (2023) Selective Guidance: Are All the Denoising Steps of Guided Diffusion Important? arXiv:2305.09847

                                              25. Zhewei Yao, Reza Yazdani Aminabadi, Olatunji Ruwase, Samyam Rajbhandari, Xiaoxia Wu, Ammar Ahmad Awan, Jeff Rasley, Minjia Zhang, Conglong Li, Connor Holmes, Zhongzhu Zhou, Michael Wyatt, Molly Smith, Lev Kurilenko, Heyang Qin, Masahiro Tanaka, Shuai Che, Shuaiwen Leon Song, Yuxiong He. (2023) DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales arXiv:2308.01320.

                                              26. Xiaoxia Wu, Zhewei Yao, Yuxiong He. (2023) ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats arXiv:2307.09782 and ENLSP2023 Workshop at NeurIPS2023 [slides]

                                              27. Zhewei Yao, Xiaoxia Wu, Conglong Li, Minjia Zhang, Heyang Qin, Olatunji Ruwase, Ammar Ahmad Awan, Samyam Rajbhandari, Yuxiong He. (2023) DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention arXiv:2309.14327

                                              28. Shuaiwen Leon Song, Bonnie Kruft, Minjia Zhang, Conglong Li, Shiyang Chen, Chengming Zhang, Masahiro Tanaka, Xiaoxia Wu, Jeff Rasley, Ammar Ahmad Awan, Connor Holmes, Martin Cai, Adam Ghanem, Zhongzhu Zhou, Yuxiong He, et al. (2023) DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies arXiv:2310.04610 [blog]

                                              29. Zhewei Yao, Reza Yazdani Aminabadi, Stephen Youn, Xiaoxia Wu, Elton Zheng, Yuxiong He. (2023) ZeroQuant-HERO: Hardware-Enhanced Robust Optimized Post-Training Quantization Framework for W8A8 Transformers arXiv:2310.17723

                                              30. Xiaoxia Wu, Haojun Xia, Stephen Youn, Zhen Zheng, Shiyang Chen, Arash Bakhtiari, Michael Wyatt, Reza Yazdani Aminabadi, Yuxiong He, Olatunji Ruwase, Leon Song, Zhewei Yao (2023) ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks arXiv:2312.08583

                                              31. Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, Shuaiwen Leon Song. (2024) FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design arXiv:2401.14112

                                              32. Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Reza Yazdani Aminadabi, Shuaiwen Leon Song, Samyam Rajbhandari, Yuxiong He. (2024) System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

                                              33. Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka, Stas Bekman, Olatunji Ruwase, Minjia Zhang. (2024) Universal Checkpointing: Efficient and Flexible Checkpointing for Large Scale Distributed Training arXiv:2406.18820

                                              34. Stas Bekman, Samyam Rajbhandari, Michael Wyatt, Jeff Rasley, Tunji Ruwase, Zhewei Yao, Aurick Qiao, Yuxiong He. (2025) Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences arXiv:2506.13996

                                              35. Tingfeng Lan, Yusen Wu, Bin Ma, Zhaoyuan Su, Rui Yang, Tekin Bicer, Masahiro Tanaka, Olatunji Ruwase, Dong Li, Yue Cheng. (2025) ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates arXiv:2505.12242

                                              36. Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song, Yun Dai, Aman Gupta, Zhipeng Wang, Hejian Sang, Shao Tang, Gregory Dexter, Sirou Zhu, Siyu Zhu, Tejas Dharamsi, Vignesh Kothapalli, Zhoutong Fu, Yihan Cao, Pin-Lun Hsu, Fedor Borisyuk, Natesh S. Pillai, Luke Simon, Rahul Mazumder.(2025) Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems EMNLP 2025

                                              37. Xinyu Lian, Masahiro Tanaka, Olatunji Ruwase, Minjia Zhang. (2026) SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips arxiv, ASPLOS 2026

                                                Videos

                                               1. DeepSpeed KDD 2020 Tutorial
                                                   1. Overview
                                                   2. ZeRO + large model training
                                                   3. 17B T-NLG demo
                                                   4. Fastest BERT training + RScan tuning
                                                   5. DeepSpeed hands on deep dive: part 1, part 2, part 3
                                                   6. FAQ
                                               2. Microsoft Research Webinar
                                                    + Registration is free and all videos are available on-demand.
                                                    + ZeRO & Fastest BERT: Increasing the scale and speed of deep learning training in DeepSpeed.
                                               3. DeepSpeed on AzureML
                                               4. Large Model Training and Inference with DeepSpeed // Samyam Rajbhandari // LLMs in Prod Conference [slides]
                                               5. Community Tutorials
                                                    + DeepSpeed: All the tricks to scale to gigantic models (Mark Saroufim)
                                                    + Turing-NLG, DeepSpeed and the ZeRO optimizer (Yannic Kilcher)
                                                    + Ultimate Guide To Scaling ML Models (The AI Epiphany)

                                        About

                                        DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

                                          www.deepspeed.ai/

                                        Topics

                                          billion-parameterscompressiondata-parallelismdeep-learninggpuinferencemachine-learningmixture-of-expertsmodel-parallelismpipeline-parallelismpytorchtrillion-parameterszero

                                        Resources

                                          Readme
                                          Apache-2.0 license

                                        Code of conduct

                                          Code of conduct

                                        Contributing

                                          Contributing

                                        Security policy

                                          Security policy
                                          Activity
                                          Custom properties

                                        Stars

                                          42.9k stars

                                        Watchers

                                          358 watching

                                        Forks

                                          4.9k forks
                                          Report repository

                                        Releases

                                        Packages

                                        Used by

                                        Contributors

                                        Languages

    Footer

        © 2026 GitHub, Inc.

      Footer navigation

        * Terms
        * Privacy
        * Security
        * Status
        * Community
        * Docs
        * Contact
        * Manage cookies
        * Do not share my personal information
      You can’t perform that action at this time.
For now, Differences are performed on text, not graphically, only the latest screenshot is available.

Screenshot requires a Content Fetcher ( Sockpuppetbrowser, selenium, etc ) that supports screenshots.