PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS.

PipeDream: Fast andEfficient Pipeline…PipeDream: Fast and Efficient Pipeline Parallel DNN TrainingMixed Precision TrainingMixed Precision TrainingPipeDream: generalizedpipeline parallelism fo…PipeDream: generalized pipeline parallelism for DNN trainingPyTorch Distributed:Experiences on…PyTorch Distributed: Experiences on Accelerating Data Parallel TrainingHigh-performance,Distributed Training of…High-performance, Distributed Training of Large-scale Deep Learning Recommendation ModelsGSPMD: General andScalable Parallelizatio…GSPMD: General and Scalable Parallelization for ML Computation GraphsDynamic TensorRematerializationDynamic Tensor RematerializationPipeTransformer:Automated Elastic…PipeTransformer: Automated Elastic Pipelining for Distributed Training of TransformersDeepViT: Towards DeeperVision TransformerDeepViT: Towards Deeper Vision TransformerAlpa: Automating Inter-and Intra-Operator…Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningMiCS: Near-linearScaling for Training…MiCS: Near-linear Scaling for Training Gigantic Model on Public CloudReducing ActivationRecomputation in Large…Reducing Activation Recomputation in Large Transformer ModelsSlapo: A ScheduleLanguage for Progressiv…Slapo: A Schedule Language for Progressive Optimization of Large Deep Learning Model TrainingCambrian-1: A FullyOpen, Vision-Centric…Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsMovie Gen: A Cast ofMedia Foundation ModelsMovie Gen: A Cast of Media Foundation ModelsWukong: Towards aScaling Law for…Wukong: Towards a Scaling Law for Large-Scale RecommendationMAGI-1: AutoregressiveVideo Generation at…MAGI-1: Autoregressive Video Generation at ScaleRELIC: Interactive VideoWorld Model with…RELIC: Interactive Video World Model with Long-Horizon MemorySeedance 1.0: Exploringthe Boundaries of Video…Seedance 1.0: Exploring the Boundaries of Video Generation ModelsAutoregressiveAdversarial…Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationQwen-Image TechnicalReportQwen-Image Technical ReportRollPacker: MitigatingLong-Tail Rollouts for…RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-TrainingScaling and evaluatingsparse autoencodersScaling and evaluating sparse autoencodersLarge Scale DiffusionDistillation via…Large Scale Diffusion Distillation via Score-Regularized Continuous-Time ConsistencyPyTorch FSDP:Experiences on Scaling…PyTorch FSDP: Experiences on Scaling Fully Sharded Data ParallelEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.