paper-with-me

홈 › Papers

Automatic Discovery of Composite SPMD Partitioning Strategies in PartIR

2022-10-07 · Sami Alabed, Dominik Grewe, Juliana Franco, Bart Chrzaszcz, Tom Natan, Tamara Norman, Norman A. Rink, Dimitrios Vytiniotis, Michael Schaarschmidt

Large neural network models are commonly trained through a combination of advanced parallelism strategies in a single program, multiple data (SPMD) paradigm. For example, training large transformer models requires combining data, model, and pipeline partitioning; and optimizer sharding techniques. However, identifying efficient combinations for many model architectures and accelerator systems requires significant manual analysis. In this work, we present an automatic partitioner that identifies these combinations through a goal-oriented search. Our key findings are that a Monte Carlo Tree Search-based partitioner leveraging partition-specific compiler analysis directly into the search and guided goals matches expert-level strategies for various models.

📄 PDF Abstract BibTeX arXiv:2210.06352

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GSPMD: General and Scalable Parallelization for ML Computation Graphs

2021-05-10 · Yuanzhong Xu, HyoukJoong Lee, Dehao Chen, Blake Hechtman 외

We present GSPMD, an automatic, compiler-based parallelization system for common machine learning computations. It allows users to write programs in the same way as for a single device, then give hints through a few anno…

Playing the Game of 2048

PartIR: Composing SPMD Partitioning Strategies for Machine Learning

2024-01-20 · Sami Alabed, Daniel Belov, Bart Chrzaszcz, Juliana Franco 외

Training of modern large neural networks (NN) requires a combination of parallelization strategies encompassing data, model, or optimizer sharding. When strategies increase in complexity, it becomes necessary for partiti…

Automap: Towards Ergonomic Automated Parallelism for ML Models

2021-12-06 · Michael Schaarschmidt, Dominik Grewe, Dimitrios Vytiniotis, Adam Paszke 외

The rapid rise in demand for training large neural network architectures has brought into focus the need for partitioning strategies, for example by using data, model, or pipeline parallelism. Implementing these methods …

Memory-efficient array redistribution through portable collective communication

2021-12-02 · Norman A. Rink, Adam Paszke, Dimitrios Vytiniotis, Georg Stefan Schmid

Modern large-scale deep learning workloads highlight the need for parallel execution across many devices in order to fit model data into hardware accelerator memories. In these settings, array redistribution may be requi…

Scaling Deep Learning Training with MPMD Pipeline Parallelism

2024-12-18 · Anxhelo Xhebraj, Sean Lee, Hanfeng Chen, Vinod Grover

We present JaxPP, a system for efficiently scaling the training of large deep learning models with flexible pipeline parallelism. We introduce a seamless programming model that allows implementing user-defined pipeline s…

Deep Learning