paper-with-me

홈 › Papers

PartIR: Composing SPMD Partitioning Strategies for Machine Learning

2024-01-20 · Sami Alabed, Daniel Belov, Bart Chrzaszcz, Juliana Franco, Dominik Grewe, Dougal Maclaurin, James Molloy, Tom Natan, Tamara Norman, Xiaoyue Pan, Adam Paszke, Norman A. Rink, Michael Schaarschmidt, Timur Sitdikov, Agnieszka Swietlik, Dimitrios Vytiniotis, Joel Wee

Training of modern large neural networks (NN) requires a combination of parallelization strategies encompassing data, model, or optimizer sharding. When strategies increase in complexity, it becomes necessary for partitioning tools to be 1) expressive, allowing the composition of simpler strategies, and 2) predictable to estimate performance analytically. We present PartIR, our design for a NN partitioning system. PartIR is focused on an incremental approach to rewriting and is hardware-and-runtime agnostic. We present a simple but powerful API for composing sharding strategies and a simulator to validate them. The process is driven by high-level programmer-issued partitioning tactics, which can be both manual and automatic. Importantly, the tactics are specified separately from the model code, making them easy to change. We evaluate PartIR on several different models to demonstrate its predictability, expressibility, and ability to reach peak performance..

📄 PDF Abstract BibTeX arXiv:2401.11202

Code (1)

openxla/shardy 공식 구현 jax

Similar Papers 제목 키워드 기반

Automatic Discovery of Composite SPMD Partitioning Strategies in PartIR

2022-10-07 · Sami Alabed, Dominik Grewe, Juliana Franco, Bart Chrzaszcz 외

Large neural network models are commonly trained through a combination of advanced parallelism strategies in a single program, multiple data (SPMD) paradigm. For example, training large transformer models requires combin…

GSPMD: General and Scalable Parallelization for ML Computation Graphs

2021-05-10 · Yuanzhong Xu, HyoukJoong Lee, Dehao Chen, Blake Hechtman 외

We present GSPMD, an automatic, compiler-based parallelization system for common machine learning computations. It allows users to write programs in the same way as for a single device, then give hints through a few anno…

Playing the Game of 2048

Automap: Towards Ergonomic Automated Parallelism for ML Models

2021-12-06 · Michael Schaarschmidt, Dominik Grewe, Dimitrios Vytiniotis, Adam Paszke 외

The rapid rise in demand for training large neural network architectures has brought into focus the need for partitioning strategies, for example by using data, model, or pipeline parallelism. Implementing these methods …

Memory-efficient array redistribution through portable collective communication

2021-12-02 · Norman A. Rink, Adam Paszke, Dimitrios Vytiniotis, Georg Stefan Schmid

Modern large-scale deep learning workloads highlight the need for parallel execution across many devices in order to fit model data into hardware accelerator memories. In these settings, array redistribution may be requi…

Structure-Preserving Margin Distribution Learning for High-Order Tensor Data with Low-Rank Decomposition

2025-09-18 · Yang Xu, Junpeng Li, Changchun Hua, Yana Yang arxiv

The Large Margin Distribution Machine (LMDM) is a recent advancement in classifier design that optimizes not just the minimum margin (as in SVM) but the entire margin distribution, thereby improving generalization. Howev…