paper-with-me

Papers

Demystifying Pipeline Parallelism: First Theory for PipeDream

2026-06-02 · Ivan Ilin, Peter Richtárik arxiv

Training modern machine learning models increasingly requires computation to be distributed across many accelerators. Data parallelism remains the default choice and is often paired with tensor-parallel sharding, but model parallelism becomes unavoidable once parameters, activations, or optimizer states no longer fit on a single device. This paper studies pipeline model parallelism through the lens of PipeDream (PD) (Harlap et al., 2018). Our first contribution is theoretical: we introduce Randomized PipeDream (RPD), a stale block-SGD abstraction that yields, to our knowledge, the first clean nonconvex convergence guarantee for a PD-style method. Our second contribution is a scaling diagnosis: we prove that the delay induced by steady-state PD grows as $S^2 - S/2 + O(1)$ for $S$ stages, so the stale-read contribution in the convergence theorem scales as $Θ(γ^2 S^4)$, equivalently as $Θ(S^4/K)$ in the tuned-rate form. Our third contribution is a comparison with LocalSGD, whose periodic model averaging trades weight staleness for synchronization bubbles. In our reported simulated-time experiments, the better-performing method depends on the objective: PD performs better on the quadratic objective and on a small language-modeling training-loss task, while for logistic regression LocalSGD becomes superior as the number of stages increases.

📄 PDF Abstract BibTeX arXiv:2606.03498

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Memory-Efficient Pipeline-Parallel DNN Training

2020-06-16 · Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen 외

Many state-of-the-art ML results have been obtained by scaling up the number of parameters in existing models. However, parameters and activations for such large models often do not fit in the memory of a single accelera…

Group-based Interleaved Pipeline Parallelism for Large-scale DNN Training

2021-09-29 · ICLR 2022 4 · Pengcheng Yang, XiaoMing Zhang, Wenpeng Zhang, Ming Yang 외

The recent trend of using large-scale deep neural networks (DNN) to boost performance has propelled the development of the parallel pipelining technique for efficient DNN training, which has resulted in the development o…

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining

2026-06-29 · Philip Zmushko, Egor Petrov, Nursultan Abdullaev, Mikhail Khrushchev 외 arxiv

Modern large-scale LLM pretraining benefits from utilizing Pipeline Parallelism; however, synchronous implementations leave GPUs idle during pipeline bubbles, wasting computational resources. Asynchronous Pipeline Parall…

GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism

2024-06-24 · Byungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim 외

Deep neural networks (DNNs) continue to grow rapidly in size, making them infeasible to train on a single device. Pipeline parallelism is commonly used in existing DNN systems to support large-scale DNN training by parti…

GPU

PipeOptim: Ensuring Effective 1F1B Schedule with Optimizer-Dependent Weight Prediction

2023-12-01 · Lei Guan, Dongsheng Li, Jiye Liang, Wenjian Wang 외

Asynchronous pipeline model parallelism with a "1F1B" (one forward, one backward) schedule generates little bubble overhead and always provides quite a high throughput. However, the "1F1B" schedule inevitably leads to we…

image-classificationImage ClassificationMachine TranslationPrediction+1