paper-with-me

홈 › Papers

RTP: Rethinking Tensor Parallelism with Memory Deduplication

2023-11-02 · Cheng Luo, Tianle Zhong, Geoffrey Fox

In the evolving landscape of neural network models, one prominent challenge stand out: the significant memory overheads associated with training expansive models. Addressing this challenge, this study delves deep into the Rotated Tensor Parallelism (RTP). RTP is an innovative approach that strategically focuses on memory deduplication in distributed training environments. It boasts of unique features like a customized communication primitive and the Flyweight Pattern initialization. Furthermore, RTP ensures a seamless overlap between partition computation and partition weight communication, optimizing the training process. Our empirical evaluations underscore RTP's efficiency, revealing that its memory consumption during distributed system training is remarkably close to the optimal - distributing the memory overhead of a single machine equitably among multiple machines. The experimental results demonstrate that RTP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of memory. Code of RTP is available at https://github.com/wdlctc/rtp.

📄 PDF Abstract BibTeX arXiv:2311.01635

Code (1)

wdlctc/rtp 공식 구현 pytorch

Similar Papers 제목 키워드 기반

TStore: Rethinking AI Model Hub with Tensor-Centric Compression

2026-04-18 · Tingfeng Lan, Zirui Wang, Yunjia Zheng, Zhaoyuan Su 외 arxiv

Modern AI models are growing rapidly in size and redundancy, leading to significant storage and distribution challenges in model hubs. We present TStore, a tensor-centric system for reducing storage overhead through fine…

Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference

2026-04-29 · Vasu Shyam, Anna Golubeva, Quentin Anthony arxiv

We present tensor and sequence parallelism (TSP), a parallel execution strategy that folds tensor parallelism and sequence parallelism onto a single device axis. In conventional multi-dimensional parallelism layouts, ten…

SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training

2024-08-19 · Kun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoğlu 외

The growth rate of the GPU memory capacity has not been able to keep up with that of the size of large language models (LLMs), hindering the model training process. In particular, activations -- the intermediate tensors …

GPULanguage ModelingLanguage ModellingLarge Language Model

Tesseract: Parallelize the Tensor Parallelism Efficiently

2021-05-30 · Boxiang Wang, Qifan Xu, Zhengda Bian, Yang You

Together with the improvements in state-of-the-art accuracies of various tasks, deep learning models are getting significantly larger. However, it is extremely difficult to implement these large models because limited GP…

GPULanguage Modelling

Scaling Neural Network Verification with Tensor Parallelism and Fully Sharded Data Parallelism

2026-06-08 · Sergei Vorobyov, Eugene Ilyushin arxiv

Formal neural network verification -- proving that a network satisfies safety properties for *all* inputs in a specified domain -- is bounded in practice by GPU memory: standard implementations of bound-propagation algor…