paper-with-me

홈 › Papers

DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72

2026-04-02 · Wanqian Li, Jintao Peng, Zongfei Jing, Tianyu Zhang, Ze Long, Xianjie Qiao, Xiaoming Chen, Dongxu Yang, Kefeng Duan, June Yang arxiv

Large language model (LLM) inference increasingly depends on multi-GPU execution, yet existing inference parallelization strategies require layer-wise inter-rank synchronization, making end-to-end performance sensitive to workload imbalance. We present DWDP (Distributed Weight Data Parallelism), an inference parallelization strategy that preserves data-parallel execution while offloading MoE weights across peer GPUs and fetching missing experts on demand. By removing collective inter-rank synchronization, DWDP allows each GPU to progress independently. We further address the practical overheads of this design with two optimizations for split-weight management and asynchronous remote-weight prefetch. Implemented in TensorRT-LLM and evaluated with DeepSeek-R1 on GB200 NVL72, DWDP improves end-to-end output TPS/GPU by 8.8% at comparable TPS/user in the 20-100 TPS/user serving range under 8K input sequence length and 1K output sequence length.

📄 PDF Abstract BibTeX arXiv:2604.01621

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism

2026-01-30 · Thalaiyasingam Ajanthan, Sameera Ramasinghe, Gil Avraham, Hadi Mohaghegh Dolatabadi 외 arxiv

Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limit…

OneFlow: Redesign the Distributed Deep Learning Framework from Scratch

2021-10-28 · Jinhui Yuan, Xinqi Li, Cheng Cheng, Juncheng Liu 외

Deep learning frameworks such as TensorFlow and PyTorch provide a productive interface for expressing and training a deep neural network (DNN) model on a single device or using data parallelism. Still, they may not be fl…

Deep Learning

A Bi-layered Parallel Training Architecture for Large-scale Convolutional Neural Networks

2018-10-17 · Jianguo Chen, Kenli Li, Kashif Bilal, Xu Zhou 외

Benefitting from large-scale training datasets and the complex training network, Convolutional Neural Networks (CNNs) are widely applied in various fields with high accuracy. However, the training process of CNNs is very…

Distributed ComputingScheduling

TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training

2025-11-12 · Houming Wu, Ling Chen arxiv

Training large language models (LLMs) is fundamentally constrained by limited device memory and costly inter-device communication. Although pipeline parallelism alleviates memory pressure by partitioning models across de…

RTP: Rethinking Tensor Parallelism with Memory Deduplication

2023-11-02 · Cheng Luo, Tianle Zhong, Geoffrey Fox

In the evolving landscape of neural network models, one prominent challenge stand out: the significant memory overheads associated with training expansive models. Addressing this challenge, this study delves deep into th…