paper-with-me

홈 › Papers

FlexShard: Flexible Sharding for Industry-Scale Sequence Recommendation Models

2023-01-08 · Geet Sethi, Pallab Bhattacharya, Dhruv Choudhary, Carole-Jean Wu, Christos Kozyrakis

Sequence-based deep learning recommendation models (DLRMs) are an emerging class of DLRMs showing great improvements over their prior sum-pooling based counterparts at capturing users' long term interests. These improvements come at immense system cost however, with sequence-based DLRMs requiring substantial amounts of data to be dynamically materialized and communicated by each accelerator during a single iteration. To address this rapidly growing bottleneck, we present FlexShard, a new tiered sequence embedding table sharding algorithm which operates at a per-row granularity by exploiting the insight that not every row is equal. Through precise replication of embedding rows based on their underlying probability distribution, along with the introduction of a new sharding strategy adapted to the heterogeneous, skewed performance of real-world cluster network topologies, FlexShard is able to significantly reduce communication demand while using no additional memory compared to the prior state-of-the-art. When evaluated on production-scale sequence DLRMs, FlexShard was able to reduce overall global all-to-all communication traffic by over 85%, resulting in end-to-end training communication latency improvements of almost 6x over the prior state-of-the-art approach.

📄 PDF Abstract BibTeX arXiv:2301.02959

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlashCP: Load-Balanced Communication-Efficient Context Parallelism for LLM Training

2026-06-07 · Zheng Wang, Eric Liu, Linan Jiang, Zhongkai Yu 외 arxiv

Context parallelism (CP) is essential for training large-scale, long-context language models, as it partitions sequences to reduce memory overhead. However, existing CP methods suffer from workload imbalance, inefficient…

veScale-FSDP: Flexible and High-Performance FSDP at Scale

2026-02-25 · Zezhou Wang, Youjie Li, Zhiqi Lin, Jiacheng Yang 외 arxiv

Fully Sharded Data Parallel (FSDP), also known as Zero Redundancy Optimizer (ZeRO), is widely used for large-scale model training, because of its memory efficiency and minimal intrusion on model code. However, existing F…

Scale MLPerf-0.6 models on Google TPU-v3 Pods

2019-09-21 · Sameer Kumar, Victor Bitorff, Dehao Chen, Chiachen Chou 외

The recent submission of Google TPU-v3 Pods to the industry wide MLPerf v0.6 training benchmark demonstrates the scalability of a suite of industry relevant ML models. MLPerf defines a suite of models, datasets and rules…

Benchmarking

ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

2025-02-28 · Hao Ge, Junda Feng, Qi Huang, Fangcheng Fu 외

Scaling long-context ability is essential for Large Language Models (LLMs). To amortize the memory consumption across multiple devices in long-context training, inter-data partitioning (a.k.a. Data Parallelism) and intra…

Pre-train and Search: Efficient Embedding Table Sharding with Pre-trained Neural Cost Models

2023-05-03 · Daochen Zha, Louis Feng, Liang Luo, Bhargav Bhushanam 외

Sharding a large machine learning model across multiple devices to balance the costs is important in distributed training. This is challenging because partitioning is NP-hard, and estimating the costs accurately and effi…