paper-with-me

Papers

Distributed Equivalent Substitution Training for Large-Scale Recommender Systems

2019-09-10 · Haidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai, Haiyang Wu, Rui Lan, Fan Li, Han Zhang, Yuekui Yang, Zhenyu Guo, Di Wang

We present Distributed Equivalent Substitution (DES) training, a novel distributed training framework for large-scale recommender systems with dynamic sparse features. DES introduces fully synchronous training to large-scale recommendation system for the first time by reducing communication, thus making the training of commercial recommender systems converge faster and reach better CTR. DES requires much less communication by substituting the weights-rich operators with the computationally equivalent sub-operators and aggregating partial results instead of transmitting the huge sparse weights directly through the network. Due to the use of synchronous training on large-scale Deep Learning Recommendation Models (DLRMs), DES achieves higher AUC(Area Under ROC). We successfully apply DES training on multiple popular DLRMs of industrial scenarios. Experiments show that our implementation outperforms the state-of-the-art PS-based training framework, achieving up to 68.7% communication savings and higher throughput compared to other PS-based recommender systems.

📄 PDF Abstract BibTeX arXiv:1909.04823

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion

2026-05-28 · Feng Han, Zhixiong Zhang, Zheming Liang, Yibin Wang 외 arxiv

Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multimodal fusion. Ideally, replacing a textua…

Multimodal ReasoningImage Captioning

An Equivalent Circuit Approach to Distributed Optimization

2023-05-24 · Aayushya Agarwal, Larry Pileggi

Distributed optimization is an essential paradigm to solve large-scale optimization problems in modern applications where big-data and high-dimensionality creates a computational bottleneck. Distributed optimization algo…

Distributed OptimizationNumerical Integration

TrainVerify: Equivalence-Based Verification for Distributed LLM Training

2025-06-19 · Yunchi Lu, Youshan Miao, Cheng Tan, Peng Huang 외

Training large language models (LLMs) at scale requires parallel execution across thousands of devices, incurring enormous computational costs. Yet, these costly distributed trainings are rarely verified, leaving them pr…

GPU

veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD

2025-09-05 · Youjie Li, Cheng Wan, Zhiqi Lin, Hongyu Zhu 외 arxiv

Large Language Models (LLMs) have scaled rapidly in size and complexity, requiring increasingly intricate parallelism for distributed training, such as 3D parallelism. This sophistication motivates a shift toward simpler…

Turk Bootstrap Word Sense Inventory 2.0: A Large-Scale Resource for Lexical Substitution

2012-05-01 · LREC 2012 5 · Chris Biemann

This paper presents the Turk Bootstrap Word Sense Inventory (TWSI) 2.0. This lexical resource, created by a crowdsourcing process using Amazon Mechanical Turk (http://www.mturk.com), encompasses a sense inventory for lex…

Machine TranslationQuestion AnsweringWord Sense Disambiguation