Speeding up Deep Learning Training by Sharing Weights and Then Unsharing
It has been widely observed that increasing deep learning model sizes often leads to significant performance improvements on a variety of natural language processing and computer vision tasks. In the meantime, however, computational costs and training time would dramatically increase when models get larger. In this paper, we propose a simple approach to speed up training for a particular kind of deep networks which contain repeated structures, such as the transformer module. In our method, we first train such a deep network with the weights shared across all the repeated layers. Once an unsharing condition is triggered, we stop weight sharing and continue training until convergence. Empirical results show that our method is able to reduce the training time of BERT by 50%. We also conduct a preliminary theoretic analysis which motivates our approach.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Speeding up Deep Model Training by Sharing Weights and Then Unsharing
We propose a simple and efficient approach for training the BERT model. Our approach exploits the special structure of BERT that contains a stack of repeated modules (i.e., transformer encoders). Our proposed approach fi…
FedMes: Speeding Up Federated Learning with Multiple Edge Servers
We consider federated learning with multiple wireless edge servers having their own local coverages. We focus on speeding up training in this increasingly practical setup. Our key idea is to utilize the devices located i…
Federated LearningAMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
Large language models (LLMs) have demonstrated remarkable capabilities in various kinds of tasks, while the billion or even trillion parameters bring storage and efficiency bottlenecks for inference. Quantization, partic…
GoSGD: Distributed Optimization for Deep Learning with Gossip Exchange
We address the issue of speeding up the training of convolutional neural networks by studying a distributed method adapted to stochastic gradient descent. Our parallel optimization setup uses several threads, each applyi…
Deep LearningDistributed OptimizationLearning Compact Neural Networks with Regularization
Proper regularization is critical for speeding up training, improving generalization performance, and learning compact models that are cost efficient. We propose and analyze regularized gradient descent algorithms for le…
Network Pruning