paper-with-me

홈 › Papers

LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression

2020-04-08 · COLING 2020 8 · Yihuan Mao, Yujing Wang, Chufan Wu, Chen Zhang, Yang Wang, Yaming Yang, Quanlu Zhang, Yunhai Tong, Jing Bai

BERT is a cutting-edge language representation model pre-trained by a large corpus, which achieves superior performances on various natural language understanding tasks. However, a major blocking issue of applying BERT to online services is that it is memory-intensive and leads to unsatisfactory latency of user requests, raising the necessity of model compression. Existing solutions leverage the knowledge distillation framework to learn a smaller model that imitates the behaviors of BERT. However, the training procedure of knowledge distillation is expensive itself as it requires sufficient training data to imitate the teacher model. In this paper, we address this issue by proposing a hybrid solution named LadaBERT (Lightweight adaptation of BERT through hybrid model compression), which combines the advantages of different model compression methods, including weight pruning, matrix factorization and knowledge distillation. LadaBERT achieves state-of-the-art accuracy on various public datasets while the training overheads can be reduced by an order of magnitude.

📄 PDF Abstract BibTeX arXiv:2004.04124

Code (0)

등록된 구현이 없습니다.

Tasks

BlockingKnowledge DistillationModel CompressionNatural Language Understanding

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…
Residual Connection 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Weight Decay 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Semi-Siamese Bi-encoder Neural Ranking Model Using Lightweight Fine-Tuning

2021-10-28 · Euna Jung, Jaekeol Choi, Wonjong Rhee

A BERT-based Neural Ranking Model (NRM) can be either a crossencoder or a bi-encoder. Between the two, bi-encoder is highly efficient because all the documents can be pre-processed before the actual query time. In this w…

Language ModelingLanguage Modelling

Governance-Aware Hybrid Fine-Tuning for Multilingual Large Language Models

2025-12-19 · Haomin Qi, Chengbo Huang, Zihan Dai, Yunkai Gao arxiv

We present a governance-aware hybrid fine-tuning framework for multilingual, low-resource adaptation of large language models. The core algorithm combines gradient-aligned low-rank updates with structured orthogonal tran…

Language Identification

Lightweight Adaptation for LLM-based Technical Service Agent: Latent Logic Augmentation and Robust Noise Reduction

2026-03-18 · Yi Yu, Junzhuo Ma, Chenghuang Shen, Xingyan Liu 외 arxiv

Adapting Large Language Models in complex technical service domains is constrained by the absence of explicit cognitive chains in human demonstrations and the inherent ambiguity arising from the diversity of valid respon…

Computational EfficiencyReinforcement LearningTrajectory Modeling

HybridCVLNet: A Hybrid CSI Feedback System and its Domain Adaptation

2023-03-30 · Haozhen Li, Xinyu Gu, Boyuan Zhang, Dongliang Li 외

Deep Learning (DL)-based channel state information (CSI) feedback is a promising technique for the transmitter to accurately acquire the CSI of massive multiple-input multiple-output (MIMO) systems. As critical concerns …

Domain AdaptationTransfer Learning

Automatic Scoring of Students' Science Writing Using Hybrid Neural Network

2023-12-02 · Ehsan Latif, Xiaoming Zhai

This study explores the efficacy of a multi-perspective hybrid neural network (HNN) for scoring student responses in science education with an analytic rubric. We compared the accuracy of the HNN model with four ML appro…

regression