paper-with-me

홈 › Papers

MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Pre-trained language models have demonstrated superior performance in various natural language processing tasks. However, these models usually contain hundreds of millions of parameters, which limits their practicality because of latency requirements in real-world applications. Existing methods train small compressed models via knowledge distillation. However, performance of these small models drops significantly compared with the pre-trained models due to their reduced model capacity. We propose MoEBERT, which uses a Mixture-of-Experts structure to increase model capacity and inference speed. We initialize MoEBERT by adapting the feed-forward neural networks in a pre-trained model into multiple experts. As such, representation power of the pre-trained model is largely retained. During inference, only one of the experts is activated, such that speed can be improved. We also propose a layer-wise distillation method to train MoEBERT. We validate the efficiency and efficacy of MoEBERT on natural language understanding and question answering tasks. Results show that the proposed method outperforms existing task-specific distillation algorithms. For example, our method outperforms previous approaches by over $2\%$ on the MNLI (mismatched) dataset. Our code will be publicly available.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationMixture-of-ExpertsNatural Language UnderstandingQuestion Answering

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

2022-04-15 · NAACL 2022 7 · Simiao Zuo, Qingru Zhang, Chen Liang, Pengcheng He 외

Pre-trained language models have demonstrated superior performance in various natural language processing tasks. However, these models usually contain hundreds of millions of parameters, which limits their practicality b…

Knowledge DistillationMixture-of-ExpertsNatural Language UnderstandingQuestion Answering

COMET: Learning Cardinality Constrained Mixture of Experts with Trees and Local Search

2023-06-05 · Shibal Ibrahim, Wenyu Chen, Hussein Hazimeh, Natalia Ponomareva 외

The sparse Mixture-of-Experts (Sparse-MoE) framework efficiently scales up model capacity in various domains, such as natural language processing and vision. Sparse-MoEs select a subset of the "experts" (thus, only a por…

Language ModelingLanguage ModellingMixture-of-ExpertsRecommendation Systems

DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models

2024-09-10 · Maryam Akhavan Aghdam, Hongpeng Jin, Yanzhao Wu

Transformer-based Mixture-of-Experts (MoE) models have been driving several recent technological advancements in Natural Language Processing (NLP). These MoE models adopt a router mechanism to determine which experts to …

Mixture-of-Experts

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations

2025-08-01 · Qiyao Xue, Yuchen Dou, Ryan Shi, Xiang Lorraine Li 외 arxiv

Hate speech detection on Chinese social networks presents distinct challenges, particularly due to the widespread use of cloaking techniques designed to evade conventional text-based detection systems. Although large lan…

Hate Speech Detection

Granger-causal Attentive Mixtures of Experts: Learning Important Features with Neural Networks

2018-02-06 · Patrick Schwab, Djordje Miladinovic, Walter Karlen

Knowledge of the importance of input features towards decisions made by machine-learning models is essential to increase our understanding of both the models and the underlying data. Here, we present a new approach to es…

Feature ImportanceMixture-of-Experts