paper-with-me

홈 › Papers

Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs

2022-06-09 · Jinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang, Hongsheng Li, Xiaogang Wang, Jifeng Dai

To build an artificial neural network like the biological intelligence system, recent works have unified numerous tasks into a generalist model, which can process various tasks with shared parameters and do not have any task-specific modules. While generalist models achieve promising results on various benchmarks, they have performance degradation on some tasks compared with task-specialized models. In this work, we find that interference among different tasks and modalities is the main factor to this phenomenon. To mitigate such interference, we introduce the Conditional Mixture-of-Experts (Conditional MoEs) to generalist models. Routing strategies under different levels of conditions are proposed to take both the training/inference cost and generalization ability into account. By incorporating the proposed Conditional MoEs, the recently proposed generalist model Uni-Perceiver can effectively mitigate the interference across tasks and modalities, and achieves state-of-the-art results on a series of downstream tasks via prompt tuning on 1% of downstream data. Moreover, the introduction of Conditional MoEs still holds the generalization ability of generalist models to conduct zero-shot inference on new tasks, e.g., video-text retrieval and video caption. Code and pre-trained generalist models shall be released.

📄 PDF Abstract BibTeX arXiv:2206.04674

Code (1)

fundamentalvision/Uni-Perceiver 공식 구현 pytorch

Tasks

Image CaptioningImage ClassificationImage-to-Text RetrievalMixture-of-ExpertsRetrievalText RetrievalVideo CaptioningVideo ClassificationVideo-Text Retrieval

Similar Papers 제목 키워드 기반

Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks

2022-11-17 · CVPR 2023 1 · Hao Li, Jinguo Zhu, Xiaohu Jiang, Xizhou Zhu 외

Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to eliminating this inconsistency is to use gene…

DecoderLanguage ModellingMulti-Task Learning

Efficient Language Modeling with Sparse all-MLP

2022-03-14 · Ping Yu, Mikel Artetxe, Myle Ott, Sam Shleifer 외

All-MLP architectures have attracted increasing interest as an alternative to attention-based models. In NLP, recent work like gMLP shows that all-MLPs can match Transformers in language modeling, but still lag behind in…

AllCommon Sense ReasoningIn-Context LearningLanguage Modeling+5

MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts

2024-10-18 · Rachel S. Y. Teo, Tan M. Nguyen

Sparse Mixture of Experts (SMoE) has become the key to unlocking unparalleled scalability in deep learning. SMoE has the potential to exponentially increase parameter count while maintaining the efficiency of the model b…

Language ModelingLanguage ModellingMixture-of-ExpertsObject Recognition

Mobile V-MoEs: Scaling Down Vision Transformers via Sparse Mixture-of-Experts

2023-09-08 · Erik Daxberger, Floris Weers, BoWen Zhang, Tom Gunter 외

Sparse Mixture-of-Experts models (MoEs) have recently gained popularity due to their ability to decouple model size from inference efficiency by only activating a small subset of the model parameters for any given input …

Mixture-of-Experts

Routers in Vision Mixture of Experts: An Empirical Study

2024-01-29 · Tianlin Liu, Mathieu Blondel, Carlos Riquelme, Joan Puigcerver

Mixture-of-Experts (MoE) models are a promising way to scale up model capacity without significantly increasing computational cost. A key component of MoEs is the router, which decides which subset of parameters (experts…

Language ModelingLanguage ModellingMixture-of-Experts