paper-with-me

홈 › Papers

Investigating Mixture of Experts in Dense Retrieval

2024-12-16 · Effrosyni Sokli, Pranav Kasela, Georgios Peikos, Gabriella Pasi

While Dense Retrieval Models (DRMs) have advanced Information Retrieval (IR), one limitation of these neural models is their narrow generalizability and robustness. To cope with this issue, one can leverage the Mixture-of-Experts (MoE) architecture. While previous IR studies have incorporated MoE architectures within the Transformer layers of DRMs, our work investigates an architecture that integrates a single MoE block (SB-MoE) after the output of the final Transformer layer. Our empirical evaluation investigates how SB-MoE compares, in terms of retrieval effectiveness, to standard fine-tuning. In detail, we fine-tune three DRMs (TinyBERT, BERT, and Contriever) across four benchmark collections with and without adding the MoE block. Moreover, since MoE showcases performance variations with respect to its parameters (i.e., the number of experts), we conduct additional experiments to investigate this aspect further. The findings show the effectiveness of SB-MoE especially for DRMs with a low number of parameters (i.e., TinyBERT), as it consistently outperforms the fine-tuned underlying model on all four benchmarks. For DRMs with a higher number of parameters (i.e., BERT and Contriever), SB-MoE requires larger numbers of training samples to yield better retrieval performance.

📄 PDF Abstract BibTeX arXiv:2412.11864

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMixture-of-ExpertsRetrieval

Methods 이 논문이 사용한 방법론

Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Weight Decay 설명 없음
WordPiece 설명 없음
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

DESIRE-ME: Domain-Enhanced Supervised Information REtrieval using Mixture-of-Experts

2024-03-20 · Pranav Kasela, Gabriella Pasi, Raffaele Perego, Nicola Tonellotto

Open-domain question answering requires retrieval systems able to cope with the diverse and varied nature of questions, providing accurate answers across a broad spectrum of query types and topics. To deal with such topi…

Information RetrievalMixture-of-ExpertsOpen-Domain Question AnsweringQuestion Answering+1

Mixture of A Million Experts

2024-07-04 · Xu Owen He

The feedforward (FFW) layers in standard transformer architectures incur a linear increase in computational costs and activation memory as the hidden layer width grows. Sparse mixture-of-experts (MoE) architectures have …

Computational EfficiencyLanguage ModelingLanguage ModellingMixture-of-Experts+1

Mixture of Parrots: Experts improve memorization more than reasoning

2024-10-24 · Samy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu 외

The Mixture-of-Experts (MoE) architecture enables a significant increase in the total number of model parameters with minimal computational overhead. However, it is not clear what performance tradeoffs, if any, exist bet…

MathMemorizationMixture-of-Experts

MODE: Mixture of Document Experts for RAG

2025-08-27 · Rahul Anand arxiv

Retrieval-Augmented Generation (RAG) often relies on large vector databases and cross-encoders tuned for large-scale corpora, which can be excessive for small, domain-specific collections. We present MODE (Mixture of Doc…

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models

2025-05-26 · Hao Kang, Zichun Yu, Chenyan Xiong

Recent large language models such as Gemini-1.5, DeepSeek-V3, and Llama-4 increasingly adopt Mixture-of-Experts (MoE) architectures, which offer strong efficiency-performance trade-offs by activating only a fraction of t…

Mixture-of-Experts