paper-with-me

홈 › Papers

Facet-Aware Multi-Head Mixture-of-Experts Model for Sequential Recommendation

2024-11-03 · Mingrui Liu, Sixiao Zhang, Cheng Long

Sequential recommendation (SR) systems excel at capturing users' dynamic preferences by leveraging their interaction histories. Most existing SR systems assign a single embedding vector to each item to represent its features, and various types of models are adopted to combine these item embeddings into a sequence representation vector to capture the user intent. However, we argue that this representation alone is insufficient to capture an item's multi-faceted nature (e.g., movie genres, starring actors). Besides, users often exhibit complex and varied preferences within these facets (e.g., liking both action and musical films in the facet of genre), which are challenging to fully represent. To address the issues above, we propose a novel structure called Facet-Aware Multi-Head Mixture-of-Experts Model for Sequential Recommendation (FAME). We leverage sub-embeddings from each head in the last multi-head attention layer to predict the next item separately. This approach captures the potential multi-faceted nature of items without increasing model complexity. A gating mechanism integrates recommendations from each head and dynamically determines their importance. Furthermore, we introduce a Mixture-of-Experts (MoE) network in each attention head to disentangle various user preferences within each facet. Each expert within the MoE focuses on a specific preference. A learnable router network is adopted to compute the importance weight for each expert and aggregate them. We conduct extensive experiments on four public sequential recommendation datasets and the results demonstrate the effectiveness of our method over existing baseline models.

📄 PDF Abstract BibTeX arXiv:2411.01457

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-ExpertsSequential Recommendation

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
MoE 설명 없음
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

A Mixture-of-Experts Model for Learning Multi-Facet Entity Embeddings

2020-12-01 · COLING 2020 8 · Rana Alshaikh, Zied Bouraoui, Shelan Jeawak, Steven Schockaert

Various methods have already been proposed for learning entity embeddings from text descriptions. Such embeddings are commonly used for inferring properties of entities, for recommendation and entity-oriented search, and…

Entity EmbeddingsMixture-of-Experts

Efficient Deweather Mixture-of-Experts with Uncertainty-aware Feature-wise Linear Modulation

2023-12-27 · Rongyu Zhang, Yulin Luo, Jiaming Liu, Huanrui Yang 외

The Mixture-of-Experts (MoE) approach has demonstrated outstanding scalability in multi-task learning including low-level upstream tasks such as concurrent removal of multiple adverse weather effects. However, the conven…

Image RestorationMixture-of-ExpertsMulti-Task Learning

MH-MoE: Multi-Head Mixture-of-Experts

2024-11-25 · Shaohan Huang, Xun Wu, Shuming Ma, Furu Wei

Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces within different experts. In this paper,…

Mixture-of-Experts

DAMEX: Dataset-aware Mixture-of-Experts for visual understanding of mixture-of-datasets

2023-11-08 · NeurIPS 2023 11 · Yash Jain, Harkirat Behl, Zsolt Kira, Vibhav Vineet

Construction of a universal detector poses a crucial question: How can we most effectively train a model on a large mixture of datasets? The answer lies in learning dataset-specific features and ensembling their knowledg…

Mixture-of-Expertsobject-detectionObject Detection

Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification

2025-11-14 · Qinghao Gao, Jiahui Qu, Wenqian Dong arxiv

Multimodal remote sensing classification often suffers from missing modalities caused by sensor failures and environmental interference, leading to severe performance degradation. In this work, we rethink missing-modalit…