paper-with-me

Papers

Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics

2025-03-30 · Jing Zhu, Mingxuan Ju, Yozen Liu, Danai Koutra, Neil Shah, Tong Zhao

Generative recommendation (GR) has become a powerful paradigm in recommendation systems that implicitly links modality and semantics to item representation, in contrast to previous methods that relied on non-semantic item identifiers in autoregressive models. However, previous research has predominantly treated modalities in isolation, typically assuming item content is unimodal (usually text). We argue that this is a significant limitation given the rich, multimodal nature of real-world data and the potential sensitivity of GR models to modality choices and usage. Our work aims to explore the critical problem of Multimodal Generative Recommendation (MGR), highlighting the importance of modality choices in GR nframeworks. We reveal that GR models are particularly sensitive to different modalities and examine the challenges in achieving effective GR when multiple modalities are available. By evaluating design strategies for effectively leveraging multiple modalities, we identify key challenges and introduce MGR-LF++, an enhanced late fusion framework that employs contrastive modality alignment and special tokens to denote different modalities, achieving a performance improvement of over 20% compared to single-modality alternatives.

📄 PDF Abstract BibTeX arXiv:2503.23333

Code (0)

등록된 구현이 없습니다.

Tasks

Recommendation Systems

Similar Papers 제목 키워드 기반

A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

2025-02-22 · Zihao Lin, Samyadeep Basu, Mohammad Beigi, Varun Manjunatha 외

The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for better control. While significant progre…

Survey

CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features

2024-10-10 · Po-han Li, Sandeep P. Chinchali, Ufuk Topcu

Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two…

Cross-Modal RetrievalGPUimage-classificationImage Classification+1

Shifting the Baseline: Single Modality Performance on Visual Navigation & QA

2018-11-01 · Jesse Thomason, Daniel Gordon, Yonatan Bisk

We demonstrate the surprising strength of unimodal baselines in multimodal domains, and make concrete recommendations for best practices in future research. Where existing work often compares against random or majority c…

Question AnsweringVisual Navigation

Shifting the Baseline: Single Modality Performance on Visual Navigation \& QA

2019-06-01 · NAACL 2019 6 · Jesse Thomason, Daniel Gordon, Yonatan Bisk

We demonstrate the surprising strength of unimodal baselines in multimodal domains, and make concrete recommendations for best practices in future research. Where existing work often compares against random or majority c…

Visual Navigation

Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs

2025-05-16 · Abhishek Dey, Aabha Bothera, Samhita Sarikonda, Rishav Aryan 외

In this paper, we study the challenges of detecting events on social media, where traditional unimodal systems struggle due to the rapid and multimodal nature of data dissemination. We employ a range of models, including…

Event Detection