paper-with-me

Papers

CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation

2025-04-14 · Junchen Fu, Yongxin Ni, Joemon M. Jose, Ioannis Arapakis, Kaiwen Zheng, Youhua Li, Xuri Ge

Multimodal Foundation Models (MFMs) excel at representing diverse raw modalities (e.g., text, images, audio, videos, etc.). As recommender systems increasingly incorporate these modalities, leveraging MFMs to generate better representations has great potential. However, their application in sequential recommendation remains largely unexplored. This is primarily because mainstream adaptation methods, such as Fine-Tuning and even Parameter-Efficient Fine-Tuning (PEFT) techniques (e.g., Adapter and LoRA), incur high computational costs, especially when integrating multiple modality encoders, thus hindering research progress. As a result, it remains unclear whether we can efficiently and effectively adapt multiple (>2) MFMs for the sequential recommendation task. To address this, we propose a plug-and-play Cross-modal Side Adapter Network (CROSSAN). Leveraging the fully decoupled side adapter-based paradigm, CROSSAN achieves high efficiency while enabling cross-modal learning across diverse modalities. To optimize the final stage of multimodal fusion across diverse modalities, we adopt the Mixture of Modality Expert Fusion (MOMEF) mechanism. CROSSAN achieves superior performance on the public datasets for adapting four foundation models with raw modalities. Performance consistently improves as more MFMs are adapted. We will release our code and datasets to facilitate future research.

📄 PDF Abstract BibTeX arXiv:2504.10307

Code (1)

col-tasas/2025-oco-with-iqcs 공식 구현

Tasks

parameter-efficient fine-tuningRecommendation SystemsSequential Recommendation

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here
Adapter 설명 없음

Similar Papers 제목 키워드 기반

DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation

2025-09-29 · Xi Chen, Hongxun Yao, Zhaopan Xu, Kui Jiang arxiv

Source-free active domain adaptation (SFADA) enhances knowledge transfer from a source model to an unlabeled target domain using limited manual labels selected via active learning. While recent domain adaptation studies …

Source-Free Domain AdaptationActive Learning

Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models

2025-01-30 · Hao Dong, Moru Liu, Kaiyang Zhou, Eleni Chatzi 외

In real-world scenarios, achieving domain adaptation and generalization poses significant challenges, as models must adapt to or generalize across unknown target distributions. Extending these capabilities to unseen mult…

Action RecognitionDomain AdaptationDomain GeneralizationSemantic Segmentation+1

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

2026-04-01 · Wish Suharitdamrong, Tony Alex, Muhammad Awais, Sara Atito arxiv

Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challeng…

Visual Grounding

Dynamic Social Interaction Mechanics CrossAnt

2018-11-17 · Samuel Gomes, Carlos Martinho, João Dias

Nowadays, big effort is being put to study gamification and how game elements can be used to engage players. In this scope, we believe there is a growing need to explore the impact game mechanics have on the players' int…

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

2026-04-13 · Samuel Cahyawijaya, Peerat Limkonchotiwat, Tack Hwa Wong, Hitesh Laxmichand Patel 외 arxiv

While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and domains, there is still no dedicated framework for assessing human-centr…