paper-with-me

홈 › Papers

Bridging Domain Gaps between Pretrained Multimodal Models and Recommendations

2025-02-21 · Wenyu Zhang, Jie Luo, Xinming Zhang, Yuan Fang

With the explosive growth of multimodal content online, pre-trained visual-language models have shown great potential for multimodal recommendation. However, while these models achieve decent performance when applied in a frozen manner, surprisingly, due to significant domain gaps (e.g., feature distribution discrepancy and task objective misalignment) between pre-training and personalized recommendation, adopting a joint training approach instead leads to performance worse than baseline. Existing approaches either rely on simple feature extraction or require computationally expensive full model fine-tuning, struggling to balance effectiveness and efficiency. To tackle these challenges, we propose \textbf{P}arameter-efficient \textbf{T}uning for \textbf{M}ultimodal \textbf{Rec}ommendation (\textbf{PTMRec}), a novel framework that bridges the domain gap between pre-trained models and recommendation systems through a knowledge-guided dual-stage parameter-efficient training strategy. This framework not only eliminates the need for costly additional pre-training but also flexibly accommodates various parameter-efficient tuning methods.

📄 PDF Abstract BibTeX arXiv:2502.15542

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal RecommendationRecommendation Systems

Similar Papers 제목 키워드 기반

TextME: Bridging Unseen Modalities Through Text Descriptions

2026-02-03 · Soyeon Hong, Jinchan Kim, Jaegook You, Seungtaek Choi 외 arxiv

Expanding multimodal representations to novel modalities is constrained by reliance on large-scale paired datasets (e.g., text-image, text-audio, text-3D, text-molecule), which are costly and often infeasible in domains …

Cross-Modal Retrieval

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

2025-06-06 · Seung-jae Lee, Paul Hongsuck Seo

Audiovisual segmentation (AVS) aims to identify visual regions corresponding to sound sources, playing a vital role in video understanding, surveillance, and human-computer interaction. Traditional AVS methods depend on …

SegmentationVideo Understanding

Bridging Dynamics Gaps via Diffusion Schrödinger Bridge for Cross-Domain Reinforcement Learning

2026-02-27 · Hanping Zhang, Yuhong Guo arxiv

Cross-domain reinforcement learning (RL) aims to learn transferable policies under dynamics shifts between source and target domains. A key challenge lies in the lack of target-domain environment interaction and reward s…

Reinforcement Learning

Bridging Domain Gaps for Fine-Grained Moth Classification Through Expert-Informed Adaptation and Foundation Model Priors

2025-08-27 · Ross J Gardiner, Guillaume Mougeot, Sareh Rowlands, Benno I Simmons 외 arxiv

Labelling images of Lepidoptera (moths) from automated camera systems is vital for understanding insect declines. However, accurate species identification is challenging due to domain shifts between curated images and no…

Knowledge Distillation

Self-Supervised Enhancement of Forward-Looking Sonar Images: Bridging Cross-Modal Degradation Gaps through Feature Space Transformation and Multi-Frame Fusion

2025-04-15 · Zhisheng Zhang, Peng Zhang, Fengxiang Wang, Liangli Ma 외

Enhancing forward-looking sonar images is critical for accurate underwater target detection. Current deep learning methods mainly rely on supervised training with simulated data, but the difficulty in obtaining high-qual…