paper-with-me

Papers

Mining Stable Preferences: Adaptive Modality Decorrelation for Multimedia Recommendation

2023-06-25 · Jinghao Zhang, Qiang Liu, Shu Wu, Liang Wang

Multimedia content is of predominance in the modern Web era. In real scenarios, multiple modalities reveal different aspects of item attributes and usually possess different importance to user purchase decisions. However, it is difficult for models to figure out users' true preference towards different modalities since there exists strong statistical correlation between modalities. Even worse, the strong statistical correlation might mislead models to learn the spurious preference towards inconsequential modalities. As a result, when data (modal features) distribution shifts, the learned spurious preference might not guarantee to be as effective on the inference set as on the training set. We propose a novel MOdality DEcorrelating STable learning framework, MODEST for brevity, to learn users' stable preference. Inspired by sample re-weighting techniques, the proposed method aims to estimate a weight for each item, such that the features from different modalities in the weighted distribution are decorrelated. We adopt Hilbert Schmidt Independence Criterion (HSIC) as independence testing measure which is a kernel-based method capable of evaluating the correlation degree between two multi-dimensional and non-linear variables. Our method could be served as a play-and-plug module for existing multimedia recommendation backbones. Extensive experiments on four public datasets and four state-of-the-art multimedia recommendation backbones unequivocally show that our proposed method can improve the performances by a large margin.

📄 PDF Abstract BibTeX arXiv:2306.14179

Code (0)

등록된 구현이 없습니다.

Tasks

Multimedia recommendation

Similar Papers 제목 키워드 기반

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

2026-04-22 · Yuanhao Chen, Peter Chin arxiv

Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions. Current approaches…

Rule Learning for Knowledge Graph Reasoning under Agnostic Distribution Shift

2025-07-07 · Shixuan Liu, Yue He, Yunfei Wang, Hao Zou 외 arxiv

Logical rule learning, a prominent category of knowledge graph (KG) reasoning methods, constitutes a critical research area aimed at learning explicit rules from observed facts to infer missing knowledge. However, like a…

Dual-Stream Cross-Modal Representation Learning via Residual Semantic Decorrelation

2025-12-08 · Xuecheng Li, Weikuan Jia, Alisher Kurbonaliev, Qurbonaliev Alisher 외 arxiv

Cross-modal learning has become a fundamental paradigm for integrating heterogeneous information sources such as images, text, and structured attributes. However, multimodal representations often suffer from modality dom…

Representation Learning

Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation

2025-11-24 · Yuyang Wanyan, Xiaoshan Yang, Weiming Dong, Changsheng Xu arxiv

In this paper, we study the challenging task of Few-Shot Video Domain Adaptation (FSVDA). The multimodal nature of videos introduces unique challenges, necessitating the simultaneous consideration of both domain alignmen…

Domain Adaptation

Adaptive and Scalable Compression of Multispectral Images using VVC

2023-01-10 · Philipp Seltsam, Priyanka Das, Mathias Wien

The VVC codec is applied to the task of multispectral image (MSI) compression using adaptive and scalable coding structures. In a 'plain' VVC approach, concepts from picture-to-picture temporal prediction are employed fo…