paper-with-me

홈 › Papers

Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning

2024-12-12 · Meng Shen, Yake Wei, Jianxiong Yin, Deepu Rajan, Di Hu, Simon See

Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on sufficient labeled data to train a well-calibrated model that can assess the uncertainty and diversity of unlabeled data. However, when assembling a dataset, labeled data are often scarce initially, leading to a cold-start problem. Additionally, most AL methods seldom address multimodal data, highlighting a research gap in this field. Our research addresses these issues by developing a two-stage method for Multi-Modal Cold-Start Active Learning (MMCSAL). Firstly, we observe the modality gap, a significant distance between the centroids of representations from different modalities, when only using cross-modal pairing information as self-supervision signals. This modality gap affects data selection process, as we calculate both uni-modal and cross-modal distances. To address this, we introduce uni-modal prototypes to bridge the modality gap. Secondly, conventional AL methods often falter in multimodal scenarios where alignment between modalities is overlooked. Therefore, we propose enhancing cross-modal alignment through regularization, thereby improving the quality of selected multimodal data pairs in AL. Finally, our experiments demonstrate MMCSAL's efficacy in selecting multimodal data pairs across three multimodal datasets.

📄 PDF Abstract BibTeX arXiv:2412.09126

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learningcross-modal alignment

Similar Papers 제목 키워드 기반

PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest

2026-03-03 · Josh Beal, Eric Kim, Jinfeng Rao, Rex Wu 외 arxiv

While multi-modal Visual Language Models (VLMs) have demonstrated significant success across various domains, the integration of VLMs into recommendation and retrieval systems remains a challenge, due to issues like trai…

Representation Learning

A Multimodal Single-Branch Embedding Network for Recommendation in Cold-Start and Missing Modality Scenarios

2024-09-26 · Christian Ganhör, Marta Moscati, Anna Hausberger, Shah Nawaz 외

Most recommender systems adopt collaborative filtering (CF) and provide recommendations based on past collective interactions. Therefore, the performance of CF algorithms degrades when few or no interactions are availabl…

Collaborative FilteringMultimodal RecommendationRecommendation Systems

Multimodal Pre-training Framework for Sequential Recommendation via Contrastive Learning

2023-03-21 · Lingzi Zhang, Xin Zhou, Zhiwei Zeng, Zhiqi Shen

Current multimodal sequential recommendation models are often unable to effectively explore and capture correlations among behavior sequences of users and items across different modalities, either neglecting correlations…

Contrastive LearningRecommendation SystemsRepresentation LearningSequential Recommendation

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

2025-02-24 · Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu 외

Multimodal alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairw…

Image GenerationRetrievalText to Image GenerationText-to-Image Generation

LRMM: Learning to Recommend with Missing Modalities

2018-08-21 · EMNLP 2018 10 · Cheng Wang, Mathias Niepert, Hui Li

Multimodal learning has shown promising performance in content-based recommendation due to the auxiliary user and item information of multiple modalities such as text and images. However, the problem of incomplete and mi…

Recommendation Systems