paper-with-me

홈 › Papers

Towards Uniformity and Alignment for Multimodal Representation Learning

2026-02-10 · Wenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu, Shujian Yu, Jan-Jakob Sonke, Efstratios Gavves arxiv

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify two conflicts in the multimodal regime, both exacerbated as the number of modalities increases: (i) an alignment-uniformity conflict, whereby the repulsion of uniformity undermines pairwise alignment, and (ii) an intra-alignment conflict, where aligning multiple modalities induces competing alignment directions. To address these issues, we propose a principled decoupling of alignment and uniformity for multimodal representations, providing a conflict-free recipe for multimodal learning that simultaneously supports discriminative and generative use cases without task-specific modules. We then provide a theoretical guarantee that our method acts as an efficient proxy for a global Hölder divergence over multiple modality distributions, and thus reduces the distribution gap among modalities. Extensive experiments on retrieval and UnCLIP-style generation demonstrate consistent gains.

📄 PDF Abstract BibTeX arXiv:2602.09507

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

2025-02-24 · Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu 외

Multimodal alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairw…

Image GenerationRetrievalText to Image GenerationText-to-Image Generation

Geodesic Multi-Modal Mixup for Robust Fine-Tuning

2022-03-08 · NeurIPS 2023 11 · Changdae Oh, Junhyuk So, Hoyoon Byun, Yongtaek Lim 외

Pre-trained multi-modal models, such as CLIP, provide transferable embeddings and show promising results in diverse applications. However, the analysis of learned multi-modal embeddings is relatively unexplored, and the …

Image Captioningzero-shot-classificationZero-Shot Learning

RAU: Towards Regularized Alignment and Uniformity for Representation Learning in Recommendation

2025-03-24 · Xi Wu, Dan Zhang, Chao Zhou, Liangwei Yang 외

Recommender systems (RecSys) have become essential in modern society, driving user engagement and satisfaction across diverse online platforms. Most RecSys focuses on designing a powerful encoder to embed users and items…

Recommendation SystemsRepresentation Learning

Generalizable Person Re-identification via Balancing Alignment and Uniformity

2024-11-18 · Yoonki Cho, Jaeyoon Kim, Woo Jae Kim, Junsik Jung 외

Domain generalizable person re-identification (DG re-ID) aims to learn discriminative representations that are robust to distributional shifts. While data augmentation is a straightforward solution to improve generalizat…

Data AugmentationGeneralizable Person Re-identificationPerson Re-Identification

Explicit Representation Alignment for Multimodal Sentiment Analysis

2026-06-08 · Baode Wang, Ziming Wang, Huacan Wang, Ronghao Chen 외 arxiv

Multimodal affective analysis aims to understand human sentiment and emotion by jointly modeling heterogeneous modalities such as text and images. However, multimodal models often fail to consistently outperform strong t…

Multimodal Sentiment Analysis