paper-with-me

Papers

Decoupling Common and Unique Representations for Multimodal Self-supervised Learning

2023-09-11 · Yi Wang, Conrad M Albrecht, Nassim Ait Ali Braham, Chenying Liu, Zhitong Xiong, Xiao Xiang Zhu

The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-modal training and modality-unique representations. We propose Decoupling Common and Unique Representations (DeCUR), a simple yet effective method for multimodal self-supervised learning. By distinguishing inter- and intra-modal embeddings through multimodal redundancy reduction, DeCUR can integrate complementary information across different modalities. We evaluate DeCUR in three common multimodal scenarios (radar-optical, RGB-elevation, and RGB-depth), and demonstrate its consistent improvement regardless of architectures and for both multimodal and modality-missing settings. With thorough experiments and comprehensive analysis, we hope this work can provide valuable insights and raise more interest in researching the hidden relationships of multimodal representations.

📄 PDF Abstract BibTeX arXiv:2309.05300

Code (2)

zhu-xlab/decur 공식 구현 pytorch
zhu-xlab/dino-mm pytorch

Tasks

Scene ClassificationSelf-Supervised LearningSemantic Segmentation

Similar Papers 제목 키워드 기반

Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation

2025-03-07 · Xinkun Wang, Yifang Wang, Senwei Liang, Feilong Tang 외

This paper discusses how ophthalmologists often rely on multimodal data to improve diagnostic accuracy. However, complete multimodal data is rare in real-world applications due to a lack of medical equipment and concerns…

DiagnosticDisentanglementfeature selectionRepresentation Learning

Self-supervised multimodal neuroimaging yields predictive representations for a spectrum of Alzheimer's phenotypes

2022-09-07 · Alex Fedorov, Eloy Geenjaar, Lei Wu, Tristan Sylvain 외

Recent neuroimaging studies that focus on predicting brain disorders via modern machine learning approaches commonly include a single modality and rely on supervised over-parameterized models.However, a single modality p…

DiagnosticSelf-Supervised Learning

Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers

2021-01-31 · Lisa Anne Hendricks, John Mellor, Rosalia Schneider, Jean-Baptiste Alayrac 외

Recently multimodal transformer models have gained popularity because their performance on language and vision tasks suggest they learn rich visual-linguistic representations. Focusing on zero-shot image retrieval tasks,…

Image RetrievalRetrievalSelf-Supervised LearningZero-shot Image Retrieval

Multimodal Understanding Through Correlation Maximization and Minimization

2023-05-04 · Yifeng Shi, Marc Niethammer

Multimodal learning has mainly focused on learning large models on, and fusing feature representations from, different modalities for better performances on downstream tasks. In this work, we take a detour from this tren…

Payload-agnostic Decoupling and Hybrid Vibration Isolation Control for a Maglev Platform with Redundant Actuation

2020-03-31

Payload-specific vibration control may be suitable for a particular task but lacks generality and transferability required for adapting to the various payload. Self-decoupling and robust vibration control are the crucial…