paper-with-me

Papers

Identifiable Shared Component Analysis of Unpaired Multimodal Mixtures

2024-09-28 · Subash Timilsina, Sagar Shrestha, Xiao Fu

A core task in multi-modal learning is to integrate information from multiple feature spaces (e.g., text and audio), offering modality-invariant essential representations of data. Recent research showed that, classical tools such as {\it canonical correlation analysis} (CCA) provably identify the shared components up to minor ambiguities, when samples in each modality are generated from a linear mixture of shared and private components. Such identifiability results were obtained under the condition that the cross-modality samples are aligned/paired according to their shared information. This work takes a step further, investigating shared component identifiability from multi-modal linear mixtures where cross-modality samples are unaligned. A distribution divergence minimization-based loss is proposed, under which a suite of sufficient conditions ensuring identifiability of the shared components are derived. Our conditions are based on cross-modality distribution discrepancy characterization and density-preserving transform removal, which are much milder than existing studies relying on independent component analysis. More relaxed conditions are also provided via adding reasonable structural constraints, motivated by available side information in various applications. The identifiability claims are thoroughly validated using synthetic and real-world data.

📄 PDF Abstract BibTeX arXiv:2409.19422

Code (1)

XiaoFuLab/Shared-Component-Analysis 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

2025-10-09 · Sharut Gupta, Shobhita Sundaram, Chenyu Wang, Stefanie Jegelka 외 arxiv

Traditional multimodal learners find unified representations for tasks like visual question answering, but rely heavily on paired datasets. However, an overlooked yet potentially powerful question is: can one leverage au…

Visual Question AnsweringRepresentation Learning

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck

2019-08-19 · ICCV 2019 10 · Shuang Ma, Daniel McDuff, Yale Song

Deep generative models have led to significant advances in cross-modal generation such as text-to-image synthesis. Training these models typically requires paired data with direct correspondence between modalities. We in…

Image GenerationSpeech Synthesis

Shared Independent Component Analysis for Multi-Subject Neuroimaging

2021-10-26 · NeurIPS 2021 12 · Hugo Richard, Pierre Ablin, Bertrand Thirion, Alexandre Gramfort 외

We consider shared response modeling, a multi-view learning problem where one wants to identify common components from multiple datasets or views. We introduce Shared Independent Component Analysis (ShICA) that models ea…

MULTI-VIEW LEARNING

Identifiable Multimodal Causal Representation Learning under Partial Latent Sharing

2026-05-18 · Manal Benhamza, Marianne Clausel, Myriam Tami arxiv

Causal representation learning (CRL) seeks to uncover meaningful latent variables and their corresponding causal structure from high-dimensional observational data. Although its significance, CRL identifiability remains …

Representation Learning

Learning Multimodal Representations for Unseen Activities

2018-06-21 · AJ Piergiovanni, Michael S. Ryoo

We present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constraints on the embedding space using paired…

General ClassificationTemporal Action Localization