paper-with-me

Papers

Importance-based Multimodal Autoencoder

2021-01-01 · Sayan Ghosh, Eugene Laksana, Louis-Philippe Morency, Stefan Scherer

Integrating information from multiple modalities (e.g., verbal, acoustic and visual data) into meaningful representations has seen great progress in recent years. However, two challenges are not sufficiently addressed by current approaches: (1) computationally efficient training of multimodal autoencoder networks which are robust in the absence of modalities, and (2) unsupervised learning of important subspaces in each modality which are correlated with other modalities. In this paper we propose the IMA (Importance-based Multimodal Autoencoder) model, a scalable model that learns modality importances and robust multimodal representations through a novel cross-covariance based loss function. We conduct experiments on MNIST-TIDIGITS a multimodal dataset of spoken and image digits,and on IEMOCAP, a multimodal emotion corpus. The IMA model is able to distinguish digits from uncorrelated noise, and word-level importances learnt that correspond to the separation between function and emotional words. The multimodal representations learnt by IMA are also competitive with state-of-the-art baseline approaches on downstream tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Variational autoencoder with weighted samples for high-dimensional non-parametric adaptive importance sampling

2023-10-13 · Julien Demange-Chryst, François Bachoc, Jérôme Morio, Timothé Krauth

Probability density function estimation with weighted samples is the main foundation of all adaptive importance sampling algorithms. Classically, a target distribution is approximated either by a non-parametric model or …

Analyzing Multimodal Integration in the Variational Autoencoder from an Information-Theoretic Perspective

2024-11-01 · Carlotta Langer, Yasmin Kim Georgie, Ilja Porohovoj, Verena Vanessa Hafner 외

Human perception is inherently multimodal. We integrate, for instance, visual, proprioceptive and tactile information into one experience. Hence, multimodal learning is of importance for building robotic systems that aim…

Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders

2026-02-08 · Sayantan Kumar, Peijie Qiu, Aristeidis Sotiras arxiv

Normative modeling learns a healthy reference distribution and quantifies subject-specific deviations to capture heterogeneous disease effects. In Alzheimers disease (AD), multimodal neuroimaging offers complementary sig…

Outlier Detection

Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization

2025-09-30 · Xintong Li, Chuhan Wang, Junda Wu, Rohan Surana 외 arxiv

Direct Preference Optimization (DPO) has recently been extended from text-only models to vision-language models. However, existing methods rely on oversimplified pairwise comparisons, generating a single negative image v…

Generative Adversarial Networks for High-Dimensional Item Factor Analysis: A Deep Adversarial Learning Algorithm

2025-02-15 · Nanyu Luo, Feng Ji

Advances in deep learning and representation learning have transformed item factor analysis (IFA) in the item response theory (IRT) literature by enabling more efficient and accurate parameter estimation. Variational Aut…

parameter estimationRepresentation Learning