paper-with-me

Papers

Joint Multimodal Learning with Deep Generative Models

2016-11-07 · Masahiro Suzuki, Kotaro Nakayama, Yutaka Matsuo

We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. Recently, some studies handle multiple modalities on deep generative models, such as variational autoencoders (VAEs). However, these models typically assume that modalities are forced to have a conditioned relation, i.e., we can only generate modalities in one direction. To achieve our objective, we should extract a joint representation that captures high-level concepts among all modalities and through which we can exchange them bi-directionally. As described herein, we propose a joint multimodal variational autoencoder (JMVAE), in which all modalities are independently conditioned on joint representation. In other words, it models a joint distribution of modalities. Furthermore, to be able to generate missing modalities from the remaining modalities properly, we develop an additional method, JMVAE-kl, that is trained by reducing the divergence between JMVAE's encoder and prepared networks of respective modalities. Our experiments show that our proposed method can obtain appropriate joint representation from multiple modalities and that it can generate and reconstruct them more properly than conventional VAEs. We further demonstrate that JMVAE can generate multiple modalities bi-directionally.

📄 PDF Abstract BibTeX arXiv:1611.01891

Code (2)

masa-su/Tars pytorch
masa-su/jmvae

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

MHVAE: a Human-Inspired Deep Hierarchical Generative Model for Multimodal Representation Learning

2020-06-04 · Miguel Vasco, Francisco S. Melo, Ana Paiva

Humans are able to create rich representations of their external reality. Their internal representations allow for cross-modality inference, where available perceptions can induce the perceptual experience of missing inp…

Representation Learning

Multimodal Generative Learning Utilizing Jensen-Shannon-Divergence

2020-06-15 · NeurIPS 2020 12 · Thomas M. Sutter, Imant Daunhawer, Julia E. Vogt

Learning from different data types is a long-standing goal in machine learning research, as multiple information sources co-occur when describing natural phenomena. However, existing generative models that approximate a …

Learning Multimodal Latent Generative Models with Energy-Based Prior

2024-09-30 · Shiyu Yuan, Jiali Cui, Hanao Li, Tian Han

Multimodal generative models have recently gained significant attention for their ability to learn representations across various modalities, enhancing joint and cross-generation coherence. However, most existing works u…

Learning Factorized Multimodal Representations

2018-06-16 · ICLR 2019 5 · Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency 외

Learning multimodal representations is a fundamentally complex research problem due to the presence of multiple heterogeneous sources of information. Although the presence of multiple modalities provides additional valua…

Representation Learning

Unified Multimodal Discrete Diffusion

2025-03-26 · Alexander Swerdlow, Mihir Prabhudesai, Siddharth Gandhi, Deepak Pathak 외

Multimodal generative models that can understand and generate across multiple modalities are dominated by autoregressive (AR) approaches, which process tokens sequentially from left to right, or top to bottom. These mode…

Image CaptioningImage GenerationQuestion AnsweringText Generation