Variational methods for Conditional Multimodal Deep Learning
In this paper, we address the problem of conditional modality learning, whereby one is interested in generating one modality given the other. While it is straightforward to learn a joint distribution over multiple modalities using a deep multimodal architecture, we observe that such models aren't very effective at conditional generation. Hence, we address the problem by learning conditional distributions between the modalities. We use variational methods for maximizing the corresponding conditional log-likelihood. The resultant deep model, which we refer to as conditional multimodal autoencoder (CMMA), forces the latent representation obtained from a single modality alone to be `close' to the joint representation obtained from multiple modalities. We use the proposed model to generate faces from attributes. We show that the faces generated from attributes using the proposed model, are qualitatively and quantitatively more representative of the attributes from which they were generated, than those obtained by other deep generative models. We also propose a secondary task, whereby the existing faces are modified by modifying the corresponding attributes. We observe that the modifications in face introduced by the proposed model are representative of the corresponding modifications in attributes.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningMultimodal Deep LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Improving VAE generations of multimodal data through data-dependent conditional priors
One of the major shortcomings of variational autoencoders is the inability to produce generations from the individual modalities of data originating from mixture distributions. This is primarily due to the use of a simpl…
Bridging the inference gap in Mutimodal Variational Autoencoders
From medical diagnosis to autonomous vehicles, critical applications rely on the integration of multiple heterogeneous data modalities. Multimodal Variational Autoencoders offer versatile and scalable methods for generat…
Autonomous VehiclesMedical DiagnosisVariational InferenceEvidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders
Discrete latent spaces in variational autoencoders have been shown to effectively capture the data distribution for many real-world problems such as natural language understanding, human intent prediction, and visual sce…
Image GenerationMotion PlanningNatural Language UnderstandingIncreasing the Generalisation Capacity of Conditional VAEs
We address the problem of one-to-many mappings in supervised learning, where a single instance has many different solutions of possibly equal cost. The framework of conditional variational autoencoders describes a class …
Structured PredictionImproving Multimodal Joint Variational Autoencoders through Normalizing Flows and Correlation Analysis
We propose a new multimodal variational autoencoder that enables to generate from the joint distribution and conditionally to any number of complex modalities. The unimodal posteriors are conditioned on the Deep Canonica…
Diversity