paper-with-me

홈 › Papers

MHVAE: a Human-Inspired Deep Hierarchical Generative Model for Multimodal Representation Learning

2020-06-04 · Miguel Vasco, Francisco S. Melo, Ana Paiva

Humans are able to create rich representations of their external reality. Their internal representations allow for cross-modality inference, where available perceptions can induce the perceptual experience of missing input modalities. In this paper, we contribute the Multimodal Hierarchical Variational Auto-encoder (MHVAE), a hierarchical multimodal generative model for representation learning. Inspired by human cognitive models, the MHVAE is able to learn modality-specific distributions, of an arbitrary number of modalities, and a joint-modality distribution, responsible for cross-modality inference. We formally derive the model's evidence lower bound and propose a novel methodology to approximate the joint-modality posterior based on modality-specific representation dropout. We evaluate the MHVAE on standard multimodal datasets. Our model performs on par with other state-of-the-art generative models regarding joint-modality reconstruction from arbitrary input modalities and cross-modality inference.

📄 PDF Abstract BibTeX arXiv:2006.02991

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Unified Cross-Modal Image Synthesis with Hierarchical Mixture of Product-of-Experts

2024-10-25 · Reuben Dorent, Nazim Haouchine, Alexandra Golby, Sarah Frisken 외

We propose a deep mixture of multimodal hierarchical variational auto-encoders called MMHVAE that synthesizes missing images from observed images in different modalities. MMHVAE's design focuses on tackling four challeng…

Image Generation

Unified Brain MR-Ultrasound Synthesis using Multi-Modal Hierarchical Representations

2023-09-15 · Reuben Dorent, Nazim Haouchine, Fryderyk Kögl, Samuel Joutard 외

We introduce MHVAE, a deep hierarchical variational auto-encoder (VAE) that synthesizes missing images from various modalities. Extending multi-modal VAEs with a hierarchical latent structure, we introduce a probabilisti…

Hierarchical Attention Fusion of Visual and Textual Representations for Cross-Domain Sequential Recommendation

2025-04-21 · Wangyu Wu, Zhenhong Chen, Siqi Song, Xianglin Qiua 외

Cross-Domain Sequential Recommendation (CDSR) predicts user behavior by leveraging historical interactions across multiple domains, focusing on modeling cross-domain preferences through intra- and inter-sequence item rel…

Decision MakingSequential Decision MakingSequential Recommendation

Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs

2026-06-11 · Huyen Vo, María Martínez-García, Isabel Valera arxiv

Existing approaches for multimodal variational autoencoders (VAEs) face a trade-off between generative quality and coherence-i.e., they struggle to generate realistic and diverse samples that, at the same time, are seman…

From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control

2025-05-28 · Jusheng Zhang, Jinzhou Tang, Sidi Liu, Mingyan Li 외

Human motion generative modeling or synthesis aims to characterize complicated human motions of daily activities in diverse real-world environments. However, current research predominantly focuses on either low-level, sh…

Motion GenerationMotion PlanningTask and Motion Planning