paper-with-me

Papers

ShaLa: Multimodal Shared Latent Space Modelling

2025-08-24 · Jiali Cui, Yan-Ying Chen, Yanxia Zhang, Matthew Klenk arxiv

This paper presents a novel generative framework for learning shared latent representations across multimodal data. Many advanced multimodal methods focus on capturing all combinations of modality-specific details across inputs, which can inadvertently obscure the high-level semantic concepts that are shared across modalities. Notably, Multimodal VAEs with low-dimensional latent variables are designed to capture shared representations, enabling various tasks such as joint multimodal synthesis and cross-modal inference. However, multimodal VAEs often struggle to design expressive joint variational posteriors and suffer from low-quality synthesis. In this work, ShaLa addresses these challenges by integrating a novel architectural inference model and a second-stage expressive diffusion prior, which not only facilitates effective inference of shared latent representation but also significantly improves the quality of downstream multimodal synthesis. We validate ShaLa extensively across multiple benchmarks, demonstrating superior coherence and synthesis quality compared to state-of-the-art multimodal VAEs. Furthermore, ShaLa scales to many more modalities while prior multimodal VAEs have fallen short in capturing the increasing complexity of the shared latent space.

📄 PDF Abstract BibTeX arXiv:2508.17376

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Shared latent subspace modelling within Gaussian-Binary Restricted Boltzmann Machines for NIST i-Vector Challenge 2014

2015-03-18 · Danila Doroshin, Alexander Yamshinin, Nikolay Lubimov, Marina Nastasenko 외

This paper presents a novel approach to speaker subspace modelling based on Gaussian-Binary Restricted Boltzmann Machines (GRBM). The proposed model is based on the idea of shared factors as in the Probabilistic Linear D…

parameter estimationSpeaker Verification

LatentUMM: Dual Latent Alignment for Unified Multimodal Models

2026-05-18 · Yinyi Luo, Wenwen Wang, Hayes Bai, Marios Savvides 외 arxiv

Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We obser…

Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision

2026-05-01 · Jiali Cui, Zhiqiang Lao, Heather Yu arxiv

Energy-based models (EBMs) are a flexible class of deep generative models and are well-suited to capture complex dependencies in multimodal data. However, learning multimodal EBM by maximum likelihood requires Markov Cha…

SanskritShala: A Neural Sanskrit NLP Toolkit with Web-Based Interface for Pedagogical and Annotation Purposes

2023-02-19 · Jivnesh Sandhan, Anshul Agarwal, Laxmidhar Behera, Tushar Sandhan 외

We present a neural Sanskrit Natural Language Processing (NLP) toolkit named SanskritShala (a school of Sanskrit) to facilitate computational linguistic analyses for several tasks such as word segmentation, morphological…

Dependency ParsingMorphological TaggingWord EmbeddingsWord Similarity

Learning Sequential Latent Variable Models from Multimodal Time Series Data

2022-04-21 · Oliver Limoyo, Trevor Ablett, Jonathan Kelly

Sequential modelling of high-dimensional data is an important problem that appears in many domains including model-based reinforcement learning and dynamics identification for control. Latent variable models applied to s…

Model-based Reinforcement LearningTime SeriesTime Series Analysis