paper-with-me

Papers

Score-Based Multimodal Autoencoder

2023-05-25 · Daniel Wesego, Pedram Rooshenas

Multimodal Variational Autoencoders (VAEs) represent a promising group of generative models that facilitate the construction of a tractable posterior within the latent space given multiple modalities. Previous studies have shown that as the number of modalities increases, the generative quality of each modality declines. In this study, we explore an alternative approach to enhance the generative performance of multimodal VAEs by jointly modeling the latent space of independently trained unimodal VAEs using score-based models (SBMs). The role of the SBM is to enforce multimodal coherence by learning the correlation among the latent variables. Consequently, our model combines a better generative quality of unimodal VAEs with coherent integration across different modalities using the latent score-based model. In addition, our approach provides the best unconditional coherence.

📄 PDF Abstract BibTeX arXiv:2305.15708

Code (1)

rooshenasgroup/sbmae 공식 구현 jax

Similar Papers 제목 키워드 기반

Comparison of Autoencoders for tokenization of ASL datasets

2025-01-12 · Vouk Praun-Petrovic, Aadhvika Koundinya, Lavanya Prahallad

Generative AI, powered by large language models (LLMs), has revolutionized applications across text, audio, images, and video. This study focuses on developing and evaluating encoder-decoder architectures for the America…

DecoderDenoisingImage ReconstructionSign Language Recognition

A Novel Autoencoders-LSTM Model for Stroke Outcome Prediction using Multimodal MRI Data

2023-03-16 · Nima Hatami, Laura Mechtouff, David Rousseau, Tae-Hee Cho 외

Patient outcome prediction is critical in management of ischemic stroke. In this paper, a novel machine learning model is proposed for stroke outcome prediction using multimodal Magnetic Resonance Imaging (MRI). The prop…

Management

Convolutional autoencoder-based multimodal one-class classification

2023-09-25 · Firas Laakom, Fahad Sohrab, Jenni Raitoharju, Alexandros Iosifidis 외

One-class classification refers to approaches of learning using data from a single class only. In this paper, we propose a deep learning one-class classification method suitable for multimodal data, which relies on two c…

ClassificationDiversityimage-classificationImage Classification+1

Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders

2026-02-08 · Sayantan Kumar, Peijie Qiu, Aristeidis Sotiras arxiv

Normative modeling learns a healthy reference distribution and quantifies subject-specific deviations to capture heterogeneous disease effects. In Alzheimers disease (AD), multimodal neuroimaging offers complementary sig…

Outlier Detection

Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection

2026-05-24 · Prabuddhi Wariyapperuma, Rajitha de Silva, Marc Hanheide, Thomas Bohné 외 arxiv

Accurate 3D bird's-eye view (BEV) object detection is essential for autonomous driving, and depends strongly on effective multimodal representations from complementary sensors such as cameras and LiDAR. Multimodal masked…

3D Object DetectionAutonomous Driving