paper-with-me

Papers

Multimodal Transformer for Parallel Concatenated Variational Autoencoders

2022-10-28 · Stephen D. Liang, Jerry M. Mendel

In this paper, we propose a multimodal transformer using parallel concatenated architecture. Instead of using patches, we use column stripes for images in R, G, B channels as the transformer input. The column stripes keep the spatial relations of original image. We incorporate the multimodal transformer with variational autoencoder for synthetic cross-modal data generation. The multimodal transformer is designed using multiple compression matrices, and it serves as encoders for Parallel Concatenated Variational AutoEncoders (PC-VAE). The PC-VAE consists of multiple encoders, one latent space, and two decoders. The encoders are based on random Gaussian matrices and don't need any training. We propose a new loss function based on the interaction information from partial information decomposition. The interaction information evaluates the input cross-modal information and decoder output. The PC-VAE are trained via minimizing the loss function. Experiments are performed to validate the proposed multimodal transformer for PC-VAE.

📄 PDF Abstract BibTeX arXiv:2210.16174

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

Normative Modeling using Multimodal Variational Autoencoders to Identify Abnormal Brain Structural Patterns in Alzheimer Disease

2021-10-10 · Sayantan Kumar, Philip Payne, Aristeidis Sotiras

Normative modelling is an emerging method for understanding the underlying heterogeneity within brain disorders like Alzheimer Disease (AD) by quantifying how each patient deviates from the expected normative pattern tha…

GPR

CoVAE: correlated multimodal generative modeling

2026-03-02 · Federico Caretti, Guido Sanguinetti arxiv

Multimodal Variational Autoencoders have emerged as a popular tool to extract effective representations from rich multimodal data. However, such models rely on fusion strategies in latent space that destroy the joint sta…

Discrete Variational Autoencoding via Policy Search

2025-09-29 · Michael Drolet, Firas Al-Hafez, Aditya Bhatt, Jan Peters 외 arxiv

Discrete latent bottlenecks in variational autoencoders (VAEs) offer high bit efficiency and can be modeled with autoregressive discrete distributions, enabling parameter-efficient multimodal search with transformers. Ho…

Image Reconstruction

A survey of multimodal deep generative models

2022-07-05 · Masahiro Suzuki, Yutaka Matsuo

Multimodal learning is a framework for building models that make predictions based on different types of modalities. Important challenges in multimodal learning are the inference of shared representations from arbitrary …

Survey

Addressing Posterior Collapse with Mutual Information for Improved Variational Neural Machine Translation

2020-07-01 · ACL 2020 6 · Arya D. McCarthy, Xi-An Li, Jiatao Gu, Ning Dong

This paper proposes a simple and effective approach to address the problem of posterior collapse in conditional variational autoencoders (CVAEs). It thus improves performance of machine translation models that use noisy …

de-enMachine TranslationNMTTranslation