paper-with-me

홈 › Papers

Diffusion Models For Multi-Modal Generative Modeling

2024-07-24 · Changyou Chen, Han Ding, Bunyamin Sisman, Yi Xu, Ouye Xie, Benjamin Z. Yao, Son Dinh Tran, Belinda Zeng

Diffusion-based generative modeling has been achieving state-of-the-art results on various generation tasks. Most diffusion models, however, are limited to a single-generation modeling. Can we generalize diffusion models with the ability of multi-modal generative training for more generalizable modeling? In this paper, we propose a principled way to define a diffusion model by constructing a unified multi-modal diffusion model in a common diffusion space. We define the forward diffusion process to be driven by an information aggregation from multiple types of task-data, e.g., images for a generation task and labels for a classification task. In the reverse process, we enforce information sharing by parameterizing a shared backbone denoising network with additional modality-specific decoder heads. Such a structure can simultaneously learn to generate different types of multi-modal data with a multi-task loss, which is derived from a new multi-modal variational lower bound that generalizes the standard diffusion model. We propose several multimodal generation settings to verify our framework, including image transition, masked-image training, joint image-label and joint image-representation generative modeling. Extensive experimental results on ImageNet indicate the effectiveness of our framework for various multi-modal generative modeling, which we believe is an important research direction worthy of more future explorations.

📄 PDF Abstract BibTeX arXiv:2407.17571

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDenoisingmultimodal generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond

2024-09-23 · Hong Chen, Xin Wang, Yuwei Zhou, Bin Huang 외

Multi-modal generative AI has received increasing attention in both academia and industry. Particularly, two dominant families of techniques are: i) The multi-modal large language model (MLLM) such as GPT-4V, which shows…

Language ModellingLarge Language ModelMixture-of-ExpertsVideo Generation

DiffSpectra: Molecular Structure Elucidation from Spectra using Diffusion Models

2025-07-09 · Liang Wang, Yu Rong, Tingyang Xu, Zhenyi Zhong 외 arxiv

Molecular structure elucidation from spectra is a fundamental challenge in molecular science. Conventional approaches rely heavily on expert interpretation and lack scalability, while retrieval-based machine learning app…

Generator Matching: Generative modeling with arbitrary Markov processes

2024-10-27 · Peter Holderrieth, Marton Havasi, Jason Yim, Neta Shaul 외

We introduce generator matching, a modality-agnostic framework for generative modeling using arbitrary Markov processes. Generators characterize the infinitesimal evolution of a Markov process, which we leverage for gene…

Image Generation

DiTS: Multimodal Diffusion Transformers Are Time Series Forecasters

2026-02-06 · Haoran Zhang, Haixuan Liu, Yong Liu, Yunzhong Qiu 외 arxiv

While generative modeling on time series facilitates more capable and flexible probabilistic forecasting, existing generative time series models do not address the multi-dimensional properties of time series data well. T…

Video Generation

Multi-modal Latent Diffusion

2023-06-07 · Mustapha Bounoua, Giulio Franzese, Pietro Michiardi

Multi-modal data-sets are ubiquitous in modern applications, and multi-modal Variational Autoencoders are a popular family of models that aim to learn a joint representation of the different modalities. However, existing…

multimodal generation