paper-with-me

홈 › Papers

Diffusion Models already have a Semantic Latent Space

2022-10-20 · Mingi Kwon, Jaeseok Jeong, Youngjung Uh

Diffusion models achieve outstanding generative performance in various domains. Despite their great success, they lack semantic latent space which is essential for controlling the generative process. To address the problem, we propose asymmetric reverse process (Asyrp) which discovers the semantic latent space in frozen pretrained diffusion models. Our semantic latent space, named h-space, has nice properties for accommodating semantic image manipulation: homogeneity, linearity, robustness, and consistency across timesteps. In addition, we introduce a principled design of the generative process for versatile editing and quality boost ing by quantifiable measures: editing strength of an interval and quality deficiency at a timestep. Our method is applicable to various architectures (DDPM++, iD- DPM, and ADM) and datasets (CelebA-HQ, AFHQ-dog, LSUN-church, LSUN- bedroom, and METFACES). Project page: https://kwonminki.github.io/Asyrp/

📄 PDF Abstract BibTeX arXiv:2210.10960

Code (1)

kwonminki/Asyrp_official 공식 구현 pytorch

Tasks

Image Manipulation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Boundary Guided Learning-Free Semantic Control with Diffusion Models

2023-02-16 · NeurIPS 2023 11 · Ye Zhu, Yu Wu, Zhiwei Deng, Olga Russakovsky 외

Applying pre-trained generative denoising diffusion models (DDMs) for downstream tasks such as image semantic editing usually requires either fine-tuning DDMs or learning auxiliary editing networks in the existing litera…

Denoising

CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning

2025-03-09 · Lei Shi, Andreas Bulling

We propose CLAD -- a Constrained Latent Action Diffusion model for vision-language procedure planning in instructional videos. Procedure planning is the challenging task of predicting intermediate actions given a visual …

DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment

2026-03-23 · Xin Cai, Zhiyuan You, Zhoutong Zhang, Tianfan Xue arxiv

Reducing token count is crucial for efficient training and inference of latent diffusion models, especially at high resolution. A common strategy is to build high-compression image tokenizers with more channels per token…

Image Generation

Decoding Diffusion: A Scalable Framework for Unsupervised Analysis of Latent Space Biases and Representations Using Natural Language Prompts

2024-10-25 · E. Zhixuan Zeng, Yuhao Chen, Alexander Wong

Recent advances in image generation have made diffusion models powerful tools for creating high-quality images. However, their iterative denoising process makes understanding and interpreting their semantic latent spaces…

DenoisingImage CaptioningImage Generation

Hierarchical Diffusion Autoencoders and Disentangled Image Manipulation

2023-04-24 · Zeyu Lu, Chengyue Wu, Xinyuan Chen, Yaohui Wang 외

Diffusion models have attained impressive visual quality for image synthesis. However, how to interpret and manipulate the latent space of diffusion models has not been extensively explored. Prior work diffusion autoenco…

Image GenerationImage ManipulationImage Reconstruction