paper-with-me

홈 › Papers

Adversarial Supervision Makes Layout-to-Image Diffusion Models Thrive

2024-01-16 · Yumeng Li, Margret Keuper, Dan Zhang, Anna Khoreva

Despite the recent advances in large-scale diffusion models, little progress has been made on the layout-to-image (L2I) synthesis task. Current L2I models either suffer from poor editability via text or weak alignment between the generated image and the input layout. This limits their usability in practice. To mitigate this, we propose to integrate adversarial supervision into the conventional training pipeline of L2I diffusion models (ALDM). Specifically, we employ a segmentation-based discriminator which provides explicit feedback to the diffusion generator on the pixel-level alignment between the denoised image and the input layout. To encourage consistent adherence to the input layout over the sampling steps, we further introduce the multistep unrolling strategy. Instead of looking at a single timestep, we unroll a few steps recursively to imitate the inference process, and ask the discriminator to assess the alignment of denoised images with the layout over a certain time window. Our experiments show that ALDM enables layout faithfulness of the generated images, while allowing broad editability via text prompts. Moreover, we showcase its usefulness for practical applications: by synthesizing target distribution samples via text control, we improve domain generalization of semantic segmentation models by a large margin (~12 mIoU points).

📄 PDF Abstract BibTeX arXiv:2401.08815

Code (1)

boschresearch/aldm 공식 구현 pytorch

Tasks

Domain GeneralizationImage GenerationLayout-to-Image GenerationSemantic Segmentation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision

2025-12-24 · Weiqi Li, Zehao Zhang, Liang Lin, Guangrun Wang arxiv

Controllability is a fundamental requirement in video synthesis, where accurate alignment with conditioning signals is essential. Existing classifier-free guidance methods typically achieve conditioning indirectly by mod…

Video Generation

LayoutDM: Transformer-based Diffusion Model for Layout Generation

2023-05-04 · CVPR 2023 1 · Shang Chai, Liansheng Zhuang, Fengying Yan

Automatic layout generation that can synthesize high-quality layouts is an important tool for graphic design in many applications. Though existing methods based on generative models such as Generative Adversarial Network…

DenoisingDiversityLayout Generationmodel

Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation

2024-08-27 · Abdelrahman Eldesokey, Peter Wonka

We propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control. Layout control has been widely studied to alleviate the shortcomings of T2I diffusion models in understanding o…

Image GenerationObjectScene Generation

ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion Model

2024-12-20 · Qi Zang, Jiayi Yang, Shuang Wang, Dong Zhao 외

Data-driven deep learning models have enabled tremendous progress in change detection (CD) with the support of pixel-level annotations. However, collecting diverse data and manually annotating them is costly, laborious, …

Change Detection

LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image Generation

2023-08-09 · Leigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie 외

In the text-to-image generation field, recent remarkable progress in Stable Diffusion makes it possible to generate rich kinds of novel photorealistic images. However, current models still face misalignment issues (e.g.,…

Image GenerationIn-Context LearningText to Image GenerationText-to-Image Generation