paper-with-me

홈 › Papers

Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt

2024-04-08 · Zhiqi Huang, Huixin Xiong, Haoyu Wang, Longguang Wang, Zhiheng Li

Text-to-image generation has witnessed great progress, especially with the recent advancements in diffusion models. Since texts cannot provide detailed conditions like object appearance, reference images are usually leveraged for the control of objects in the generated images. However, existing methods still suffer limited accuracy when the relationship between the foreground and background is complicated. To address this issue, we develop a framework termed Mask-ControlNet by introducing an additional mask prompt. Specifically, we first employ large vision models to obtain masks to segment the objects of interest in the reference image. Then, the object images are employed as additional prompts to facilitate the diffusion model to better understand the relationship between foreground and background regions during image generation. Experiments show that the mask prompts enhance the controllability of the diffusion model to maintain higher fidelity to the reference image while achieving better image quality. Comparison with previous text-to-image generation methods demonstrates our method's superior quantitative and qualitative performance on the benchmark datasets.

📄 PDF Abstract BibTeX arXiv:2404.05331

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models

2023-05-25 · NeurIPS 2023 11 · Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao 외

Text-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. However, despite their success, text descri…

All

Active Learning Inspired ControlNet Guidance for Augmenting Semantic Segmentation Datasets

2025-03-12 · Hannah Kniesel, Pedro Hermosilla, Timo Ropinski

Recent advances in conditional image generation from diffusion models have shown great potential in achieving impressive image quality while preserving the constraints introduced by the user. In particular, ControlNet en…

Active LearningConditional Image GenerationImage GenerationSemantic Segmentation+1

E-Commerce Inpainting with Mask Guidance in Controlnet for Reducing Overcompletion

2024-09-15 · Guandong Li

E-commerce image generation has always been one of the core demands in the e-commerce field. The goal is to restore the missing background that matches the main product given. In the post-AIGC era, diffusion models are p…

Image Generation

OmniControlNet: Dual-stage Integration for Conditional Image Generation

2024-06-09 · Yilin Wang, Haiyang Xu, Xiang Zhang, Zeyuan Chen 외

We provide a two-way integration for the widely adopted ControlNet by integrating external condition generation algorithms into a single dense prediction method and incorporating its individually trained image generation…

Conditional Image GenerationImage GenerationText to Image GenerationText-to-Image Generation

Enhancing Prompt Following with Visual Control Through Training-Free Mask-Guided Diffusion

2024-04-23 · Hongyu Chen, Yiqi Gao, Min Zhou, Peng Wang 외

Recently, integrating visual controls into text-to-image~(T2I) models, such as ControlNet method, has received significant attention for finer control capabilities. While various training-free methods make efforts to enh…

AttributeObject