Single-Reference Text-to-Image Manipulation with Dual Contrastive Denoising Score
Large-scale text-to-image generative models have shown remarkable ability to synthesize diverse and high-quality images. However, it is still challenging to directly apply these models for editing real images for two reasons. First, it is difficult for users to come up with a perfect text prompt that accurately describes every visual detail in the input image. Second, while existing models can introduce desirable changes in certain regions, they often dramatically alter the input content and introduce unexpected changes in unwanted regions. To address these challenges, we present Dual Contrastive Denoising Score, a simple yet powerful framework that leverages the rich generative prior of text-to-image diffusion models. Inspired by contrastive learning approaches for unpaired image-to-image translation, we introduce a straightforward dual contrastive loss within the proposed framework. Our approach utilizes the extensive spatial information from the intermediate representations of the self-attention layers in latent diffusion models without depending on auxiliary networks. Our method achieves both flexible content modification and structure preservation between input and output images, as well as zero-shot image-to-image translation. Through extensive experiments, we show that our approach outperforms existing methods in real image editing while maintaining the capability to directly utilize pretrained text-to-image diffusion models without further training.
Code (0)
등록된 구현이 없습니다.
Tasks
Image-to-Image TranslationContrastive LearningImage ManipulationImage EditingSimilar Papers 제목 키워드 기반
Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences
Current approaches to aligning large language models (LLMs) aggregate diverse human preferences into a single reward signal, effectively optimizing for a hypothetical ``average user'' who represents no real person partic…
3DSwapping: Texture Swapping For 3D Object From Single Reference Image
3D texture swapping allows for the customization of 3D object textures, enabling efficient and versatile visual transformations in 3D editing. While no dedicated method exists, adapted 2D editing and text-driven 3D editi…
Zero-Shot Color Image Manipulation Localization via Noise Residual Artifact Pattern Analysis
Digital cameras embed device-specific artifacts into every acquired image through demosaicing, in-camera post-processing, and lossy compression. These traces constitute a forensic signal that can be exploited to assess i…
Image Manipulation LocalizationHighly Personalized Text Embedding for Image Manipulation by Stable Diffusion
Diffusion models have shown superior performance in image generation and manipulation, but the inherent stochasticity presents challenges in preserving and manipulating image content and identity. While previous approach…
Image GenerationImage ManipulationVariable population manipulations of reallocation rules in economies with single-peaked preferences
In a one-commodity economy with single-peaked preferences and individual endowments, we study different ways in which reallocation rules can be strategically distorted by affecting the set of active agents. We introduce …