paper-with-me

Papers

SEGA: Instructing Text-to-Image Models using Semantic Guidance

2023-01-28 · NeurIPS 2023 11 · Manuel Brack, Felix Friedrich, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, Kristian Kersting

Text-to-image diffusion models have recently received a lot of interest for their astonishing ability to produce high-fidelity images from text only. However, achieving one-shot generation that aligns with the user's intent is nearly impossible, yet small changes to the input prompt often result in very different images. This leaves the user with little semantic control. To put the user in control, we show how to interact with the diffusion process to flexibly steer it along semantic directions. This semantic guidance (SEGA) generalizes to any generative architecture using classifier-free guidance. More importantly, it allows for subtle and extensive edits, changes in composition and style, as well as optimizing the overall artistic conception. We demonstrate SEGA's effectiveness on both latent and pixel-based diffusion models such as Stable Diffusion, Paella, and DeepFloyd-IF using a variety of tasks, thus providing strong evidence for its versatility, flexibility, and improvements over existing methods.

📄 PDF Abstract BibTeX arXiv:2301.12247

Code (1)

ml-research/semantic-image-editing pytorch

Similar Papers 제목 키워드 기반

SegAttnGAN: Text to Image Generation with Segmentation Attention

2020-05-25 · Yuchuan Gou, Qiancheng Wu, Minghao Li, Bo Gong 외

In this paper, we propose a novel generative network (SegAttnGAN) that utilizes additional segmentation information for the text-to-image synthesis task. As the segmentation data introduced to the model provides useful g…

Image GenerationSegmentationText to Image GenerationText-to-Image Generation

Semantics-Enhanced Adversarial Nets for Text-to-Image Synthesis

2019-10-01 · ICCV 2019 10 · Hongchen Tan, Xiuping Liu, Xin Li, Yi Zhang 외

This paper presents a new model, Semantics-enhanced Generative Adversarial Network (SEGAN), for fine-grained text-to-image generation. We introduce two modules, a Semantic Consistency Module (SCM) and an Attention Compet…

Generative Adversarial NetworkImage GenerationText to Image GenerationText-to-Image Generation

The Stable Artist: Steering Semantics in Diffusion Latent Space

2022-12-12 · Manuel Brack, Patrick Schramowski, Felix Friedrich, Dominik Hintersdorf 외

Large, text-conditioned generative diffusion models have recently gained a lot of attention for their impressive performance in generating high-fidelity images from text alone. However, achieving high-quality results is …

Image Generation

Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization

2025-05-20 · Yuanyuan Chang, Yinghua Yao, Tao Qin, Mengmeng Wang 외

Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by th…

AttributeDisentanglementImage Generation

PoseGAN: A Pose-to-Image Translation Framework for Camera Localization

2020-06-23 · Kanglin Liu, Qing Li, Guoping Qiu

Camera localization is a fundamental requirement in robotics and computer vision. This paper introduces a pose-to-image translation framework to tackle the camera localization problem. We present PoseGANs, a conditional …

Camera LocalizationPose EstimationTranslation