paper-with-me

Papers

Robust Text-driven Image Editing Method that Adaptively Explores Directions in Latent Spaces of StyleGAN and CLIP

2023-04-03 · Tsuyoshi Baba, Kosuke Nishida, Kyosuke Nishida

Automatic image editing has great demands because of its numerous applications, and the use of natural language instructions is essential to achieving flexible and intuitive editing as the user imagines. A pioneering work in text-driven image editing, StyleCLIP, finds an edit direction in the CLIP space and then edits the image by mapping the direction to the StyleGAN space. At the same time, it is difficult to tune appropriate inputs other than the original image and text instructions for image editing. In this study, we propose a method to construct the edit direction adaptively in the StyleGAN and CLIP spaces with SVM. Our model represents the edit direction as a normal vector in the CLIP space obtained by training a SVM to classify positive and negative images. The images are retrieved from a large-scale image corpus, originally used for pre-training StyleGAN, according to the CLIP similarity between the images and the text instruction. We confirmed that our model performed as well as the StyleCLIP baseline, whereas it allows simple inputs without increasing the computational time.

📄 PDF Abstract BibTeX arXiv:2304.00964

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Feedforward Network A Feedforward Network, or a Multilayer Perceptron (MLP), is a neural network with solely densely connected layers. This is the classic neural network architecture of the…
R1 Regularization R_INLINE_MATH_1 Regularization is a regularization technique and gradient penalty for training [generative adversarial…
SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Adaptive Instance Normalization 설명 없음

Similar Papers 제목 키워드 기반

TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation

2025-09-26 · Qihang Wang, Yaxiong Wang, Lechao Cheng, Zhun Zhong arxiv

This paper explores image editing under the joint control of text and drag interactions. While recent advances in text-driven and drag-driven editing have achieved remarkable progress, they suffer from complementary limi…

Image ManipulationImage Editing

Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

2023-05-29 · Fu-Yun Wang, Wenshuo Chen, Guanglu Song, Han-Jia Ye 외

Leveraging large-scale image-text datasets and advancements in diffusion models, text-driven generative models have made remarkable strides in the field of image generation and editing. This study explores the potential …

DenoisingImage GenerationText-to-Video EditingVideo Generation

Affective Image Editing: Shaping Emotional Factors via Text Descriptions

2025-05-24 · Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang 외

In daily life, images as common affective stimuli have widespread applications. Despite significant progress in text-driven image editing, there is limited work focusing on understanding users' emotional requests. In thi…

IE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception Alignment

2025-01-17 · Shangkun Sun, Bowen Qu, Xiaoyu Liang, Songlin Fan 외

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image…

Image Generation

HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads

2024-11-22 · Yu Xu, Fan Tang, Juan Cao, Yuxin Zhang 외

Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based…

Image Generationtext-guided-image-editing