paper-with-me

홈 › Papers

iEdit: Localised Text-guided Image Editing with Weak Supervision

2023-05-10 · Rumeysa Bodur, Erhan Gundogdu, Binod Bhattarai, Tae-Kyun Kim, Michael Donoser, Loris Bazzani

Diffusion models (DMs) can generate realistic images with text guidance using large-scale datasets. However, they demonstrate limited controllability in the output space of the generated images. We propose a novel learning method for text-guided image editing, namely \texttt{iEdit}, that generates images conditioned on a source image and a textual edit prompt. As a fully-annotated dataset with target images does not exist, previous approaches perform subject-specific fine-tuning at test time or adopt contrastive learning without a target image, leading to issues on preserving the fidelity of the source image. We propose to automatically construct a dataset derived from LAION-5B, containing pseudo-target images with their descriptive edit prompts given input image-caption pairs. This dataset gives us the flexibility of introducing a weakly-supervised loss function to generate the pseudo-target image from the latent noise of the source image conditioned on the edit prompt. To encourage localised editing and preserve or modify spatial structures in the image, we propose a loss function that uses segmentation masks to guide the editing during training and optionally at inference. Our model is trained on the constructed dataset with 200K samples and constrained GPU resources. It shows favourable results against its counterparts in terms of image fidelity, CLIP alignment score and qualitatively for editing both generated and real images.

📄 PDF Abstract BibTeX arXiv:2305.05947

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDescriptiveGPUtext-guided-image-editing

Methods 이 논문이 사용한 방법론

Test 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

2024-02-20 · Jianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo 외

Recent advances in text-guided video editing have showcased promising results in appearance editing (e.g., stylization). However, video motion editing in the temporal dimension (e.g., from eating to waving), which distin…

Video Editing

Evaluating Image Editing with LLMs: A Comprehensive Benchmark and Intermediate-Layer Probing Approach

2026-03-20 · Shiqi Gao, Zitong Xu, Kang Fu, Huiyu Duan 외 arxiv

Evaluating text-guided image editing (TIE) methods remains a challenging problem, as reliable assessment should simultaneously consider perceptual quality, alignment with textual instructions, and preservation of origina…

Image Editing

MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks

2025-09-18 · Mingsong Li, Lin Liu, Hongjun Wang, Haoxing Chen 외 arxiv

Current instruction-based image editing (IBIE) methods struggle with challenging editing tasks, as both editing types and sample counts of existing datasets are limited. Moreover, traditional dataset construction often c…

Style TransferImage Editing

OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision

2024-11-11 · Cong Wei, Zheyang Xiong, Weiming Ren, Xinrun Du 외

Instruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or manually annotated image editing pairs. However, these methods remain far from…

FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing

2024-07-25 · Gwanhyeong Koo, Sunjae Yoon, Ji Woo Hong, Chang D. Yoo

Current image editing methods primarily utilize DDIM Inversion, employing a two-branch diffusion approach to preserve the attributes and layout of the original image. However, these methods encounter challenges with non-…

Text-based Image Editing