paper-with-me

Papers

Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

2025-06-25 · Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java, Silky Singh, Tarun Ram Menta, Surgan Jandial, Balaji Krishnamurthy

Text-to-Image Diffusion models have enabled a wide array of image editing applications. However, capturing all types of edits through text alone can be challenging and cumbersome. The ambiguous nature of certain image edits is better expressed through an exemplar pair, i.e., a pair of images depicting an image before and after an edit respectively. In this work, we tackle exemplar-based image editing -- the task of transferring an edit from an exemplar pair to a content image(s), by leveraging pretrained text-to-image diffusion models and multimodal VLMs. Even though our end-to-end pipeline is optimization-free, our experiments demonstrate that it still outperforms baselines on multiple types of edits while being ~4x faster.

📄 PDF Abstract BibTeX arXiv:2506.20155

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

StyleBooth: Image Style Editing with Multimodal Instruction

2024-04-18 · Zhen Han, Chaojie Mao, Zeyinzi Jiang, Yulin Pan 외

Given an original image, image editing aims to generate an image that align with the provided instruction. The challenges are to accept multimodal inputs as instructions and a scarcity of high-quality training data, incl…

ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models

2024-11-06 · Ashutosh Srivastava, Tarun Ram Menta, Abhinav Java, Avadhoot Jadhav 외

Modern Text-to-Image (T2I) Diffusion models have revolutionized image editing by enabling the generation of high-quality photorealistic images. While the de facto method for performing edits with T2I models is through te…

PIXELS: Progressive Image Xemplar-based Editing with Latent Surgery

2025-01-16 · Shristi Das Biswas, Matthew Shreve, Xuelu Li, Prateek Singhal 외

Recent advancements in language-guided diffusion models for image editing are often bottle-necked by cumbersome prompt engineering to precisely articulate desired changes. An intuitive alternative calls on guidance from …

Image GenerationPrompt Engineering

Paint by Example: Exemplar-based Image Editing with Diffusion Models

2022-11-23 · CVPR 2023 1 · Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang 외

Language-guided image editing has achieved great success recently. In this paper, for the first time, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervi…

Image GenerationImage Manipulation

FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing

2024-12-26 · Wanglong Lu, Jikai Wang, Xiaogang Jin, Xianta Jiang 외

Existing facial editing methods have achieved remarkable results, yet they often fall short in supporting multimodal conditional local facial editing. One of the significant evidences is that their output image quality d…

AttributeFacial Editing