paper-with-me

Papers

Referring Image Editing: Object-level Image Editing via Referring Expressions

2024-01-01 · CVPR 2024 1 · Chang Liu, Xiangtai Li, Henghui Ding

Significant advancements have been made in image editing with the recent advance of the Diffusion model. However most of the current methods primarily focus on global or subject-level modifications and often face limitations when it comes to editing specific objects when there are other objects coexisting in the scene given solely textual prompts. In response to this challenge we introduce an object-level generative task called Referring Image Editing (RIE) which enables the identification and editing of specific source objects in an image using text prompts. To tackle this task effectively we propose a tailored framework called ReferDiffusion. It aims to disentangle input prompts into multiple embeddings and employs a mixed-supervised multi-stage training strategy. To facilitate further research in this domain we introduce the RefCOCO-Edit dataset comprising images editing prompts source object segmentation masks and reference edited images for training and evaluation. Our extensive experiments demonstrate the effectiveness of our approach in identifying and editing target objects while conventional general image editing and region-based image editing methods have difficulties in this challenging task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Segmentation

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation

2026-05-13 · Jingxuan He, Xiyu Wang, Yunke Wang, Mengyu Zheng 외 arxiv

Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which implicitly requires identifying where an edi…

Image SegmentationImage Editing

RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions

2025-06-03 · Bimsara Pathiraja, Maitreya Patel, Shivam Singh, Yezhou Yang 외

Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing mult…

Referring ExpressionSynthetic Data Generation

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

2026-07-20 · Dingyun Zhang, Lixue Gong, Wei Liu hf

In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting vide…

Referring Expression SegmentationImage Editing

FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension

2024-09-23 · Junzhuo Liu, Xuzheng Yang, Weiwei Li, Peng Wang

Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. Consequently, it serves …

Image ComprehensionReferring ExpressionReferring Expression ComprehensionVisual Reasoning

Referring Layer Decomposition

2026-02-22 · Fangyi Chen, Yaojie Shen, Lu Xu, Ye Yuan 외 arxiv

Precise, object-aware control over visual content is essential for advanced image editing and compositional generation. Yet, most existing approaches operate on entire images holistically, limiting the ability to isolate…

Zero-shot GeneralizationImage Editing