paper-with-me

Papers

RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions

2025-06-03 · Bimsara Pathiraja, Maitreya Patel, Shivam Singh, Yezhou Yang, Chitta Baral

Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility.

📄 PDF Abstract BibTeX arXiv:2506.03448

Code (0)

등록된 구현이 없습니다.

Tasks

Referring ExpressionSynthetic Data Generation

Similar Papers 제목 키워드 기반

MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing

2026-04-06 · Ziqian Liu, Stephan Alaniz arxiv

Instruction-guided image editing has seen remarkable progress with models like FLUX.2 and Qwen-Image-Edit, yet they still struggle with complex scenarios with multiple similar instances each requiring individual edits. W…

Image Editing

UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

2025-04-22 · Tsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu 외

Text-to-Image (T2I) diffusion models have shown impressive results in generating visually compelling images following user prompts. Building on this, various methods further fine-tune the pre-trained T2I model for specif…

Depth EstimationImage Generation

MIRA: Multimodal Iterative Reasoning Agent for Image Editing

2025-11-26 · Ziyun Zeng, Hang Hua, Jiebo Luo arxiv

Instruction-guided image editing offers an intuitive way for users to edit images with natural language. However, diffusion-based editing models often struggle to accurately interpret complex user instructions, especiall…

Multimodal ReasoningImage Editing

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation

2026-05-13 · Jingxuan He, Xiyu Wang, Yunke Wang, Mengyu Zheng 외 arxiv

Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which implicitly requires identifying where an edi…

Image SegmentationImage Editing

InstructBrush: Learning Attention-based Instruction Optimization for Image Editing

2024-03-27 · Ruoyu Zhao, Qingnan Fan, Fei Kou, Shuai Qin 외

In recent years, instruction-based image editing methods have garnered significant attention in image editing. However, despite encompassing a wide range of editing priors, these methods are helpless when handling editin…