paper-with-me

홈 › Papers

Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions

2025-05-25 · Chenrui Ma, Xi Xiao, Tianyang Wang, Yanning Shen

Current text-driven image editing methods typically follow one of two directions: relying on large-scale, high-quality editing pair datasets to improve editing precision and diversity, or exploring alternative dataset-free techniques. However, constructing large-scale editing datasets requires carefully designed pipelines, is time-consuming, and often results in unrealistic samples or unwanted artifacts. Meanwhile, dataset-free methods may suffer from limited instruction comprehension and restricted editing capabilities. Faced with these challenges, the present work develops a novel paradigm for instruction-driven image editing that leverages widely available and enormous text-image pairs, instead of relying on editing pair datasets. Our approach introduces a multi-scale learnable region to localize and guide the editing process. By treating the alignment between images and their textual descriptions as supervision and learning to generate task-specific editing regions, our method achieves high-fidelity, precise, and instruction-consistent image editing. Extensive experiments demonstrate that the proposed approach attains state-of-the-art performance across various tasks and benchmarks, while exhibiting strong adaptability to various types of generative models.

📄 PDF Abstract BibTeX arXiv:2505.19352

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

POEM: Precise Object-level Editing via MLLM control

2025-04-10 · Marco Schouten, Mehmet Onurcan Kaya, Serge Belongie, Dim P. Papadopoulos

Diffusion models have significantly improved text-to-image generation, producing high-quality, realistic images from textual descriptions. Beyond generation, object-level image editing remains a challenging problem, requ…

Image GenerationObjectObject LocalizationText-based Image Editing+2

InstructionBench: An Instructional Video Understanding Benchmark

2025-04-07 · Haiwan Wei, Yitian Yuan, Xiaohan Lan, Wei Ke 외

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce Inst…

Common Sense ReasoningMultiple-choiceVideo Understanding

SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

2024-05-07 · Yuying Ge, Sijie Zhao, Chen Li, Yixiao Ge 외

In this technical report, we introduce SEED-Data-Edit: a unique hybrid dataset for instruction-guided image editing, which aims to facilitate image manipulation using open-form language. SEED-Data-Edit is composed of thr…

Image ManipulationLanguage ModelingLanguage ModellingLarge Language Model+1

CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

2026-05-30 · Fangzhou Lin, Peiran Li, Lingyu Xu, Wenjing Chen 외 arxiv

Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture the diversity of real-image tasks in profe…

Image Editing

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

2026-08-26 · Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun 외 arxiv

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by au…

Image Editing