paper-with-me

홈 › Papers

EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling

2025-09-28 · Xin Luo, Jiahao Wang, Chenyuan Wu, Shitao Xiao, Xiyan Jiang, Defu Lian, Jiajun Zhang, Dong Liu, Zheng liu arxiv

Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce a desired result. Reinforcement Learning (RL) offers a promising solution, but its adoption in image editing has been severely hindered by the lack of a high-fidelity, efficient reward signal. In this work, we present a comprehensive methodology to overcome this barrier, centered on the development of a state-of-the-art, specialized reward model. We first introduce EditReward-Bench, a comprehensive benchmark to systematically evaluate reward models on editing quality. Building on this benchmark, we develop EditScore, a series of reward models (7B-72B) for evaluating the quality of instruction-guided image editing. Through meticulous data curation and filtering, EditScore effectively matches the performance of learning proprietary VLMs. Furthermore, coupled with an effective self-ensemble strategy tailored for the generative nature of EditScore, our largest variant even surpasses GPT-5 in the benchmark. We then demonstrate that a high-fidelity reward model is the key to unlocking online RL for image editing. Our experiments show that, while even the largest open-source VLMs fail to provide an effective learning signal, EditScore enables efficient and robust policy optimization. Applying our framework to a strong base model, OmniGen2, results in a final model that shows a substantial and consistent performance uplift. Overall, this work provides the first systematic path from benchmarking to reward modeling to RL training in image editing, showing that a high-fidelity, domain-specialized reward model is the key to unlocking the full potential of RL in this domain.

📄 PDF Abstract BibTeX arXiv:2509.23909

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Editing

Similar Papers 제목 키워드 기반

EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

2026-04-09 · Xiangyuan Wang, Honghao Cai, Yunhao Bai, Chao Hui 외 arxiv

High-quality source-target image pairs with precise editing instructions are essential for instruction-guided image editing, yet constructing such training triplets at scale remains costly. Recent pipelines often rely on…

Image Editing

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

2026-02-07 · Yancheng Long, Yankai Yang, Hongyang Wei, Wei Chen 외 arxiv

Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and fine-grained reward signals. Existing evaluators frequently struggle wit…

Reinforcement LearningSpatial ReasoningImage Editing

Video Editing via Factorized Diffusion Distillation

2024-03-14 · Uriel Singer, Amit Zohar, Yuval Kirstain, Shelly Sheynin 외

We introduce Emu Video Edit (EVE), a model that establishes a new state-of-the art in video editing without relying on any supervised video editing data. To develop EVE we separately train an image editing adapter and a …

Video EditingVideo Generation

VINCIE: Unlocking In-context Image Editing from Video

2025-06-12 · Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao 외

In-context image editing aims to modify images based on a contextual sequence comprising text and previously generated images. Existing methods typically depend on task-specific pipelines and expert models (e.g., segment…

PredictionSegmentationStory Generation

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

2025-11-18 · Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu 외 arxiv

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipel…

Reinforcement LearningImage GenerationImage Editing