paper-with-me

홈 › Papers

EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing

2025-09-30 · Keming Wu, Sicong Jiang, Max Ku, Ping Nie, Minghao Liu, Wenhu Chen arxiv

Recently, we have witnessed great progress in image editing with natural language instructions. Several closed-source models like GPT-Image-1, Seedream, and Google-Nano-Banana have shown highly promising progress. However, the open-source models are still lagging. The main bottleneck is the lack of a reliable reward model to scale up high-quality synthetic training data. To address this critical bottleneck, we built EditReward, trained with our new large-scale human preference dataset, meticulously annotated by trained experts following a rigorous protocol containing over 200K preference pairs. EditReward demonstrates superior alignment with human preferences in instruction-guided image editing tasks. Experiments show that EditReward achieves state-of-the-art human correlation on established benchmarks such as GenAI-Bench, AURORA-Bench, ImagenHub, and our new EditReward-Bench, outperforming a wide range of VLM-as-judge models. Furthermore, we use EditReward to select a high-quality subset from the existing noisy ShareGPT-4o-Image dataset. We train Step1X-Edit on the selected subset, which shows significant improvement over training on the full set. This demonstrates EditReward's ability to serve as a reward model to scale up high-quality training data for image editing. Furthermore, its strong alignment suggests potential for advanced applications like reinforcement learning-based post-training and test-time scaling of image editing models. EditReward with its training dataset will be released to help the community build more high-quality image editing training datasets.

📄 PDF Abstract BibTeX arXiv:2509.26346

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Editing

Similar Papers 제목 키워드 기반

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

2026-05-13 · Xuehai Bai, Yang Shi, Yi-Fan Zhang, Xuanyu Zhu 외 arxiv

Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, …

Instruction FollowingVisual ReasoningImage Editing

RewardHarness: Self-Evolving Agentic Post-Training

2026-05-09 · Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei 외 arxiv

Evaluating instruction-guided image edits requires rewards that reflect subtle human preferences, yet current reward models typically depend on large-scale preference annotation and additional model training. This create…

EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling

2025-09-28 · Xin Luo, Jiahao Wang, Chenyuan Wu, Shitao Xiao 외 arxiv

Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce a desired result. Reinforcement Learning…

Reinforcement LearningImage Editing

BLEUBERI: BLEU is a surprisingly effective reward for instruction following

2025-05-16 · Yapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh 외

Reward models are central to aligning LLMs with human preferences, but they are costly to train, requiring large-scale human-labeled preference data and powerful pretrained LLM backbones. Meanwhile, the increasing availa…

Instruction FollowingSynthetic Data Generation

SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning

2026-02-07 · Yancheng Long, Yankai Yang, Hongyang Wei, Wei Chen 외 arxiv

Online Reinforcement Learning (RL) offers a promising avenue for complex image editing but is currently constrained by the scarcity of reliable and fine-grained reward signals. Existing evaluators frequently struggle wit…

Reinforcement LearningSpatial ReasoningImage Editing