paper-with-me

홈 › Papers

Multi-Reward as Condition for Instruction-based Image Editing

2024-11-06 · Xin Gu, Ming Li, Libo Zhang, Fan Chen, Longyin Wen, Tiejian Luo, Sijie Zhu

High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable Diffusion, DALL-E) which are not trained for image editing. Accordingly, these datasets suffer from inaccurate instruction following, poor detail preserving, and generation artifacts. In this paper, we propose to address the training data quality issue with multi-perspective reward data instead of refining the ground-truth image quality. 1) we first design a quantitative metric system based on best-in-class LVLM (Large Vision Language Model), i.e., GPT-4o in our case, to evaluate the generation quality from 3 perspectives, namely, instruction following, detail preserving, and generation quality. For each perspective, we collected quantitative score in $0\sim 5$ and text descriptive feedback on the specific failure points in ground-truth edited images, resulting in a high-quality editing reward dataset, i.e., RewardEdit20K. 2) We further proposed a novel training framework to seamlessly integrate the metric output, regarded as multi-reward, into editing models to learn from the imperfect training triplets. During training, the reward scores and text descriptions are encoded as embeddings and fed into both the latent space and the U-Net of the editing models as auxiliary conditions. During inference, we set these additional conditions to the highest score with no text description for failure points, to aim at the best generation outcome. Experiments indicate that our multi-reward conditioned model outperforms its no-reward counterpart on two popular editing pipelines, i.e., InsPix2Pix and SmartEdit. The code and dataset will be released.

📄 PDF Abstract BibTeX arXiv:2411.04713

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveInstruction Following

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
U-Net 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

2026-02-19 · Qiucheng Wu, Jing Shi, Simon Jenni, Kushal Kafle 외 arxiv

Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promisin…

Reinforcement LearningMultimodal ReasoningImage Editing

HIVE: Harnessing Human Feedback for Instructional Visual Editing

2023-03-16 · CVPR 2024 1 · Shu Zhang, Xinyi Yang, Yihao Feng, Can Qin 외

Incorporating human feedback has been shown to be crucial to align text generated by large language models to human preferences. We hypothesize that state-of-the-art instructional image editing models, where outputs are …

Text-based Image Editing

Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing

2026-04-10 · Zhuohan Ouyang, Zhe Qian, Wenhuo Cui, Chaoqun Wang arxiv

Instruction-guided image editing requires balancing target modification with non-target preservation. Recently, flow-based models have emerged as a strong and increasingly adopted backbone for instruction-guided image ed…

Instruction FollowingImage Editing

EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling

2025-09-28 · Xin Luo, Jiahao Wang, Chenyuan Wu, Shitao Xiao 외 arxiv

Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce a desired result. Reinforcement Learning…

Reinforcement LearningImage Editing

EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing

2025-09-30 · Keming Wu, Sicong Jiang, Max Ku, Ping Nie 외 arxiv

Recently, we have witnessed great progress in image editing with natural language instructions. Several closed-source models like GPT-Image-1, Seedream, and Google-Nano-Banana have shown highly promising progress. Howeve…

Reinforcement LearningImage Editing