paper-with-me

홈 › Papers

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

2025-10-19 · Zongjian Li, Zheyuan Liu, Qihui Zhang, Bin Lin, Feize Wu, Shenghai Yuan, Zhiyuan Yan, Yang Ye, Wangbo Yu, Yuwei Niu, Shaodong Wang, Xinhua Cheng, Li Yuan arxiv

Instruction-based image editing has achieved remarkable progress; however, models solely trained via supervised fine-tuning often overfit to annotated patterns, hindering their ability to explore and generalize beyond training distributions. To this end, we introduce Edit-R1, a novel post-training framework for instruction-based image editing based on policy optimization. Specifically, we utilize Diffusion Negative-aware Finetuning (DiffusionNFT), a likelihood-free policy optimization method consistent with the flow matching forward process, thereby enabling the use of higher-order samplers and more efficient training. Another key challenge here is the absence of a universal reward model, resulting from the diverse nature of editing instructions and tasks. To bridge this gap, we employ a Multimodal Large Language Model (MLLM) as a unified, training-free reward model, leveraging its output logits to provide fine-grained feedback. Furthermore, we carefully design a low-variance group filtering mechanism to reduce MLLM scoring noise and stabilize optimization. \texttt{UniWorld-V2}, trained with this framework, achieves \textbf{state-of-the-art} results on the ImgEdit and GEdit-Bench benchmarks, scoring 4.49 and 7.83, respectively. Crucially, our framework is model-agnostic, delivering substantial performance gains when applied to diverse base models like Qwen-Image-Edit and FLUX-Kontext, demonstrating its wide applicability. Code and models are publicly available to support further research.

📄 PDF Abstract BibTeX arXiv:2510.16888

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

UniWorld-Design: From Pixel Generation to Layer-Native Design

2026-08-04 · Zongjian Li, Zhiyuan Yan, Chenxu Bai, Chen Chen 외 hf

We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editin…

Image Generation

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

2026-08-05 · Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou 외 hf

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance use…

Novel View Synthesis

UniWorld: Autonomous Driving Pre-training via World Models

2023-08-14 · Chen Min, Dawei Zhao, Liang Xiao, Yiming Nie 외

In this paper, we draw inspiration from Alberto Elfes' pioneering work in 1989, where he introduced the concept of the occupancy grid as World Models for robots. We imbue the robot with a spatial-temporal world model, te…

3D Object DetectionAutonomous Drivingmotion predictionobject-detection+1

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework

2026-03-21 · Guanzhou Chen, Erfei Cui, Changyao Tian, Danni Yang 외 arxiv

Instruction-based image editing has emerged as a key capability for unified multimodal models (UMMs), yet constructing large-scale, diverse, and high-quality editing datasets without costly proprietary APIs remains chall…

Image Editing

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

2025-06-03 · Bin Lin, Zongjian Li, Xinhua Cheng, Yuwei Niu 외

Although existing unified models achieve strong performance in vision-language understanding and text-to-image generation, they remain limited in addressing image perception and manipulation -- capabilities increasingly …

Image EditingImage GenerationImage Manipulation+2