paper-with-me

홈 › Papers

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

2026-06-01 · Jiahui Huang, Yasi Zhang, Tianyu Chen, Shu Wang, Jianwen Xie, Oscar Leong, Mingyuan Zhou, Nanzhu Wang, Ying Nian Wu arxiv

Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing--the natural interactive setting where a user iteratively refines an image based on the model's own previous outputs. This failure stems from the all-or-nothing requirement, where a single failed turn compromises the entire sequence, and error propagation, where exposure bias leads to compounding editing errors. To address these challenges, we introduce MT-EditFlow, a flow-matching reinforcement learning framework designed to optimize reward signals for sequential image editing. MT-EditFlow integrates a multi-turn perspective with a multi-reward formulation to provide a unified structure applicable to both GRPO and NFT-based reinforcement learning methods. We systematically analyze and optimize the reward signal by investigating effective scoring strategies for turn-level aggregation, VLM reasoning modes to trade off reward bias and variance, and advantage fusion levels to prevent reward hacking. Our findings reveal that broadcasting the aggregated advantage across the entire editing trajectory effectively bridges the gap between local planning and global multi-turn task success. Extensive experiments demonstrate that MT-EditFlow significantly improves performance across diverse base models. Notably, it boosts FLUX.1-Kontext-dev by 6.85 points in turn-3 overall performance, surpassing state-of-the-art open-source models such as Qwen-Image-Edit. By maintaining high marginal success rates and reducing exposure bias, MT-EditFlow provides a foundation for more reliable and natural human-AI collaboration in visual content creation.

📄 PDF Abstract BibTeX arXiv:2606.01985

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Editing

Similar Papers 제목 키워드 기반

EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

2026-08-04 · Rui Nie, Chuang Wang, Haitao Zhou, Jiahe Song 외 arxiv

Controllable local editing of 3D assets requires precise target localization and appropriate visual guidance. However, existing methods lack a simple yet accurate way to obtain 3D masks and struggle to achieve the desire…

Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing

2026-06-04 · Yuxiao Ye, Haoran He, Fangyuan Kong, Xintao Wang 외 arxiv

Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain confined to single-turn settings, overlooking the more realistic scenar…

Reinforcement LearningInstruction FollowingImage GenerationImage Editing

Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing

2025-08-09 · Shichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang 외 arxiv

Text-to-image generation tasks have driven remarkable advances in diverse media applications, yet most focus on single-turn scenarios and struggle with iterative, multi-turn creative tasks. Recent dialogue-based systems …

Text-to-Image GenerationImage Editing

CHATEDIT: Towards Multi-turn Interactive Facial Image Editing via Dialogue

2023-03-20 · Xing Cui, Zekun Li, Peipei Li, Yibo Hu 외

This paper explores interactive facial image editing via dialogue and introduces the ChatEdit benchmark dataset for evaluating image editing and conversation abilities in this context. ChatEdit is constructed from the Ce…

AttributeFacial EditingResponse Generation

ImgEdit: A Unified Image Editing Dataset and Benchmark

2025-05-26 · Yang Ye, Xianyi He, Zongjian Li, Bin Lin 외

Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quali…

Image EditingImage GenerationLanguage Modeling+3