paper-with-me

홈 › Papers

Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

2025-07-07 · Chun-Hsiao Yeh, Yilin Wang, Nanxuan Zhao, Richard Zhang, Yuheng Li, Yi Ma, Krishna Kumar Singh arxiv

Recent diffusion-based image editing methods have significantly advanced text-guided tasks but often struggle to interpret complex, indirect instructions. Moreover, current models frequently suffer from poor identity preservation, unintended edits, or rely heavily on manual masks. To address these challenges, we introduce X-Planner, a Multimodal Large Language Model (MLLM)-based planning system that effectively bridges user intent with editing model capabilities. X-Planner employs chain-of-thought reasoning to systematically decompose complex instructions into simpler, clear sub-instructions. For each sub-instruction, X-Planner automatically generates precise edit types and segmentation masks, eliminating manual intervention and ensuring localized, identity-preserving edits. Additionally, we propose a novel automated pipeline for generating large-scale data to train X-Planner which achieves state-of-the-art results on both existing benchmarks and our newly introduced complex editing benchmark.

📄 PDF Abstract BibTeX arXiv:2507.05259

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing

2025-12-18 · Tianyuan Qu, Lei Ke, Xiaohang Zhan, Longxiang Tang 외 arxiv

Instruction-based image editing enables natural-language control over visual modifications, yet existing models falter under Instruction-Visual Complexity (IV-Complexity), where intricate instructions meet cluttered or a…

Reinforcement LearningImage Editing

V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes

2025-03-13 · YanMing Zhang, Jun-Kun Chen, Jipeng Lyu, Yu-Xiong Wang

This paper introduces V$^2$Edit, a novel training-free framework for instruction-guided video and 3D scene editing. Addressing the critical challenge of balancing original content preservation with editing task fulfillme…

3D scene EditingDenoisingVideo Editing

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

2026-08-20 · Haoxiang Cao, Jiajiong Cao, Xuanpu Zhang, Changqian Yu 외 arxiv

Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with …

Reinforcement LearningImage Editing

Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling

2026-02-09 · Ruijie Ye, Jiayi Zhang, Zhuoxin Liu, Zihao Zhu 외 arxiv

We study instruction-based image editing under professional workflows and identify three persistent challenges: (i) editors often over-edit, modifying content beyond the user's intent; (ii) existing models are largely si…

Instruction FollowingImage Editing

From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing

2026-05-14 · Anirudh Sundara Rajan, Krishna Kumar Singh, Yong Jae Lee arxiv

Modern image editing models produce realistic results but struggle with abstract, multi step instructions (e.g., ``make this advertisement more vegetarian-friendly''). Prior agent based methods decompose such tasks but r…

Image Editing