paper-with-me

홈 › Papers

CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning

2026-01-31 · Minheng Ni, Yutao Fan, Zhengyuan Yang, Yeli Shen, Yuxiang Wei, Yaowen Zhang, Lijuan Wang, Lei Zhang, Wangmeng Zuo arxiv

Recent advances in large multimodal models (LMMs) have enabled instruction-based image editing, allowing users to modify visual content via natural language descriptions. However, existing approaches often struggle with high-level semantic reasoning and visual consistency, particularly under ambiguous or complex instructions. To address these challenges, we propose CoEditor++, a cognitively structured, training-free framework that decomposes editing into "what to edit" and "how to edit" through two cognitive stages with a reflective self-selection mechanism, enabling robust, fine-grained, and interpretable editing. Built entirely from open-sourced components, CoEditor++ requires no additional training or fine-tuning, ensuring transparency and cross-domain applicability. We evaluate CoEditor++ on SmartEdit, a widely used benchmark for general editing, and AltBear, a privacy and compliance-oriented benchmark. Experimental results show that CoEditor++ achieves state-of-the-art performance in both general editing and responsible editing tasks compared with open-sourced models that require training on specialized editing datasets maintaining significantly higher visual consistency. When compared with closed-source models such as Nano Banana Pro or GPT-4o, CoEditor++ preserves comparable instruction following while still substantially outperforming them in visual consistency. Extensive ablation studies confirm that the effectiveness of CoEditor++ benefits from its structured cognitive design rather than any specific model component. Our findings suggest the potential toward cognitive-centric instruction-based image editing.

📄 PDF Abstract BibTeX arXiv:2603.05518

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingImage Editing

Similar Papers 제목 키워드 기반

Responsible Visual Editing

2024-04-08 · Minheng Ni, Yeli Shen, Lei Zhang, WangMeng Zuo

With recent advancements in visual synthesis, there is a growing risk of encountering images with detrimental effects, such as hate, discrimination, or privacy violations. The research on transforming harmful images into…

Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing

2023-05-29 · Jiayi Wei, Greg Durrett, Isil Dillig

Developers often dedicate significant time to maintaining and refactoring existing code. However, most prior work on generative models for code focuses solely on creating new code, overlooking the distinctive needs of ed…

Code CompletionEDIT TaskLanguage Modelling

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

2025-05-22 · Yongliang Wu, Zonghui Li, Xinting Hu, Xinyu Ye 외

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based re…

BenchmarkingDiagnostic

Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

2025-07-02 · Qingdong He, Xueqin Chen, Chaoyi Wang, Yanjie Pan 외 arxiv

Instruction-based image editing (IIE) has advanced rapidly with the success of diffusion models. However, existing efforts primarily focus on simple and explicit instructions to execute editing operations such as adding,…

Zero-shot GeneralizationVisual ReasoningImage Editing

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

2025-12-05 · Hongyu Li, Manyuan Zhang, Dian Zheng, Ziyu Guo 외 arxiv

Instruction-based image editing has emerged as a prominent research area, which, benefiting from image generation foundation models, have achieved high aesthetic quality, making instruction-following capability the prima…

Reinforcement LearningImage GenerationImage Editing