paper-with-me

홈 › Papers

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

2025-06-15 · Chenglin Wang, Yucheng Zhou, Qianning Wang, Zhe Wang, Kai Zhang

Text-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ``chain'' instructions where operations are interdependent. Current models struggle with these intricate directives, and existing benchmarks inadequately evaluate such capabilities. Specifically, they often overlook multi-instruction and chain-instruction complexities, and common consistency metrics are flawed. To address this, we introduce ComplexBench-Edit, a novel benchmark designed to systematically assess model performance on complex, multi-instruction, and chain-dependent image editing tasks. ComplexBench-Edit also features a new vision consistency evaluation method that accurately assesses non-modified regions by excluding edited areas. Furthermore, we propose a simple yet powerful Chain-of-Thought (CoT)-based approach that significantly enhances the ability of existing models to follow complex instructions. Our extensive experiments demonstrate ComplexBench-Edit's efficacy in differentiating model capabilities and highlight the superior performance of our CoT-based method in handling complex edits. The data and code are released at https://github.com/llllly26/ComplexBench-Edit.

📄 PDF Abstract BibTeX arXiv:2506.12830

Code (1)

llllly26/complexbench-edit 공식 구현 pytorch

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Benchmarking Complex Instruction-Following with Multiple Constraints Composition

2024-07-04 · Bosi Wen, Pei Ke, Xiaotao Gu, Lindong Wu 외

Instruction following is one of the fundamental capabilities of large language models (LLMs). As the ability of LLMs is constantly improving, they have been increasingly applied to deal with complex human instructions in…

BenchmarkingInstruction Following

CompBench: Benchmarking Complex Instruction-guided Image Editing

2025-05-18 · Bohan Jia, Wenxuan Huang, Yuntian Tang, Junbo Qiao 외

While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. T…

BenchmarkingInstruction Following

MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance

2026-02-08 · Xuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan 외 arxiv

Recent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-world applications that require complex …

Image Editing

MIGE: A Unified Framework for Multimodal Instruction-Based Image Generation and Editing

2025-02-28 · Xueyun Tian, Wei Li, Bingbing Xu, Yige Yuan 외

Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high…

Image GenerationTransfer Learning

An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing

2025-08-24 · Zihan Liang, Jiahao Sun, Haoran Ma arxiv

Despite the remarkable capabilities of text-to-image (T2I) generation models, real-world applications often demand fine-grained, iterative image editing that existing methods struggle to provide. Key challenges include g…

Scene UnderstandingImage Editing