paper-with-me

Papers

TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering

2025-12-18 · Rui Gui, Yang Wan, Haochen Han, Dongxing Mao, Fangming Liu, Min Li, Alex Jinpeng Wang arxiv

Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. However, text editing within images remains largely unexplored, as it requires generating legible characters while preserving semantic, geometric, and contextual coherence. To fill this gap, we introduce TextEditBench, a comprehensive evaluation benchmark that explicitly focuses on text-centric regions in images. Beyond basic pixel manipulations, our benchmark emphasizes reasoning-intensive editing scenarios that require models to understand physical plausibility, linguistic meaning, and cross-modal dependencies. We further propose a novel evaluation dimension, Semantic Expectation (SE), which measures reasoning ability of model to maintain semantic consistency, contextual coherence, and cross-modal alignment during text editing. Extensive experiments on state-of-the-art editing systems reveal that while current models can follow simple textual instructions, they still struggle with context-dependent reasoning, physical consistency, and layout-aware integration. By focusing evaluation on this long-overlooked yet fundamental capability, TextEditBench establishes a new testing ground for advancing text-guided image editing and reasoning in multimodal generation.

📄 PDF Abstract BibTeX arXiv:2512.16270

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generationImage Editing

Similar Papers 제목 키워드 기반

Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning

2025-12-10 · Xinyu Liu, Hangjie Yuan, Yujie Wei, Jiazheng Xing 외 arxiv

Unified video models exhibit strong capabilities in understanding and generation, yet they struggle with reason-informed visual editing even when equipped with powerful internal vision-language models (VLMs). We attribut…

Video Generation

Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

2025-07-02 · Qingdong He, Xueqin Chen, Chaoyi Wang, Yanjie Pan 외 arxiv

Instruction-based image editing (IIE) has advanced rapidly with the success of diffusion models. However, existing efforts primarily focus on simple and explicit instructions to execute editing operations such as adding,…

Zero-shot GeneralizationVisual ReasoningImage Editing

MultiMedEdit: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA

2025-08-09 · Shengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan 외 arxiv

Knowledge editing (KE) provides a scalable approach for updating factual knowledge in large language models without full retraining. While previous studies have demonstrated effectiveness in general domains and medical Q…

knowledge editingVisual Reasoning

Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

2025-04-03 · Xiangyu Zhao, Peiyuan Zhang, Kexian Tang, Hao Li 외

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preservi…

BenchmarkingLogical Reasoning

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

2026-04-16 · Yixuan Ding, Wei Huang, Ruijie Quan, Xiaojuan Qi 외 arxiv

Diffusion-based image editing has achieved strong visual fidelity under natural language instructions, yet most existing systems still operate at the level of surface instruction following, without reasoning about the im…

Instruction FollowingImage Editing