paper-with-me

홈 › Papers

GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing

2026-03-22 · Zifeng Zhu, Jiaming Han, Jiaxiang Zhao, Minnan Luo, Xiangyu Yue arxiv

While Diffusion Large Language Models (DLLMs) have demonstrated remarkable capabilities in multi-modal generation, performing precise, training-free image editing remains an open challenge. Unlike continuous diffusion models, the discrete tokenization inherent in DLLMs hinders the application of standard noise inversion techniques, often leading to structural degradation during editing. In this paper, we introduce GIDE (Grounded Inversion for DLLM Image Editing), a unified framework designed to bridge this gap. GIDE incorporates a novel Discrete Noise Inversion mechanism that accurately captures latent noise patterns within the discrete token space, ensuring high-fidelity reconstruction. We then decompose the editing pipeline into grounding, inversion, and refinement stages. This design enables GIDE supporting various editing instructions (text, point and box) and operations while strictly preserving the unedited background. Furthermore, to overcome the limitations of existing single-step evaluation protocols, we introduce GIDE-Bench, a rigorous benchmark comprising 805 compositional editing scenarios guided by diverse multi-modal inputs. Extensive experiments on GIDE-Bench demonstrate that GIDE significantly outperforms prior training-free methods, improving Semantic Correctness by 51.83% and Perceptual Quality by 50.39%. Additional evaluations on ImgEdit-Bench confirm its broad applicability, demonstrating consistent gains over trained baselines and yielding photorealistic consistency on par with leading models.

📄 PDF Abstract BibTeX arXiv:2603.21176

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

Achieving Scalable Robot Autonomy via neurosymbolic planning using lightweight local LLM

2025-05-13 · Nicholas Attolino, Alessio Capitanelli, Fulvio Mastrogiovanni

PDDL-based symbolic task planning remains pivotal for robot autonomy yet struggles with dynamic human-robot collaboration due to scalability, re-planning demands, and delayed plan availability. Although a few neurosymbol…

16k8kTask Planning

AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation

2026-08-21 · Xiangfei Sheng, Weidong Zou, Tianjiao Gu, Zhichao Yang 외 arxiv

Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. While existing AI-generated image (AGI) evaluation benchmarks have mad…

LogiDebrief: A Signal-Temporal Logic based Automated Debriefing Approach with Large Language Models Integration

2025-05-06 · Zirong Chen, Ziyan An, Jennifer Reynolds, Kristin Mullen 외

Emergency response services are critical to public safety, with 9-1-1 call-takers playing a key role in ensuring timely and effective emergency operations. To ensure call-taking performance consistency, quality assurance…

LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs

2025-06-17 · Xiaoran Liu, Zhigeng Liu, Zengfeng Huang, Qipeng Guo 외

Large Language Diffusion Models, or diffusion LLMs, have emerged as a significant focus in NLP research, with substantial effort directed toward understanding their scalability and downstream task performance. However, t…

ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement

2025-12-15 · Zhihang Liu, Xiaoyi Bao, Pandeng Li, Junjie Zhou 외 arxiv

While existing generation and unified models excel at general image generation, they struggle with tasks requiring deep reasoning, planning, and precise data-to-visual mapping abilities beyond general scenarios. To push …

Image Generation