paper-with-me

홈 › Papers

WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing

2026-07-23 · Ming Hu, Mingyu Dou, Jianfu Yin, Miaomiao Zhang, Cong Hu, Yao Wang, Bingliang Hu, Quan Wang arxiv

Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing one-step editing methods primarily rely on text conditioning for semantic transformation, lacking explicit spatial control over \textit{where} to edit. More importantly, even when spatial constraints are introduced, these methods often struggle to achieve strong and stable semantic modifications within the target regions. In this work, we revisit one-step image editing from a spatially controlled perspective and identify two key challenges: discovering editable regions and achieving effective localized semantic transformation. We reveal that existing methods perform global semantic transport, which limits high-intensity local editing under the one-step setting. To address this issue, we propose \textbf{WhereEdit}, a framework that reformulates one-step editing as localized adaptive editing. WhereEdit automatically identifies semantically relevant regions from internal model features and applies adaptive local modulation to enhance target-region editing while preserving non-target areas and structural consistency. Experiments on the PIE-Bench benchmark demonstrate that WhereEdit consistently outperforms existing one-step image editing methods, achieving superior editing quality while maintaining the efficiency of one-step generation. Additional experiments with region-level supervision further highlight the importance of explicit spatial reasoning for high-quality one-step image editing.

📄 PDF Abstract BibTeX arXiv:2607.20883

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 112

Tasks

Spatial ReasoningImage Editing

Similar Papers 제목 키워드 기반

FENeRF: Face Editing in Neural Radiance Fields

2021-11-30 · CVPR 2022 1 · Jingxiang Sun, Xuan Wang, Yong Zhang, Xiaoyu Li 외

Previous portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view c…

3D-Aware Image SynthesisImage Generation

LatentEditor: Text Driven Local Editing of 3D Scenes

2023-12-14 · Umar Khalid, Hasan Iqbal, Nazmul Karim, Jing Hua 외

While neural fields have made significant strides in view synthesis and scene reconstruction, editing them poses a formidable challenge due to their implicit encoding of geometry and texture information from multi-view i…

3D scene EditingDenoisingNeRF

DesignEdit: Multi-Layered Latent Decomposition and Fusion for Unified & Accurate Image Editing

2024-03-21 · Yueru Jia, Yuhui Yuan, Aosong Cheng, Chuke Wang 외

Recently, how to achieve precise image editing has attracted increasing attention, especially given the remarkable success of text-to-image generation models. To unify various spatial-aware image editing abilities into o…

Image Generationspatial-aware image editingText to Image GenerationText-to-Image Generation

MaTe3D: Mask-guided Text-based 3D-aware Portrait Editing

2023-12-12 · Kangneng Zhou, Daiheng Gao, Xuan Wang, Jie Zhang 외

3D-aware portrait editing has a wide range of applications in multiple fields. However, current approaches are limited due that they can only perform mask-guided or text-based editing. Even by fusing the two procedures i…

Blended Latent Diffusion under Attention Control for Real-World Video Editing

2024-09-05 · Deyin Liu, Lin Yuanbo Wu, Xianghua Xie

Due to lack of fully publicly available text-to-video models, current video editing methods tend to build on pre-trained text-to-image generation models, however, they still face grand challenges in dealing with the loca…

Image GenerationText to Image GenerationText-to-Image GenerationVideo Editing