paper-with-me

Papers

CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing

2025-08-15 · Zhe Zhu, Honghua Chen, Peng Li, Mingqiang Wei arxiv

Text-driven 3D editing seeks to modify 3D scenes according to textual descriptions, and most existing approaches tackle this by adapting pre-trained 2D image editors to multi-view inputs. However, without explicit control over multi-view information exchange, they often fail to maintain cross-view consistency, leading to insufficient edits and blurry details. We introduce CoreEditor, a novel framework for consistent text-to-3D editing. The key innovation is a correspondence-constrained attention mechanism that enforces precise interactions between pixels expected to remain consistent throughout the diffusion denoising process. Beyond relying solely on geometric alignment, we further incorporate semantic similarity estimated during denoising, enabling more reliable correspondence modeling and robust multi-view editing. In addition, we design a selective editing pipeline that allows users to choose preferred results from multiple candidates, offering greater flexibility and user control. Extensive experiments show that CoreEditor produces high-quality, 3D-consistent edits with sharper details, significantly outperforming prior methods.

📄 PDF Abstract BibTeX arXiv:2508.11603

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing

2024-06-13 · Jiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao 외

Video editing is an emerging task, in which most current methods adopt the pre-trained text-to-image (T2I) diffusion model to edit the source video in a zero-shot manner. Despite extensive efforts, maintaining the tempor…

DenoisingGPUVideo Editing

Edicho: Consistent Image Editing in the Wild

2024-12-30 · Qingyan Bai, Hao Ouyang, Yinghao Xu, Qiuyu Wang 외

As a verified need, consistent editing across in-the-wild images remains a technical challenge arising from various unmanageable factors, like object poses, lighting conditions, and photography environments. Edicho steps…

Denoising

Efficient-NeRF2NeRF: Streamlining Text-Driven 3D Editing with Multiview Correspondence-Enhanced Diffusion Models

2023-12-13 · Liangchen Song, Liangliang Cao, Jiatao Gu, Yifan Jiang 외

The advancement of text-driven 3D content editing has been blessed by the progress from 2D generative diffusion models. However, a major obstacle hindering the widespread adoption of 3D content editing is its time-intens…

GPU

TokenFlow: Consistent Diffusion Features for Consistent Video Editing

2023-07-19 · Michal Geyer, Omer Bar-Tal, Shai Bagon, Tali Dekel

The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual quality and user control over the generated conte…

Video Editing

SyncNoise: Geometrically Consistent Noise Prediction for Text-based 3D Scene Editing

2024-06-25 · Ruihuang Li, Liyi Chen, Zhengqiang Zhang, Varun Jampani 외

Text-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achie…

3D scene EditingImage Generation