paper-with-me

홈 › Papers

CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion

2024-12-02 · CVPR 2025 1 · Kai He, Chin-Hsuan Wu, Igor Gilitschenski

Recent advances in 3D representations, such as Neural Radiance Fields and 3D Gaussian Splatting, have greatly improved realistic scene modeling and novel-view synthesis. However, achieving controllable and consistent editing in dynamic 3D scenes remains a significant challenge. Previous work is largely constrained by its editing backbones, resulting in inconsistent edits and limited controllability. In our work, we introduce a novel framework that first fine-tunes the InstructPix2Pix model, followed by a two-stage optimization of the scene based on deformable 3D Gaussians. Our fine-tuning enables the model to "learn" the editing ability from a single edited reference image, transforming the complex task of dynamic scene editing into a simple 2D image editing process. By directly learning editing regions and styles from the reference, our approach enables consistent and precise local edits without the need for tracking desired editing regions, effectively addressing key challenges in dynamic scene editing. Then, our two-stage optimization progressively edits the trained dynamic scene, using a designed edited image buffer to accelerate convergence and improve temporal consistency. Compared to state-of-the-art methods, our approach offers more flexible and controllable local scene editing, achieving high-quality and consistent results.

📄 PDF Abstract BibTeX arXiv:2412.01792

Code (0)

등록된 구현이 없습니다.

Tasks

3D scene EditingNovel View Synthesis

Similar Papers 제목 키워드 기반

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

2025-12-19 · Yun He, Francesco Pittaluga, Ziyu Jiang, Matthias Zwicker 외 arxiv

LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represents each video as an explicit 3D scene graph, decomposing the scene into…

Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints

2023-10-05 · Chuan Fang, Yuan Dong, Kunming Luo, Xiaotao Hu 외

Text-driven 3D indoor scene generation is useful for gaming, the film industry, and AR/VR applications. However, existing methods cannot faithfully capture the room layout, nor do they allow flexible editing of individua…

Layout GenerationScene GenerationText to 3D

Tuning-Free Visual Customization via View Iterative Self-Attention Control

2024-06-10 · Xiaojie Li, Chenghao Gu, Shuzhao Xie, Yunpeng Bai 외

Fine-Tuning Diffusion Models enable a wide range of personalized generation and editing applications on diverse visual modalities. While Low-Rank Adaptation (LoRA) accelerates the fine-tuning process, it still requires m…

Denoising

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

2026-07-10 · Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park arxiv

Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted…

Virtual Try-onImage Editing

BlobCtrl: A Unified and Flexible Framework for Element-level Image Generation and Editing

2025-03-17 · Yaowei Li, Lingen Li, Zhaoyang Zhang, Xiaoyu Li 외

Element-level visual manipulation is essential in digital content creation, but current diffusion-based methods lack the precision and flexibility of traditional tools. In this work, we introduce BlobCtrl, a framework th…

Computational EfficiencyData AugmentationDiversityImage Generation